Library

ggml v0.24.0

Versionv0.24.0
Stars★ 15,358
Released2026-09-14

Tensor library for machine learning

Release notes

## Overview

This release focuses on expanding backend coverage and robustness, with a new precision-control API, major Vulkan/SYCL/Hexagon/OpenCL work, and numerous correctness and performance fixes across CPU, CUDA, Metal, and other backends.

### API changes

- Expanded `ggml_prec` with `GGML_PREC_BF16`, `F16`, `Q8`, and `Q4`, and deprecated `GGML_PREC_DEFAULT` ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)).
- Added `ggml_prec_set_acc()` and `ggml_prec_set_src()` to control accumulator and per-source precision for `MUL_MAT`/`MUL_MAT_ID` and flash attention ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)).
- Deprecated `ggml_mul_mat_set_prec()` and `ggml_flash_attn_ext_set_prec()` in favor of the new precision API ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)).

### Core changes

- Added precision op-params layout for `MUL_MAT` and `MUL_MAT_ID` in `ggml-impl.h` ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)).
- Backend scheduler no longer forces an extra split for backend inputs and skips zero-sized MoE `ids` tensors ([llama/28387](https://github.com/ggml-org/llama.cpp/pull/28387), [llama/28739](https://github.com/ggml-org/llama.cpp/pull/28739)).
- Added PCH/unity-build support to speed up builds ([llama/28091](https://github.com/ggml-org/llama.cpp/pull/28091)).
- Fixed MSVC+Clang `ggml_vld1q_u32` and added a missing header for `gguf.cpp` ([llama/28284](https://github.com/ggml-org/llama.cpp/pull/28284), [llama/28566](https://github.com/ggml-org/llama.cpp/pull/28566)).

### Backend changes

#### CPU

- Disabled PCH and fixed `CACHE_LINE_SIZE` ambiguity causing heap corruption ([llama/28882](https://github.com/ggml-org/llama.cpp/pull/28882)).
- Added s390x repack support for `q4_0` and Q1_0 vector intrinsics; guarded VXE-only helpers and added non-VXE tests ([llama/28667](https://github.com/ggml-org/llama.cpp/pull/28667), [llama/28606](https://github.com/ggml-org/llama.cpp/pull/28606), [llama/28775](https://github.com/ggml-org/llama.cpp/pull/28775), [llama/28776](https://github.com/ggml-org/llama.cpp/pull/28776)).
- Added PCH support to `ggml-cpu` ([llama/28091](https://github.com/ggml-org/llama.cpp/pull/28091)).
- Fixed MSVC+Clang `ggml_vld1q_u32` ([llama/28284](https://github.com/ggml-org/llama.cpp/pull/28284)).

#### CUDA / HIP / MUSA

- Added AMD GCN-specific MMQ config table and gfx90c HIP support ([llama/27841](https://github.com/ggml-org/llama.cpp/pull/27841), [llama/26454](https://github.com/ggml-org/llama.cpp/pull/26454)).
- Fall back to F32 on devices without BF16 hardware acceleration ([llama/28846](https://github.com/ggml-org/llama.cpp/pull/28846)).
- Reworked flash attention quant compile flags (`GGML_FA_QUANTS`) and tuned flash attention for gfx1201 ([llama/28079](https://github.com/ggml-org/llama.cpp/pull/28079), [llama/28102](https://github.com/ggml-org/llama.cpp/pull/28102)).
- Fixed divergent barrier in f16 flash attention and races in MoE MMID/MMF kernels ([llama/27870](https://github.com/ggml-org/llama.cpp/pull/27870), [llama/28475](https://github.com/ggml-org/llama.cpp/pull/28475)).
- Tuned MoE MMQ N-tile sizing for RDNA3 and added branchless Q4_K/Q5_K unpack with L2 prefetch ([llama/24546](https://github.com/ggml-org/llama.cpp/pull/24546), [llama/28552](https://github.com/ggml-org/llama.cpp/pull/28552), [llama/26705](https://github.com/ggml-org/llama.cpp/pull/26705)).
- Reverted HIP `prop.integrated` restore ([llama/28604](https://github.com/ggml-org/llama.cpp/pull/28604)).

#### Metal

- Reworked fusion into a single-source table with debug support ([llama/28164](https://github.com/ggml-org/llama.cpp/pull/28164)).
- Added flash attention vec tuning for M2 Max and M3 ([llama/28458](https://github.com/ggml-org/llama.cpp/pull/28458), [llama/28396](https://github.com/ggml-org/llama.cpp/pull/28396)).
- Fixed idle threads in IQ `mul_mv` kernels and skipped empty MoE token tiles ([llama/28692](https:/…

Share this resource


Discovered 2026-09-15 Source GitHub Archive 2026-09 →