ggml v0.24.0
Permanent link:
cppdashboard.dev/r/2026/09/ggml-v0-24-0Tensor library for machine learning
Release notes
## Overview This release focuses on expanding backend coverage and robustness, with a new precision-control API, major Vulkan/SYCL/Hexagon/OpenCL work, and numerous correctness and performance fixes across CPU, CUDA, Metal, and other backends. ### API changes - Expanded `ggml_prec` with `GGML_PREC_BF16`, `F16`, `Q8`, and `Q4`, and deprecated `GGML_PREC_DEFAULT` ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)). - Added `ggml_prec_set_acc()` and `ggml_prec_set_src()` to control accumulator and per-source precision for `MUL_MAT`/`MUL_MAT_ID` and flash attention ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)). - Deprecated `ggml_mul_mat_set_prec()` and `ggml_flash_attn_ext_set_prec()` in favor of the new precision API ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)). ### Core changes - Added precision op-params layout for `MUL_MAT` and `MUL_MAT_ID` in `ggml-impl.h` ([llama/26675](https://github.com/ggml-org/llama.cpp/pull/26675)). - Backend scheduler no longer forces an extra split for backend inputs and skips zero-sized MoE `ids` tensors ([llama/28387](https://github.com/ggml-org/llama.cpp/pull/28387), [llama/28739](https://github.com/ggml-org/llama.cpp/pull/28739)). - Added PCH/unity-build support to speed up builds ([llama/28091](https://github.com/ggml-org/llama.cpp/pull/28091)). - Fixed MSVC+Clang `ggml_vld1q_u32` and added a missing header for `gguf.cpp` ([llama/28284](https://github.com/ggml-org/llama.cpp/pull/28284), [llama/28566](https://github.com/ggml-org/llama.cpp/pull/28566)). ### Backend changes #### CPU - Disabled PCH and fixed `CACHE_LINE_SIZE` ambiguity causing heap corruption ([llama/28882](https://github.com/ggml-org/llama.cpp/pull/28882)). - Added s390x repack support for `q4_0` and Q1_0 vector intrinsics; guarded VXE-only helpers and added non-VXE tests ([llama/28667](https://github.com/ggml-org/llama.cpp/pull/28667), [llama/28606](https://github.com/ggml-org/llama.cpp/pull/28606), [llama/28775](https://github.com/ggml-org/llama.cpp/pull/28775), [llama/28776](https://github.com/ggml-org/llama.cpp/pull/28776)). - Added PCH support to `ggml-cpu` ([llama/28091](https://github.com/ggml-org/llama.cpp/pull/28091)). - Fixed MSVC+Clang `ggml_vld1q_u32` ([llama/28284](https://github.com/ggml-org/llama.cpp/pull/28284)). #### CUDA / HIP / MUSA - Added AMD GCN-specific MMQ config table and gfx90c HIP support ([llama/27841](https://github.com/ggml-org/llama.cpp/pull/27841), [llama/26454](https://github.com/ggml-org/llama.cpp/pull/26454)). - Fall back to F32 on devices without BF16 hardware acceleration ([llama/28846](https://github.com/ggml-org/llama.cpp/pull/28846)). - Reworked flash attention quant compile flags (`GGML_FA_QUANTS`) and tuned flash attention for gfx1201 ([llama/28079](https://github.com/ggml-org/llama.cpp/pull/28079), [llama/28102](https://github.com/ggml-org/llama.cpp/pull/28102)). - Fixed divergent barrier in f16 flash attention and races in MoE MMID/MMF kernels ([llama/27870](https://github.com/ggml-org/llama.cpp/pull/27870), [llama/28475](https://github.com/ggml-org/llama.cpp/pull/28475)). - Tuned MoE MMQ N-tile sizing for RDNA3 and added branchless Q4_K/Q5_K unpack with L2 prefetch ([llama/24546](https://github.com/ggml-org/llama.cpp/pull/24546), [llama/28552](https://github.com/ggml-org/llama.cpp/pull/28552), [llama/26705](https://github.com/ggml-org/llama.cpp/pull/26705)). - Reverted HIP `prop.integrated` restore ([llama/28604](https://github.com/ggml-org/llama.cpp/pull/28604)). #### Metal - Reworked fusion into a single-source table with debug support ([llama/28164](https://github.com/ggml-org/llama.cpp/pull/28164)). - Added flash attention vec tuning for M2 Max and M3 ([llama/28458](https://github.com/ggml-org/llama.cpp/pull/28458), [llama/28396](https://github.com/ggml-org/llama.cpp/pull/28396)). - Fixed idle threads in IQ `mul_mv` kernels and skipped empty MoE token tiles ([llama/28692](https:/…
Share this resource