ggml v0.26.0
Permanent link:
cppdashboard.dev/r/2026/10/ggml-v0-26-0Tensor library for machine learning
Release notes
## Overview ggml v0.26.0 brings a new `alloc_buffer_n`/`get_alloc_size_n` buffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes. ### API changes - Add `alloc_buffer_n` and `get_alloc_size_n` to the buffer type interface with public APIs `ggml_backend_buft_alloc_buffer_n`/`ggml_backend_buft_get_alloc_size_n` for multi-buffer allocation and size planning ([llama/23671](https://github.com/ggml-org/llama.cpp/pull/23671)) - Add new public header `include/ggml-zdnn.h` for the zDNN backend ([llama/29541](https://github.com/ggml-org/llama.cpp/pull/29541)) ### Core changes - Speed up model loading and fix a hang on crafted GGUFs with very large KV dimensions ([llama/29598](https://github.com/ggml-org/llama.cpp/pull/29598)) - Harden tensor/GGUF size validation: fix integer overflows and reject tensor sizes that wrap after padding ([llama/29384](https://github.com/ggml-org/llama.cpp/pull/29384), [llama/26979](https://github.com/ggml-org/llama.cpp/pull/26979)) - Collect all input tensors into `graph_inputs` up front and require them to be `GGML_OP_NONE`, fixing spurious graph re-reserves with pipeline parallelism ([llama/29634](https://github.com/ggml-org/llama.cpp/pull/29634), [llama/29647](https://github.com/ggml-org/llama.cpp/pull/29647)) - Meta backend: clear inactive AllReduce shards with FILL instead of SCALE ([llama/29793](https://github.com/ggml-org/llama.cpp/pull/29793)) - ggml-quants: avoid invalid rounding in the qkx3 scale search ([llama/29817](https://github.com/ggml-org/llama.cpp/pull/29817)) - Enable Windows ARM64 builds with MSVC cl.exe ([llama/28362](https://github.com/ggml-org/llama.cpp/pull/28362)) ### Backend changes #### CPU - Support BF16/FP16/FP32 K tails in tinyBLAS on x86 ([llama/29806](https://github.com/ggml-org/llama.cpp/pull/29806)) - BF16 op support: add unary, GLU, binary and scale ops (with CUDA) and accept BF16 in src1 of mul_mat ([llama/29675](https://github.com/ggml-org/llama.cpp/pull/29675), [llama/28937](https://github.com/ggml-org/llama.cpp/pull/28937)) - Tiled mul_mat for k-quants via int8 unpack tiles and 16x16 microkernels ([llama/27851](https://github.com/ggml-org/llama.cpp/pull/27851)) - Enable tiled flash attention for non-vector-multiple head dims on x86 ([llama/29423](https://github.com/ggml-org/llama.cpp/pull/29423)) - Accumulate f16 dot products in f32 on AVX512-FP16 ([llama/29545](https://github.com/ggml-org/llama.cpp/pull/29545)) - Add Q8_0 IME1 matrix kernel for SpacemiT X60 ([llama/28479](https://github.com/ggml-org/llama.cpp/pull/28479)) - Fix soft_max_back wrong output when dst aliases src1 ([llama/27096](https://github.com/ggml-org/llama.cpp/pull/27096)) #### CUDA / HIP - Add model-driven W4A4 (NVFP4/MXFP4) mul_mat path with `llama_prec_policy` ([llama/24364](https://github.com/ggml-org/llama.cpp/pull/24364)) - NVFP4: optimize MMQ accumulation and handle the compute type on the cuBLAS path ([llama/29857](https://github.com/ggml-org/llama.cpp/pull/29857), [llama/29173](https://github.com/ggml-org/llama.cpp/pull/29173)) - Fuse shared experts into MMVQ ([llama/29184](https://github.com/ggml-org/llama.cpp/pull/29184)) - Use MMVF for thin f16/bf16 mul_mat at small batch size ([llama/29633](https://github.com/ggml-org/llama.cpp/pull/29633)) - FlashAttention: prefer whole-tile scheduling for two-stage kernels, tune fp16 tile configs for head sizes 40-112, fix 2 broken Volta cases ([llama/29435](https://github.com/ggml-org/llama.cpp/pull/29435), [llama/26289…
Share this resource