onnxruntime v1.30.0
Permanent link:
cppdashboard.dev/r/2026/09/onnxruntime-v1-30-0ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Release notes
ONNX Runtime 1.30.0 expands generative AI inference, improves CPU and GPU performance, adds Go bindings, and strengthens runtime reliability. These notes cover changes since ONNX Runtime 1.29.1. ## Highlights - Expanded CUDA inference support with variable-length causal convolution for continuous batching, speculative decoding in paged XQA, and INT4 paged KV caches with per-channel scales ([#32168](https://github.com/microsoft/onnxruntime/pull/32168), [#32340](https://github.com/microsoft/onnxruntime/pull/32340), [#32515](https://github.com/microsoft/onnxruntime/pull/32515)). - Improved WebGPU PagedAttention, added GPT-OSS support and INT8 KV-cache block quantization, and extended convolution optimizations ([#31727](https://github.com/microsoft/onnxruntime/pull/31727), [#32277](https://github.com/microsoft/onnxruntime/pull/32277), [#32284](https://github.com/microsoft/onnxruntime/pull/32284), [#32420](https://github.com/microsoft/onnxruntime/pull/32420)). - Added fused CPU LinearAttention kernels for AVX-512, Arm64 NEON, and SVE, plus AVX2 LayerNorm/RMSNorm acceleration ([#31674](https://github.com/microsoft/onnxruntime/pull/31674), [#31973](https://github.com/microsoft/onnxruntime/pull/31973), [#32178](https://github.com/microsoft/onnxruntime/pull/32178), [#32356](https://github.com/microsoft/onnxruntime/pull/32356)). - Added Go bindings for the ONNX Runtime C API and DeepSeek Engram contrib operators ([#29615](https://github.com/microsoft/onnxruntime/pull/29615), [#32268](https://github.com/microsoft/onnxruntime/pull/32268)). ## Announcements & Compatibility - FP4 QMoE kernels are now enabled by default in CUDA builds, with Windows build support added in this release. Source builds can opt out with `-Donnxruntime_USE_FP4_QMOE=OFF` ([#32096](https://github.com/microsoft/onnxruntime/pull/32096), [#32163](https://github.com/microsoft/onnxruntime/pull/32163)). - CUDA fpA-intB builds now default to a compact kernel set for FP16 activations, INT4/INT8 weights, scale-only quantization, and `block_size=32`. Set `-Donnxruntime_USE_FPA_INTB_GEMM_FULL=ON` when building from source to retain the full kernel set, including BF16, zero-point, bias, larger-block-size, and native Hopper variants ([#32324](https://github.com/microsoft/onnxruntime/pull/32324)). - CPU FP16 `Gemm` and `MatMul` execution is gated on hardware acceleration. CPU-assigned FP16 nodes without a matching kernel now fall back to FP32 ([#32301](https://github.com/microsoft/onnxruntime/pull/32301), [#32197](https://github.com/microsoft/onnxruntime/pull/32197)). - WebGPU plugin EP packaging now supports Linux AArch64. Plugin versions were advanced to WebGPU 0.4.0 and CUDA 0.2 ([#32287](https://github.com/microsoft/onnxruntime/pull/32287), [#31960](https://github.com/microsoft/onnxruntime/pull/31960), [#31970](https://github.com/microsoft/onnxruntime/pull/31970)). ## Security & Reliability ### Model Loading, Memory, and Input Validation - Limited nested model-graph depth and canonicalized external-data locations to harden model loading ([#32344](https://github.com/microsoft/onnxruntime/pull/32344), [#32135](https://github.com/microsoft/onnxruntime/pull/32135)). - Added checked rounding for BFC arena allocations and fixed prepacked-weight reference lifetimes ([#32010](https://github.com/microsoft/onnxruntime/pull/32010), [#32040](https://github.com/microsoft/onnxruntime/pull/32040)). - Strengthened shape, rank, and parameter validation for `Split`, `Scan`, `GatherND`, `ScatterND`, `SpaceToDepth`/`DepthToSpace`, `Crop`, `Conv`, `Normalizer`, and pooling ([#29461](https://github.com/microsoft/onnxruntime/pull/29461), [#31668](https://github.com/microsoft/onnxruntime/pull/31668), [#32034](https://github.com/microsoft/onnxruntime/pull/32034), [#32039](https://github.com/microsoft/onnxruntime/pull/32039), [#32076](https://github.com/microsoft/onnxruntime/pull/32076), [#32157](https://github.com/microsoft/onnxruntime/pull/32157), [#32160](https://github…
Share this resource