onnxruntime v1.31.0
Permanent link:
cppdashboard.dev/r/2026/10/onnxruntime-v1-31-0ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Release notes
ONNX Runtime 1.31.0 stabilizes model-package and EPContext data APIs, improves CPU model loading and quantized MoE inference, expands device-based execution-provider selection, and strengthens model-loading and runtime reliability. These notes cover changes since ONNX Runtime 1.30.0. CUDA and WebGPU kernel updates are summarized here; detailed provider notes are covered by their separate plugin EP releases. ## Highlights - Promoted the model-package API and EPContext data callbacks to stable C and C++ APIs, with EPContext callback support added across language bindings ([#33166](https://github.com/microsoft/onnxruntime/pull/33166), [#32265](https://github.com/microsoft/onnxruntime/pull/32265)). - Reduced CPU session-initialization overhead by parallelizing eligible weight prepacking, and accelerated block-wise INT4/INT8 QMoE experts with MLAS QNBit kernels and grouped expert dispatch ([#31691](https://github.com/microsoft/onnxruntime/pull/31691), [#32644](https://github.com/microsoft/onnxruntime/pull/32644), [#32668](https://github.com/microsoft/onnxruntime/pull/32668)). - Added x86 FP16 LayerNorm/RMSNorm acceleration, Arm KleidiAI SVE2.1 FP16 GEMM support, and additional RISC-V vector kernels ([#32715](https://github.com/microsoft/onnxruntime/pull/32715), [#32670](https://github.com/microsoft/onnxruntime/pull/32670), [#32710](https://github.com/microsoft/onnxruntime/pull/32710)). - Enabled CoreML participation in device-based EP selection, including Apple Neural Engine selection through `PREFER_NPU` ([#31975](https://github.com/microsoft/onnxruntime/pull/31975)). - Added workspace-memory accounting and verification for constrained-memory graph partitioning, alongside broader model, graph, and tensor validation ([#31962](https://github.com/microsoft/onnxruntime/pull/31962), [#32189](https://github.com/microsoft/onnxruntime/pull/32189)). ## Announcements & Compatibility - [CUDA Plugin EP v0.3.0](https://github.com/microsoft/onnxruntime/releases/tag/plugin-ep-cuda/v0.3.0) and WebGPU Plugin EP v0.5.0 are matched with ONNX Runtime 1.31.0. For **contrib-op compatibility**, core and the plugin must use the same contributed-operator schemas; building both from the **same ONNX Runtime revision** is recommended. Schema mismatches can cause incorrect execution or crashes and may not be detected during plugin registration. - **`GatherBlockQuantized` behavior change:** out-of-range indices now produce zero output slices across CPU, CUDA, and WebGPU for both integer and floating-point quantized formats. This deliberately differs from ONNX `Gather` error semantics ([#32480](https://github.com/microsoft/onnxruntime/pull/32480)). - **Experimental API migration:** model-package consumers should use `OrtApi::GetModelPackageApi` and the stable `OrtModelPackageApi` table. EPContext callback registration now uses stable APIs; the corresponding experimental lookup entries were removed. The stable native callback contract carries callback/state pointers, with payload limits enforced by application callbacks, bindings, or EPs rather than a native read-options policy ([#33166](https://github.com/microsoft/onnxruntime/pull/33166), [#32265](https://github.com/microsoft/onnxruntime/pull/32265)). - **ACL EP is deprecated** and will be removed in a future release. Source builds using `--use_acl` now receive a deprecation warning ([#32564](https://github.com/microsoft/onnxruntime/pull/32564)). - The ONNX dependency remains at **1.22.0**. The initial ONNX 1.23 integration was reverted pending upstream fixes; this release does not ship that integration's new opset or FLOAT6 support ([#33073](https://github.com/microsoft/onnxruntime/pull/33073)). - Supported native builds now default to **1DS telemetry**. Windows TraceLogging remains selectable with `--use_windows_telemetry` or `onnxruntime_USE_WINDOWS_TELEMETRY=ON`; telemetry can be disabled with `--no_telemetry` ([#32384](https://github.com/microsoft/onnxruntime/pull/32384)). ## New F…
Share this resource