nccl v2.32.3-1
Permanent link:
cppdashboard.dev/r/2026/09/nccl-v2-32-3-1Optimized primitives for collective multi-GPU communication
Release notes
## Vera Rubin Support - Adds initial Rubin platform support including support for sm107, CX9 rail and plane detection, and MPS+MLoPart. - 2.32.3 focuses on new functionalities and does not contain performance model tuning for Rubin. This will be part of the next release to improve out-of-box experience for Rubin users. ## Device API Enhancements - Adds Compute Fabric Transport (CFT) counted-write and wait support. - Adds socket-based GIN support, enabling custom kernel development with TCP sockets. - Adds support in GDAKI for LAG-aware QP assignment based on the context ID. ([Github PR #2315](https://github.com/NVIDIA/nccl/pull/2315)) - Adds NCCL_WIN_REGISTER_GIN so users can register a window only for GIN usage. - Optimize GIN performance by skipping mcst operation when possible. ## Collectives and Runtime Enhancements - Adds a ring-based hierarchical copy-engine AllGather implementation selectable with NCCL_HIER_CE_COLL_AG_RAIL_RING_ENABLE. ([Github PR #2299](https://github.com/NVIDIA/nccl/pull/2299)) - Improves Blackwell symmetric AllGather performance, resource overhead modeling and kernel selection with a new cost model. - Adds optional TLS encryption for NCCL-owned socket traffic when NCCL is built with OpenSSL3, configured through the ncclSetEncryption API. - Adds ncclCollConfig_t::launchCompletionEvent, allowing callers to observe kernel-launch completion. - Adds ncclNvlsHostMode_t to ncclConfig_t so applications can disable host NVLS collectives per communicator. - Optimizes NVLS slot consumption to avoid resource exhaustion issues when using multiple communicators with NVLS. ## Diagnostics and Profiling - Adds the ATTN log level for important non-fatal conditions, such as configuration fallbacks and plugin initialization failures. - Expands RAS capability with GPU-resident progress counters and a watchdog DMA mirror for diagnosing stalled collective kernels. - Adds additional RAS diagnostics functionalities for NVLink/NIC state and speed, PCI and GDR configuration, Xid/SXid events and others. - Reports degraded NVLink fabric bandwidth with an ATTN message during communicator initialization. - Reports mismatched NCCL Git revisions across communicator ranks during initialization. ## Other Improvements - Fixes PAT+NVLS performance drops on H100 platforms. - Avoid address space exhaustion when repeatedly registering symmetric windows backed by a single physical memory allocation. - Reduces communicator initialization overhead at large scale by avoiding scans of inactive proxy poll descriptors. - Adds an experimental built-in NetworkDirect transport on Windows, with automatic socket fallback. - EFA team improved EFA GDA support with gin.get API and others. ([Github PR #2382](https://github.com/NVIDIA/nccl/pull/2382))([Github PR #2398](https://github.com/NVIDIA/nccl/pull/2398))([Github PR #2405](https://github.com/NVIDIA/nccl/pull/2405)) - Reports Inspector ring-buffer drops and operation counts in JSON. ([Github PR #2304](https://github.com/NVIDIA/nccl/pull/2304)) - NCCL now supports communicators using multiple MIG instances. - Generates llms.txt for agents to better read NCCL documentation. ## Bug Fixes - Fixes profiler overhead when no profiler plugin is loaded. ([Github Issue #2355](https://github.com/NVIDIA/nccl/issues/2355)) - Fixes P2P IPC registration reuse producing out-of-bounds remote addresses when a registered allocation spans multiple cuMem segments. ([Github PR #2362](https://github.com/NVIDIA/nccl/pull/2362)) - Fixes PAT connection-setup deadlocks when runtime connection is disabled and nodes have uneven local-rank counts. ([Github Issue #2385](https://github.com/NVIDIA/nccl/issues/2385)) - Fixes non-thread-safe token parsing that could corrupt concurrent configuration parsing. ([Github Issue #2361](https://github.com/NVIDIA/nccl/issues/2361)) - Fixes B40 TMA symmetric kernels…
Share this resource