Library

nccl v2.32.3-1

Versionv2.32.3-1
RepositoryNVIDIA/nccl ↗
Stars★ 5,100
Released2026-09-17

Optimized primitives for collective multi-GPU communication

Release notes

 ## Vera Rubin Support

  - Adds initial Rubin platform support including support for sm107, CX9 rail and plane detection, and MPS+MLoPart.
  - 2.32.3 focuses on new functionalities and does not contain performance model tuning for Rubin. This will be part of the next release to improve out-of-box experience for Rubin users.

  ## Device API Enhancements

  - Adds Compute Fabric Transport (CFT) counted-write and wait support.
  - Adds socket-based GIN support, enabling custom kernel development with TCP sockets.
  - Adds support in GDAKI for LAG-aware QP assignment based on the context ID. ([Github PR #2315](https://github.com/NVIDIA/nccl/pull/2315))
  - Adds NCCL_WIN_REGISTER_GIN so users can register a window only for GIN usage.
  - Optimize GIN performance by skipping mcst operation when possible.

  ## Collectives and Runtime Enhancements

  - Adds a ring-based hierarchical copy-engine AllGather implementation selectable with NCCL_HIER_CE_COLL_AG_RAIL_RING_ENABLE. ([Github PR #2299](https://github.com/NVIDIA/nccl/pull/2299))
  - Improves Blackwell symmetric AllGather performance, resource overhead modeling and kernel selection with a new cost model.
  - Adds optional TLS encryption for NCCL-owned socket traffic when NCCL is built with OpenSSL3, configured through the ncclSetEncryption API.
  - Adds ncclCollConfig_t::launchCompletionEvent, allowing callers to observe kernel-launch completion.
  - Adds ncclNvlsHostMode_t to ncclConfig_t so applications can disable host NVLS collectives per communicator.
  - Optimizes NVLS slot consumption to avoid resource exhaustion issues when using multiple communicators with NVLS.

  ## Diagnostics and Profiling

  - Adds the ATTN log level for important non-fatal conditions, such as configuration fallbacks and plugin initialization failures.
  - Expands RAS capability with GPU-resident progress counters and a watchdog DMA mirror for diagnosing stalled collective kernels.
  - Adds additional RAS diagnostics functionalities for NVLink/NIC state and speed, PCI and GDR configuration, Xid/SXid events and others.
  - Reports degraded NVLink fabric bandwidth with an ATTN message during communicator initialization.
  - Reports mismatched NCCL Git revisions across communicator ranks during initialization.

  ## Other Improvements

  - Fixes PAT+NVLS performance drops on H100 platforms.
  - Avoid address space exhaustion when repeatedly registering symmetric windows backed by a single physical memory allocation.
  - Reduces communicator initialization overhead at large scale by avoiding scans of inactive proxy poll descriptors.
  - Adds an experimental built-in NetworkDirect transport on Windows, with automatic socket fallback.
  - EFA team improved EFA GDA support with gin.get API and others. ([Github PR #2382](https://github.com/NVIDIA/nccl/pull/2382))([Github PR #2398](https://github.com/NVIDIA/nccl/pull/2398))([Github PR #2405](https://github.com/NVIDIA/nccl/pull/2405))
  - Reports Inspector ring-buffer drops and operation counts in JSON. ([Github PR #2304](https://github.com/NVIDIA/nccl/pull/2304))
  - NCCL now supports communicators using multiple MIG instances.
  - Generates  llms.txt for agents to better read NCCL documentation.

  ## Bug Fixes

  - Fixes profiler overhead when no profiler plugin is loaded. ([Github Issue #2355](https://github.com/NVIDIA/nccl/issues/2355))
  - Fixes P2P IPC registration reuse producing out-of-bounds remote addresses when a registered allocation spans multiple cuMem segments. ([Github PR #2362](https://github.com/NVIDIA/nccl/pull/2362))
  - Fixes PAT connection-setup deadlocks when runtime connection is disabled and nodes have uneven local-rank counts. ([Github Issue #2385](https://github.com/NVIDIA/nccl/issues/2385))
  - Fixes non-thread-safe token parsing that could corrupt concurrent configuration parsing. ([Github Issue #2361](https://github.com/NVIDIA/nccl/issues/2361))
  - Fixes B40 TMA symmetric kernels…

Share this resource


Discovered 2026-09-18 Source GitHub Archive 2026-09 →