AI RACE— The AI Race
New Models

NVIDIA Releases Nemotron 3 Diarization, an Open-Weight Model Tracking Up to Eight Speakers

NVIDIA has launched Nemotron 3 Diarization, a 100-million-parameter open-weight model that tracks up to eight overlapping speakers in real-time and recorded audio, securing the top spot on VoiceArena's Diarization-Bench.

09/23/2026, 20:17
New Models

NVIDIA Debuts Open-Weight Diarization Model

On September 23, 2026, NVIDIA researchers Francesco Ciannella, Maryam Motamedi, Taejin Park, Ivan Medennikov, and Adi Margolin introduced Nemotron 3 Diarization, an open-weight, 100-million-parameter speech model designed to identify who spoke when in multi-party conversations. Released under the OpenMDW License Agreement v1.1, the model is built for Linux environments running NVIDIA Ampere, Hopper, or Blackwell GPUs.

The system expands upon NVIDIA’s prior four-speaker Streaming Sortformer architecture by extending simultaneous tracking to eight anonymous speaker channels. In benchmark testing on VoiceArena’s initial Diarization-Bench—evaluating 139 English-language conversations spanning roughly 22 hours—Nemotron 3 Diarization placed first among 12 evaluated systems across 17 configurations. Scoring with overlapping speech, system-generated speech activity detection, and zero collar tolerance, it registered a 14.72% Diarization Error Rate (DER), outperforming the second-place system's 19.3% DER by a 24% relative margin.

Architecture, Benchmarks, and Low-Latency Performance

Nemotron 3 Diarization processes 16 kHz single-channel audio across .wav, .flac, .opus, and .mp3 formats. Incoming audio is converted into Mel-spectrogram features at 10 ms frame steps, which are stacked eightfold into 80 ms frames before passing to a 31-layer Transformer encoder equipped with Rotary Positional Embeddings (RoPE). A Conv1D layer upsamples the output back to 10 ms intervals, outputting a floating-point tensor that tracks speaker activity probabilities across eight channels simultaneously to preserve speech overlap.

To maintain consistent speaker IDs during live streaming without re-evaluating permutations for every incoming slice, the model uses a Sortformer-based arrival-order scheme. It retains past context via an Arrival-Order Speaker Cache (AOSC), a first-in, first-out (FIFO) queue for recent frames, and configurable lookahead right context. Developers can adjust input-buffer latency across four primary configurations: an offline-style 30.4 seconds, a low-latency 1.04 seconds, a very low latency 0.64 seconds, and an ultra-low latency 0.32 seconds (with 80 ms theoretically supported, though 0.32 seconds is the recommended floor).

Training incorporated public and commercial datasets, including multi-speaker audio licensed from David AI covering both natural speech and synthetic mixtures across 21 languages. Incorporating the David AI datasets reduced compound DER by 0.77 absolute points, lowering it from 11.19% to 10.42%.

Compared to NVIDIA’s previous baseline model (diar_streaming_sortformer_4spk-v2.1) across eight public benchmarks at a 1.04-second latency buffer, the new model averaged an unweighted 41.0% relative reduction in DER. Improvements ranged from 9.0% on CALLHOME-Part2 up to 65.2% on NOTSOFAR1 MHM. On the DIHARD III benchmark, DER fell from 19.60% to 13.18% at 1.04-second latency and from 19.09% to 12.73% at 30.4-second latency. Hardware throughput tests on an NVIDIA RTX PRO 5000 using BF16 precision and torch.compile() at batch size 32 showed processing speeds reaching an 865× real-time factor speedup (RTFx) at 1.04 seconds of latency and 15,113× RTFx in the 30.4-second offline configuration.

Ecosystem Integrations and Deployment Caveats

Because standalone diarization produces speaker timestamps rather than speech-to-text transcriptions, downstream systems must pair the model with an automatic speech recognition (ASR) engine—such as NVIDIA's Parakeet-TDT 0.6B v3 or Nemotron ASR 3.5—to produce attributed text. Ecosystem partners are already packaging the model for specialized workflows: Argmax integrated Nemotron 3 Diarization into its Argmax Pro SDK 3 to provide real-time on-device attribution and a pre-diarized transcription API, while cloud deployment workflows were highlighted for Baseten and DigitalOcean.

NVIDIA notes several operational boundaries for real-world production. Nemotron 3 Diarization caps attribution at eight speakers; conversations exceeding that limit will suffer from missed speech or misassigned speaker channels. Acoustic challenges like extreme room reverberation, far-field recording, background noise, or significant domain shift can also induce speaker confusion, false alarms, and timestamp boundary inaccuracies.

Related stories