Nvidia releases free 100M‑parameter Nemotron 3 Diarization model, identifies up to eight speakers in real time
The new open‑weight model hits a 14.7% diarization error rate and beats its predecessor by roughly 41%, according to the VoiceArena benchmark.
What Nvidia announced
On September 27, 2026, Nvidia unveiled Nemotron 3 Diarization, a speaker‑segmentation model whose weights are publicly available. The 100‑million‑parameter system can distinguish as many as eight concurrent speakers and flag moments when multiple voices overlap.
Model specs and benchmark performance
Nemotron 3 processes both recorded and live audio streams, assigning generic labels such as “speaker_2” to each voice. It can be combined with speech‑recognition engines like Parakeet to produce transcriptions that include speaker tags. The model supports four audio‑buffer configurations, ranging from 30.4 seconds down to 0.32 seconds; shorter buffers generally lower accuracy.
In the VoiceArena Diarization Benchmark v1, Nemotron 3 achieved a diarization error rate (DER) of 14.7%, placing it at the top of the leaderboard. On the stricter Diarization‑Bench test, it recorded a 14.72% error rate, outpacing the runner‑up’s 19.3%. The benchmark penalizes overlapping speech and even minor misalignments at speaker switches. Compared with Nvidia’s earlier Streaming Sortformer, Nemotron 3 reduces error by an average of 41% across eight test scenarios when using a 1.04‑second buffer.
Implications for speech AI
By offering a free, high‑performing diarization model, Nvidia lowers the barrier for developers building multi‑speaker transcription pipelines. The ability to handle up to eight speakers and detect overlaps makes it suitable for meetings, podcasts, and other multi‑party audio contexts, especially when paired with existing recognizers. The strong benchmark results suggest a shift toward more accurate, real‑time speaker‑tracking solutions in the broader AI‑speech ecosystem.



