NVIDIA clarifies the parameters of the Nemotron 3 Diarization speaker recognition model
NVIDIA has published additional technical details on the open-weight model Nemotron 3 Diarization: according to the company, it is 41% more accurate than its predecessor and leads the Diarization-Bench leaderboard with an error rate of 14.72% (DER).
NVIDIA has added details to the previously released Nemotron 3 Diarization model, an open-weight model with roughly 100 million parameters designed to recognize who is speaking in a conversation and when. The model supports up to eight speakers, works with both live and recorded audio, and handles overlapping speech. On the Diarization-Bench leaderboard of the Voice Arena service, it took first place with an error rate (DER) of 14.72%, while the second-best system achieves 19.3%. According to NVIDIA, the model is on average 41% more accurate than its predecessor Streaming Sortformer (support for four speakers) across eight test scenarios when using a 1.04-second buffer.
The newly published details clarify the technical parameters: the audio buffer can be set to four levels ranging from 30.4 to 0.32 seconds, with a shorter buffer usually reducing accuracy. The error rate increases with a higher number of conversation participants, strong background noise, or reverberation. The model can be combined with the Parakeet speech recognition tool to create transcripts with speaker tags, but this involves only anonymous labels such as "speaker_2", not the identification of a specific person.
According to an earlier description, the model was trained on public and licensed data, including recordings of real multi-speaker conversations licensed from David AI, and on simulated multilingual mixtures covering 21 languages. Architecturally, it processes 16kHz audio using a Mel-spectrogram and a 31-layer Transformer with rotary positional encoding (RoPE), and uses AOSC memory mechanisms and a FIFO queue for streaming.
Why it matters
An open and freely available speaker diarization model lowers the barrier to deploying transcripts with speaker tags in applications such as call centers, meetings, or voice agents, without having to pay for a proprietary API — the cost is operating it yourself and integrating it with an ASR tool.
Release card
Nemotron 3 Diarization
NVIDIA
- Specifications
- about 100 million parameters
- Inputs
- input is audio, output is anonymous speaker timestamps
- Availability
- the model weights are freely downloadable
- Streamed speaker recognition in live and recorded conversations with up to eight speakers, including overlapping speech
- Combination with automatic speech recognition (ASR) into a pipeline for transcription with speaker highlighting
- Adjustable latency of streamed processing for various operational scenarios
- Does not determine the specific identity of a person, only anonymous speaker channels (e.g. speaker_2)
- Performance may decline for unusually long recordings or for audio with strong noise, reverberation, distant capture, or domain shift
According to the source, the model builds on the earlier NVIDIA Streaming Sortformer, which supported four speakers, and extends support to eight speakers with improved accuracy and throughput.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
What was added since the original report
Verified updates
-
41% better error rate than Streaming Sortformer (specific percentage); The model can be combined with Parakeet to create transcripts with speaker tags; Audio buffer adjustable to four levels (30.4–0.32 seconds); Error rate increases with more participants, strong noise, and hares of sound; The Voice Arena benchmark is strict – it counts overlapping speech and misalignments
- 41% better error rate than Streaming Sortformer (specific percentage)
- The model can be combined with Parakeet to create transcripts with speaker tags
- Audio buffer adjustable to four levels (30.4–0.32 seconds)
- Error rate increases with more participants, strong noise, and hares of sound
- The Voice Arena benchmark is strict – it counts overlapping speech and misalignments
Two audiences, two different impacts
What this means
For individuals
Developers working with audio transcripts gain a freely available real-time speaker diarization tool that can be combined with a speech recognition model and deployed without paying for a proprietary API.
For a business
Companies that process calls, meetings, or podcasts (transcription, conversation analytics, voice agents) can replace or supplement a proprietary speaker recognition solution with an open model with freely available weights, which reduces licensing costs, but requires their own deployment and infrastructure.
DevelopmentCheck the original
Event sources
independently confirmed · 2 publishers, 1 independent. We count feeds from the same owner only once.