Skip to content
worth noting New models verified update

NVIDIA clarifies the parameters of the Nemotron 3 Diarization speaker recognition model

independently confirmed updated September 27, 2026

NVIDIA has published additional technical details on the open-weight model Nemotron 3 Diarization: according to the company, it is 41% more accurate than its predecessor and leads the Diarization-Bench leaderboard with an error rate of 14.72% (DER).

NVIDIA has added details to the previously released Nemotron 3 Diarization model, an open-weight model with roughly 100 million parameters designed to recognize who is speaking in a conversation and when. The model supports up to eight speakers, works with both live and recorded audio, and handles overlapping speech. On the Diarization-Bench leaderboard of the Voice Arena service, it took first place with an error rate (DER) of 14.72%, while the second-best system achieves 19.3%. According to NVIDIA, the model is on average 41% more accurate than its predecessor Streaming Sortformer (support for four speakers) across eight test scenarios when using a 1.04-second buffer.

The newly published details clarify the technical parameters: the audio buffer can be set to four levels ranging from 30.4 to 0.32 seconds, with a shorter buffer usually reducing accuracy. The error rate increases with a higher number of conversation participants, strong background noise, or reverberation. The model can be combined with the Parakeet speech recognition tool to create transcripts with speaker tags, but this involves only anonymous labels such as "speaker_2", not the identification of a specific person.

According to an earlier description, the model was trained on public and licensed data, including recordings of real multi-speaker conversations licensed from David AI, and on simulated multilingual mixtures covering 21 languages. Architecturally, it processes 16kHz audio using a Mel-spectrogram and a 31-layer Transformer with rotary positional encoding (RoPE), and uses AOSC memory mechanisms and a FIFO queue for streaming.

What changed

Why it matters

An open and freely available speaker diarization model lowers the barrier to deploying transcripts with speaker tags in applications such as call centers, meetings, or voice agents, without having to pay for a proprietary API — the cost is operating it yourself and integrating it with an ASR tool.

Release card

Nemotron 3 Diarization

NVIDIA

open weights
Specifications
about 100 million parameters
Inputs
input is audio, output is anonymous speaker timestamps
Availability
the model weights are freely downloadable
According to the sources, it is suitable for
  • Streamed speaker recognition in live and recorded conversations with up to eight speakers, including overlapping speech
  • Combination with automatic speech recognition (ASR) into a pipeline for transcription with speaker highlighting
  • Adjustable latency of streamed processing for various operational scenarios
Documented limits
  • Does not determine the specific identity of a person, only anonymous speaker channels (e.g. speaker_2)
  • Performance may decline for unusually long recordings or for audio with strong noise, reverberation, distant capture, or domain shift

According to the source, the model builds on the earlier NVIDIA Streaming Sortformer, which supported four speakers, and extends support to eight speakers with improved accuracy and throughput.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

What was added since the original report

Verified updates

  1. New verified information

    41% better error rate than Streaming Sortformer (specific percentage); The model can be combined with Parakeet to create transcripts with speaker tags; Audio buffer adjustable to four levels (30.4–0.32 seconds); Error rate increases with more participants, strong noise, and hares of sound; The Voice Arena benchmark is strict – it counts overlapping speech and misalignments

    • 41% better error rate than Streaming Sortformer (specific percentage)
    • The model can be combined with Parakeet to create transcripts with speaker tags
    • Audio buffer adjustable to four levels (30.4–0.32 seconds)
    • Error rate increases with more participants, strong noise, and hares of sound
    • The Voice Arena benchmark is strict – it counts overlapping speech and misalignments

Two audiences, two different impacts

What this means

01

For individuals

Developers working with audio transcripts gain a freely available real-time speaker diarization tool that can be combined with a speech recognition model and deployed without paying for a proprietary API.

What to do Try the freely available weights of the Nemotron 3 Diarization model combined with an ASR tool such as Parakeet to create transcripts with speaker tags.
More practical updates →
02

For a business

Companies that process calls, meetings, or podcasts (transcription, conversation analytics, voice agents) can replace or supplement a proprietary speaker recognition solution with an open model with freely available weights, which reduces licensing costs, but requires their own deployment and infrastructure.

Development
What to decide Test the Nemotron 3 Diarization model on a sample of your own recordings (calls, meetings) and compare the error rate with the speaker recognition solution currently in use.
More business impacts →
Audio ML lokální inference Nemotron 3 Diarization NVIDIA NVIDIA Nemotron 3 Diarization Open-weight model open weights Speaker diarization Speech processing Voice Arena

Check the original

Event sources

independently confirmed · 2 publishers, 1 independent. We count feeds from the same owner only once.

2
Hugging Face Blog primary source · first detected **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** The Decoder (daily AI news) independent context Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time