NVIDIA clarifies the parameters of the Nemotron 3 Diarization speaker recognition model
NVIDIA has published additional technical details on the open-weight model Nemotron 3 Diarization: according to the company, it is 41% more accurate than its predecessor and leads the Diarization-Bench leaderboard with an error rate of 14.72% (DER).
New: 41% better error rate than Streaming Sortformer (specific percentage); The model can be combined with Parakeet to create transcripts with speaker tags; Audio buffer adjustable to four levels (30.4–0.32 seconds); Error rate increases with more participants, strong noise, and hares of sound; The Voice Arena benchmark is strict – it counts overlapping speech and misalignments