Nemotron-3-Diarization
by nvidia
Who spoke when, live or offline, for up to eight speakers, in a 100M-parameter model
nvidia/Nemotron-3-DiarizationDeploy Nemotron-3-Diarization
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Nemotron 3 Diarization works out who spoke when in a recording, for up to eight speakers, live or after the fact. NVIDIA released it on 23 September 2026 under the OpenMDW 1.1 license, which allows commercial use. It has about 100M parameters.
On DIHARD III, an 11-domain benchmark, the card reports a 12.73% diarization error rate offline, against 19.09% for NVIDIA's previous streaming Sortformer. With five to nine speakers the error is 27.58%, against 40.21%. On CALLHOME telephone calls it reports 9.10%.
One checkpoint covers streaming from 0.32 seconds of input latency up to an offline 30.4 second buffer, and chunked inference removes any length limit. It labels speakers as speaker 1, speaker 2 and so on; it does not identify people. The scores are self-reported.
Architecture
A 31-layer Transformer encoder with rotary position embeddings reads 10 ms mel-spectrogram features stacked down to 80 ms frames, and a Conv1D layer upsamples its predictions back to 10 ms. The output is a probability of activity for each of eight speaker channels per frame. Following Sortformer, channels are ordered by when each speaker first appears, which resolves the label-permutation problem. For streaming it keeps an arrival-order speaker cache and a FIFO queue of recent frames, so speaker identities hold across chunks. Training started from a NEST self-supervised checkpoint and used about 10,000 hours of real conversations plus 82,611 hours of simulated multi-speaker mixtures.
Mixpeek SDK Integration
# Run diarization yourself, then store each speaker turn as a text object (with the
# transcript for that span) so a retriever can find who said what, and when.
import requests
for turn in turns: # [{"speaker": "speaker_1", "start": 12.4, "end": 18.9, "text": "..."}]
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={
"blobs": [{"property": "transcript", "type": "text", "data": turn["text"]}],
"metadata": {"recording": "s3://calls/2026-09-29-acme.mp4", "speaker": turn["speaker"],
"start_s": turn["start"], "end_s": turn["end"]},
},
)Capabilities
- Speaker diarization for up to eight speakers, with overlapping speech
- Streaming with input latency from 0.32 s, or offline with a 30.4 s buffer
- Recordings of any length through chunked inference
- Output resolution configurable in 10 ms steps; runs in NeMo, Transformers or NeMo-Speech.cpp
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| DIHARD III eval, full | DER (lower is better) | 12.73% | Model card: nvidia/Nemotron-3-Diarization (self-reported, offline; previous Sortformer 4spk-v2.1: 19.09%) |
| DIHARD III eval, 5-9 speakers | DER | 27.58% | Model card (self-reported, offline; previous: 40.21%) |
| CALLHOME Part 2, full | DER | 9.10% | Model card (self-reported, offline; previous: 10.32%) |
| DIHARD III eval, full, 0.32 s latency | DER | 13.55% | Model card (self-reported, streaming) |
Performance
About 100M parameters, optimized for NVIDIA GPUs. We have not measured speed.
Common Pipeline Companions
Frequently Asked Questions
What is speaker diarization?
Splitting a recording by who is speaking: the output says speaker 1 talked from 0.5 to 12.6 seconds, speaker 2 from 12.6 to 20.1, and so on. It does not name the people. Paired with a transcript, it turns one block of text into a record of who said what.
How many speakers can Nemotron 3 Diarization handle?
Up to eight in one recording, including overlapping speech. Accuracy drops as the count rises: on DIHARD III the card reports 9.13% diarization error for recordings with one to four speakers and 27.58% for five to nine.
Can Nemotron 3 Diarization run in real time?
Yes. One checkpoint runs streaming with an input buffer as short as 0.32 seconds, or offline with a 30.4 second buffer. On DIHARD III the streaming setting at 0.32 seconds scores 13.55% error against 12.73% offline.
How do I search speaker-labelled recordings with Mixpeek?
Store each speaker turn as a text object with its transcript, the speaker label and the start time in metadata, then search with a text retriever and filter by speaker, as in the example on this page.
Specification
Research Paper
Nemotron 3 Diarization (Hugging Face blog)
arxiv.orgBuild a pipeline with Nemotron-3-Diarization
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free