NEWVectors or files. Pick a path.Start →
    Models/nvidia/Nemotron-3-Diarization
    OpenMDW-1.1

    Nemotron-3-Diarization

    by nvidia

    Who spoke when, live or offline, for up to eight speakers, in a 100M-parameter model

    Identifiers
    Model ID
    nvidia/Nemotron-3-Diarization
    Feature URI

    Deploy Nemotron-3-Diarization

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Nemotron 3 Diarization works out who spoke when in a recording, for up to eight speakers, live or after the fact. NVIDIA released it on 23 September 2026 under the OpenMDW 1.1 license, which allows commercial use. It has about 100M parameters.

    On DIHARD III, an 11-domain benchmark, the card reports a 12.73% diarization error rate offline, against 19.09% for NVIDIA's previous streaming Sortformer. With five to nine speakers the error is 27.58%, against 40.21%. On CALLHOME telephone calls it reports 9.10%.

    One checkpoint covers streaming from 0.32 seconds of input latency up to an offline 30.4 second buffer, and chunked inference removes any length limit. It labels speakers as speaker 1, speaker 2 and so on; it does not identify people. The scores are self-reported.

    Architecture

    A 31-layer Transformer encoder with rotary position embeddings reads 10 ms mel-spectrogram features stacked down to 80 ms frames, and a Conv1D layer upsamples its predictions back to 10 ms. The output is a probability of activity for each of eight speaker channels per frame. Following Sortformer, channels are ordered by when each speaker first appears, which resolves the label-permutation problem. For streaming it keeps an arrival-order speaker cache and a FIFO queue of recent frames, so speaker identities hold across chunks. Training started from a NEST self-supervised checkpoint and used about 10,000 hours of real conversations plus 82,611 hours of simulated multi-speaker mixtures.

    Mixpeek SDK Integration

    # Run diarization yourself, then store each speaker turn as a text object (with the
    # transcript for that span) so a retriever can find who said what, and when.
    import requests
    
    for turn in turns:  # [{"speaker": "speaker_1", "start": 12.4, "end": 18.9, "text": "..."}]
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={
                "blobs": [{"property": "transcript", "type": "text", "data": turn["text"]}],
                "metadata": {"recording": "s3://calls/2026-09-29-acme.mp4", "speaker": turn["speaker"],
                             "start_s": turn["start"], "end_s": turn["end"]},
            },
        )

    Capabilities

    • Speaker diarization for up to eight speakers, with overlapping speech
    • Streaming with input latency from 0.32 s, or offline with a 30.4 s buffer
    • Recordings of any length through chunked inference
    • Output resolution configurable in 10 ms steps; runs in NeMo, Transformers or NeMo-Speech.cpp

    Use Cases on Mixpeek

    Labelling meeting and call recordings by speaker before they are searched
    Live captions that show which person is talking
    Podcast and interview archives where each guest's answers must be found
    Speaker-tagged transcripts when paired with a speech recognizer

    Benchmarks

    DatasetMetricScoreSource
    DIHARD III eval, fullDER (lower is better)12.73%Model card: nvidia/Nemotron-3-Diarization (self-reported, offline; previous Sortformer 4spk-v2.1: 19.09%)
    DIHARD III eval, 5-9 speakersDER27.58%Model card (self-reported, offline; previous: 40.21%)
    CALLHOME Part 2, fullDER9.10%Model card (self-reported, offline; previous: 10.32%)
    DIHARD III eval, full, 0.32 s latencyDER13.55%Model card (self-reported, streaming)

    Performance

    Input Size16 kHz mono audio (.wav, .flac, .opus, .mp3); no length limit with chunked inference
    Embedding Dimn/a (outputs per-speaker activity over time, up to 8 speakers)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    About 100M parameters, optimized for NVIDIA GPUs. We have not measured speed.

    Frequently Asked Questions

    What is speaker diarization?

    Splitting a recording by who is speaking: the output says speaker 1 talked from 0.5 to 12.6 seconds, speaker 2 from 12.6 to 20.1, and so on. It does not name the people. Paired with a transcript, it turns one block of text into a record of who said what.

    How many speakers can Nemotron 3 Diarization handle?

    Up to eight in one recording, including overlapping speech. Accuracy drops as the count rises: on DIHARD III the card reports 9.13% diarization error for recordings with one to four speakers and 27.58% for five to nine.

    Can Nemotron 3 Diarization run in real time?

    Yes. One checkpoint runs streaming with an input buffer as short as 0.32 seconds, or offline with a 30.4 second buffer. On DIHARD III the streaming setting at 0.32 seconds scores 13.55% error against 12.73% offline.

    How do I search speaker-labelled recordings with Mixpeek?

    Store each speaker turn as a text object with its transcript, the speaker label and the start time in metadata, then search with a text retriever and filter by speaker, as in the example on this page.

    Specification

    Organizationnvidia
    Retriever-
    Parameters100M
    LicenseOpenMDW-1.1
    Downloads/moN/A
    Likes535

    Research Paper

    Nemotron 3 Diarization (Hugging Face blog)

    arxiv.org

    Build a pipeline with Nemotron-3-Diarization

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free