NEWVectors or files. Pick a path.Start →
    Models/moondream/parakeet-ultra
    CC-BY-4.0

    parakeet-ultra

    by moondream

    A post-trained Parakeet 0.6B: lower word error in 25 languages, in noise and on long recordings

    Identifiers
    Model ID
    moondream/parakeet-ultra
    Feature URI

    Deploy parakeet-ultra

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Parakeet Ultra turns speech into text with timestamps in 25 European languages. It is Moondream's post-trained version of NVIDIA's parakeet-tdt-0.6b-v3, with the same architecture, tokenizer and 0.6B parameters, released in September 2026 under CC-BY-4.0.

    On the card's benchmarks it beats the original everywhere it was tested: 5.80% word error rate against 6.26% on the seven English sets of the Open ASR Leaderboard, 9.55% against 11.62% across 25 FLEURS languages, 5.82% against 6.72% with background noise, and 1.94% against 2.71% on full-length TED talks. On AMI meeting recordings it reports 9.77% against 10.86%.

    It returns segment and word timestamps and splits long recordings at pauses with its own voice-activity head. The scores are self-reported, measured in Moondream's Photon runtime.

    Architecture

    A Parakeet TDT (token-and-duration transducer) model: a FastConformer encoder feeding a transducer decoder that predicts each token together with how many frames it spans, which is where the timestamps come from. It keeps the original's 0.6B parameters and tokenizer and adds a small voice-activity head on the encoder's subsampler, which the Photon runtime uses to cut long audio into segments of at most 30 seconds. Post-training by Moondream improved accuracy across languages, noise conditions and long-form audio without changing the architecture.

    Mixpeek SDK Integration

    # Transcribe with the model yourself, then store each timestamped segment as a text
    # object so a Mixpeek text collection can embed it and a retriever can find the moment.
    import requests
    
    for seg in segments:  # [{"start": 12.4, "end": 18.9, "text": "..."}]
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={
                "blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
                "metadata": {"recording": "s3://calls/2026-09-29-acme.mp4",
                             "start_s": seg["start"], "end_s": seg["end"]},
            },
        )

    Capabilities

    • Speech to text in 25 European languages, including English
    • Segment and word timestamps
    • Built-in voice-activity head that splits long recordings at pauses
    • CC-BY-4.0 license, same as the NVIDIA original

    Use Cases on Mixpeek

    Transcribing meeting and call recordings so they can be searched by what was said
    Timestamped subtitles and transcripts for a video archive
    Multilingual European content where one model covers every language
    High-volume batch transcription where throughput per GPU matters

    Benchmarks

    DatasetMetricScoreSource
    Open ASR Leaderboard, 7 English setsWER (lower is better)5.80%Model card: moondream/parakeet-ultra (self-reported; original 6.26%)
    AMI meetingsWER9.77%Model card (self-reported; original 10.86%)
    FLEURS, 25 languagesWER9.55%Model card (self-reported; original 11.62%)
    TED-LIUM 3, 11 full talksWER1.94%Model card (self-reported; original 2.71%)

    Performance

    Input SizeSpeech audio of any length; split at pauses into segments of at most 30 seconds
    Embedding Dimn/a (outputs text with segment and word timestamps)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The card reports 9,743x real time on LibriSpeech test-clean on one NVIDIA B200 with 128 requests in flight, running in Moondream's Photon runtime. We have not measured it.

    Frequently Asked Questions

    How is Parakeet Ultra different from NVIDIA's parakeet-tdt-0.6b-v3?

    It is the same architecture, tokenizer and size, post-trained by Moondream. On the card's benchmarks it has a lower word error rate on every group: 5.80% against 6.26% on the Open ASR Leaderboard's English sets, 9.55% against 11.62% across 25 FLEURS languages, and 1.94% against 2.71% on long TED talks.

    Does Parakeet Ultra give word-level timestamps?

    Yes. Its transcribe call returns one segment per sentence with start and end times, and word-level start and end times when asked for. That is what lets a search result jump to the moment a phrase was said.

    Is Parakeet Ultra good for meeting recordings?

    The card reports 9.77% word error rate on AMI, a meeting-recording test set, and 8.48% on a cleaned version of it, both lower than the original model. Meetings with crosstalk remain harder than read speech for any model, so test on your own calls.

    How do I search transcripts from Parakeet Ultra with Mixpeek?

    Store each timestamped segment as a text object with the recording and its start time in metadata, and search it with a text retriever, as in the example on this page. To have Mixpeek transcribe the recordings itself, index them with its multimodal extractor instead.

    Specification

    Organizationmoondream
    Retriever-
    Parameters0.6B
    LicenseCC-BY-4.0
    Downloads/moN/A
    Likes53

    Research Paper

    Introducing Parakeet Redux and Ultra (Moondream)

    arxiv.org

    Build a pipeline with parakeet-ultra

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free