parakeet-ultra
by moondream
A post-trained Parakeet 0.6B: lower word error in 25 languages, in noise and on long recordings
moondream/parakeet-ultraDeploy parakeet-ultra
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Parakeet Ultra turns speech into text with timestamps in 25 European languages. It is Moondream's post-trained version of NVIDIA's parakeet-tdt-0.6b-v3, with the same architecture, tokenizer and 0.6B parameters, released in September 2026 under CC-BY-4.0.
On the card's benchmarks it beats the original everywhere it was tested: 5.80% word error rate against 6.26% on the seven English sets of the Open ASR Leaderboard, 9.55% against 11.62% across 25 FLEURS languages, 5.82% against 6.72% with background noise, and 1.94% against 2.71% on full-length TED talks. On AMI meeting recordings it reports 9.77% against 10.86%.
It returns segment and word timestamps and splits long recordings at pauses with its own voice-activity head. The scores are self-reported, measured in Moondream's Photon runtime.
Architecture
A Parakeet TDT (token-and-duration transducer) model: a FastConformer encoder feeding a transducer decoder that predicts each token together with how many frames it spans, which is where the timestamps come from. It keeps the original's 0.6B parameters and tokenizer and adds a small voice-activity head on the encoder's subsampler, which the Photon runtime uses to cut long audio into segments of at most 30 seconds. Post-training by Moondream improved accuracy across languages, noise conditions and long-form audio without changing the architecture.
Mixpeek SDK Integration
# Transcribe with the model yourself, then store each timestamped segment as a text
# object so a Mixpeek text collection can embed it and a retriever can find the moment.
import requests
for seg in segments: # [{"start": 12.4, "end": 18.9, "text": "..."}]
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={
"blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
"metadata": {"recording": "s3://calls/2026-09-29-acme.mp4",
"start_s": seg["start"], "end_s": seg["end"]},
},
)Capabilities
- Speech to text in 25 European languages, including English
- Segment and word timestamps
- Built-in voice-activity head that splits long recordings at pauses
- CC-BY-4.0 license, same as the NVIDIA original
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Open ASR Leaderboard, 7 English sets | WER (lower is better) | 5.80% | Model card: moondream/parakeet-ultra (self-reported; original 6.26%) |
| AMI meetings | WER | 9.77% | Model card (self-reported; original 10.86%) |
| FLEURS, 25 languages | WER | 9.55% | Model card (self-reported; original 11.62%) |
| TED-LIUM 3, 11 full talks | WER | 1.94% | Model card (self-reported; original 2.71%) |
Performance
The card reports 9,743x real time on LibriSpeech test-clean on one NVIDIA B200 with 128 requests in flight, running in Moondream's Photon runtime. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
How is Parakeet Ultra different from NVIDIA's parakeet-tdt-0.6b-v3?
It is the same architecture, tokenizer and size, post-trained by Moondream. On the card's benchmarks it has a lower word error rate on every group: 5.80% against 6.26% on the Open ASR Leaderboard's English sets, 9.55% against 11.62% across 25 FLEURS languages, and 1.94% against 2.71% on long TED talks.
Does Parakeet Ultra give word-level timestamps?
Yes. Its transcribe call returns one segment per sentence with start and end times, and word-level start and end times when asked for. That is what lets a search result jump to the moment a phrase was said.
Is Parakeet Ultra good for meeting recordings?
The card reports 9.77% word error rate on AMI, a meeting-recording test set, and 8.48% on a cleaned version of it, both lower than the original model. Meetings with crosstalk remain harder than read speech for any model, so test on your own calls.
How do I search transcripts from Parakeet Ultra with Mixpeek?
Store each timestamped segment as a text object with the recording and its start time in metadata, and search it with a text retriever, as in the example on this page. To have Mixpeek transcribe the recordings itself, index them with its multimodal extractor instead.
Specification
Research Paper
Introducing Parakeet Redux and Ultra (Moondream)
arxiv.orgBuild a pipeline with parakeet-ultra
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free