parakeet-tdt-0.6b-v3
by nvidia
600M multilingual ASR with 25-language support and automatic language detection
nvidia/parakeet-tdt-0.6b-v3mixpeek://transcription@v1/nvidia_parakeet_tdt_v3Overview
Parakeet TDT 0.6B v3 is NVIDIA's multilingual speech-to-text model built on the FastConformer-TDT architecture and trained on over 670,000 hours of audio from NVIDIA's Granary dataset. It extends the English-only v2 to 25 European languages with automatic language detection, achieving a 6.34% average WER on the HuggingFace Open ASR Leaderboard while maintaining among the highest throughput of any multilingual model.
On Mixpeek, Parakeet TDT powers cost-efficient multilingual transcription pipelines where Whisper-class accuracy is needed at lower compute cost. Its 600M parameter count and FastConformer architecture deliver excellent throughput for batch processing large audio and video archives across European languages.
Architecture
FastConformer encoder with Token-and-Duration Transducer (TDT) decoder. 600M parameters. Uses a unified SentencePiece tokenizer with 8,192-token vocabulary. Supports audio up to 3 hours via local attention mode. Automatic language identification across 25 languages.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so parakeet-tdt-0.6b-v3 runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- 25 European languages with automatic detection
- 1.93% WER on LibriSpeech test-clean
- 6.34% average WER on Open ASR Leaderboard
- Audio up to 3 hours via local attention mode
- Word-level timestamps included
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| LibriSpeech test-clean | WER | 1.93% | NVIDIA, 2025: Model Card |
| LibriSpeech test-other | WER | 3.59% | NVIDIA, 2025: Model Card |
| Open ASR Leaderboard (avg) | WER | 6.34% | NVIDIA, 2025: Model Card |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient Multilingual ASR
arxiv.orgBuild a pipeline with parakeet-tdt-0.6b-v3
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free