NEWVectors or files. Pick a path.Start →
    Models/Speech & Audio/nvidia/parakeet-tdt-0.6b-v3
    NeMoTranscriptioncc-by-4.0

    parakeet-tdt-0.6b-v3

    by nvidia

    600M multilingual ASR with 25-language support and automatic language detection

    730Kdl/month
    1,070likes
    627Mparams
    Identifiers
    Model ID
    nvidia/parakeet-tdt-0.6b-v3
    Feature URI
    mixpeek://transcription@v1/nvidia_parakeet_tdt_v3

    Overview

    Parakeet TDT 0.6B v3 is NVIDIA's multilingual speech-to-text model built on the FastConformer-TDT architecture and trained on over 670,000 hours of audio from NVIDIA's Granary dataset. It extends the English-only v2 to 25 European languages with automatic language detection, achieving a 6.34% average WER on the HuggingFace Open ASR Leaderboard while maintaining among the highest throughput of any multilingual model.

    On Mixpeek, Parakeet TDT powers cost-efficient multilingual transcription pipelines where Whisper-class accuracy is needed at lower compute cost. Its 600M parameter count and FastConformer architecture deliver excellent throughput for batch processing large audio and video archives across European languages.

    Architecture

    FastConformer encoder with Token-and-Duration Transducer (TDT) decoder. 600M parameters. Uses a unified SentencePiece tokenizer with 8,192-token vocabulary. Supports audio up to 3 hours via local attention mode. Automatic language identification across 25 languages.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so parakeet-tdt-0.6b-v3 runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The model produces text, so it lands in payload. Give the
              // collection a text vector index and embed that text to make it
              // searchable rather than only filterable.
              payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
              vectors: { "text-embedding": embeddingOfModelOutput },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // universal_extractor@v1 runs google/gemini-embedding-2
    // (3072-d) over a bucket, with no inference of your own.

    Capabilities

    • 25 European languages with automatic detection
    • 1.93% WER on LibriSpeech test-clean
    • 6.34% average WER on Open ASR Leaderboard
    • Audio up to 3 hours via local attention mode
    • Word-level timestamps included

    Use Cases on Mixpeek

    Multilingual video transcription for European content libraries at scale
    Batch audio processing of podcasts and meetings across 25 languages
    Cost-efficient ASR pipeline replacing Whisper for European language content

    Benchmarks

    DatasetMetricScoreSource
    LibriSpeech test-cleanWER1.93%NVIDIA, 2025: Model Card
    LibriSpeech test-otherWER3.59%NVIDIA, 2025: Model Card
    Open ASR Leaderboard (avg)WER6.34%NVIDIA, 2025: Model Card

    Performance

    Input SizeVariable-length audio (up to 3 hours)
    GPU Latency~3s / minute of audio (A100)
    GPU Throughput~20x realtime (A100)
    GPU Memory~2.5 GB

    Specification

    FrameworkNeMo
    Organizationnvidia
    FeatureTranscription
    Outputtext + timestamps
    Modalitiesvideo, audio
    RetrieverTranscript Search
    Parameters627M
    Licensecc-by-4.0
    Downloads/mo730K
    Likes1,070

    Research Paper

    Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient Multilingual ASR

    arxiv.org

    Build a pipeline with parakeet-tdt-0.6b-v3

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free