NEWVectors or files. Pick a path.Start →
    Models/moondream/parakeet-redux
    CC-BY-4.0

    parakeet-redux

    by moondream

    A 178 MB, 1.58-bit Parakeet that transcribes 25 languages at 113 times real time on eight CPU cores

    214likes
    0.6B (1.58-bit encoder)params
    Identifiers
    Model ID
    moondream/parakeet-redux
    Feature URI

    Deploy parakeet-redux

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Parakeet Redux is a 1.58-bit version of NVIDIA's parakeet-tdt-0.6b-v3: every encoder weight is -1, 0 or +1, so the model fits in 178 MB. Moondream released it in September 2026 under CC-BY-4.0, alongside the full-precision Parakeet Ultra.

    The card reports 113 times real time on eight x86 CPU cores and stays within 0.3 points of the original's English word error rate (6.55% against 6.26%). It beats the original on the 25-language FLEURS set (10.56% against 11.62%) and on long talks, and does worst in background noise (9.04% against 6.72%).

    It returns segment and word timestamps. The scores are self-reported and the speeds were measured in Moondream's Photon runtime.

    Architecture

    A Parakeet TDT model: a FastConformer encoder feeds a transducer decoder that predicts each token with the number of frames it spans, which gives the timestamps. Redux keeps the original architecture and tokenizer and quantizes every encoder weight to three values, about 1.58 bits each. Moondream's Photon runtime reads the packed weights directly with AVX-512 VNNI on x86, NEON on ARM and Metal on Apple GPUs.

    Mixpeek SDK Integration

    # Transcribe with Parakeet Redux, then store each timestamped segment as a text object
    # so a Mixpeek text collection can embed it and a retriever can find the moment.
    import moondream as md
    import requests
    
    with md.photon("moondream/parakeet-redux", device="cpu") as speech:
        segments = speech.transcribe(audio="call.wav", timestamps="segment")["segments"]
    
    for seg in segments:
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={"blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
                  "metadata": {"recording": "s3://calls/2026-10-05-acme.wav", "start_s": seg["start"], "end_s": seg["end"]}},
        )

    Capabilities

    • Speech to text in 25 European languages, including English
    • Segment and word timestamps
    • Runs on CPUs (x86 and ARM) and Apple silicon through Moondream's Photon runtime
    • CC-BY-4.0, same as the NVIDIA original

    Use Cases on Mixpeek

    Transcribing calls and meetings on ordinary CPU servers or laptops
    Timestamped transcripts for an audio or video archive, without a GPU
    On-device transcription where download size matters

    Benchmarks

    DatasetMetricScoreSource
    Open ASR Leaderboard, 7 English setsWER (lower is better)6.55%Model card: moondream/parakeet-redux (self-reported; original 6.26%)
    FLEURS, 25 languagesWER10.56%Model card (self-reported; original 11.62%)
    TED-LIUM long-formWER2.51%Model card (self-reported; original 2.71%)
    LibriSpeech test-clean, 8 x86 coresReal-time factor113xModel card (self-reported, Photon; parakeet.cpp 45x)

    Performance

    Input SizeSpeech audio files
    Embedding Dimn/a (outputs text with segment and word timestamps)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    178 MB of weights against 1.2 GB for the original. The card's speed figures are one utterance at a time in Moondream's Photon runtime on an AMD EPYC 9575F. Noisy audio is its weak spot: 9.04% against 6.72% across nine MUSAN noise conditions. We have not measured it.

    Frequently Asked Questions

    Does Parakeet Redux need a GPU?

    No. It is built for CPUs and Apple silicon. On eight x86 cores the card reports 113 times real time, 2.5 times the fastest other Parakeet runtime it measured on the same machine.

    How much accuracy does Parakeet Redux give up?

    On the card's English sets it scores 6.55% word error against the original's 6.26%, and it does better on the 25-language FLEURS set (10.56% against 11.62%) and on long talks. It loses the most in background noise: 9.04% against 6.72%.

    What is the difference between Parakeet Redux and Parakeet Ultra?

    Redux is the 1.58-bit, 178 MB version for CPUs and Apple silicon. Ultra keeps full precision for GPUs and is more accurate than the original on every benchmark the card lists.

    How do I search Parakeet Redux transcripts with Mixpeek?

    Store each timestamped segment as a text object with the recording and its start time in metadata, then search it with a text retriever, as in the example on this page.

    Specification

    Organizationmoondream
    Retriever-
    Parameters0.6B (1.58-bit encoder)
    LicenseCC-BY-4.0
    Downloads/moN/A
    Likes214

    Research Paper

    Introducing Parakeet Redux and Ultra (Moondream)

    arxiv.org

    Build a pipeline with parakeet-redux

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free