NEWVectors or files. Pick a path.Start →
    Models/FermionResearch/Phonon-2
    CC-BY-4.0

    Phonon-2

    by FermionResearch

    English speech-to-text in a 164 MB download, about 2-bit, that keeps the accuracy of its 2.5 GB Parakeet teacher

    141likes
    0.6B (about 2.1-bit encoder)params
    Identifiers
    Model ID
    FermionResearch/Phonon-2
    Feature URI

    Deploy Phonon-2

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Phonon-2 is an English speech-to-text model that fits in a 164 MB download. Fermion Research compressed NVIDIA's parakeet-tdt-0.6b-v3 with quantization-aware training so its encoder holds each weight at one of five learned levels, about 2.1 bits, and released it on 28 September 2026 under CC-BY-4.0.

    Across the Open ASR Leaderboard's seven English sets the card reports 5.21% average word error, against 4.96% for its 2.5 GB full-precision teacher and 6.58% for Whisper large-v3-turbo. It beats the teacher on AMI meetings (9.37% against 9.42%) and VoxPopuli (2.46% against 3.19%).

    It runs on Apple silicon, CPUs and GPUs, with word timestamps. It is English only, and the scores are self-reported.

    Architecture

    A Parakeet TDT (token-and-duration transducer) model, the architecture of NVIDIA's parakeet-tdt-0.6b-v3: a FastConformer encoder feeds a transducer decoder that predicts each token together with how many frames it spans, which is where word timestamps come from. Quantization-aware training holds each encoder weight at one of five learned levels, about 2.1 bits, which shrinks the download from 2.5 GB to 164 MB. The tokenizer and output conventions (punctuation, casing, numerals) are the original's. It runs through MLX on Apple silicon and through Fermion's engines on CPUs and CUDA.

    Mixpeek SDK Integration

    # Transcribe with Phonon-2 (phonon transcribe recording.wav --json gives word times),
    # then store each timestamped segment as a text object so a retriever can find the moment.
    import requests
    
    for seg in segments:  # [{"start": 12.4, "end": 18.9, "text": "..."}]
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={
                "blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
                "metadata": {"recording": "s3://calls/2026-10-01-acme.wav",
                             "start_s": seg["start"], "end_s": seg["end"]},
            },
        )

    Capabilities

    • English speech to text with punctuation, casing and numerals
    • Word-level start and end times with --json
    • Runs on Apple silicon (MLX), Linux and Windows CPUs, and NVIDIA GPUs via Docker
    • CC-BY-4.0 weights, same as the NVIDIA original; command line and code under Apache 2.0

    Use Cases on Mixpeek

    Transcribing meetings and calls on a laptop without sending audio anywhere
    Timestamped transcripts for a video archive so a search can jump to what was said
    High-volume batch transcription on CPUs where a GPU is not available
    Dictation and note-taking apps that need a small on-device model

    Benchmarks

    DatasetMetricScoreSource
    Open ASR Leaderboard, 7 English setsWER (lower is better)5.21%Model card: FermionResearch/Phonon-2 (self-reported; 2.5 GB teacher 4.96%)
    AMI meetingsWER9.37%Model card (self-reported; teacher 9.42%)
    VoxPopuliWER2.46%Model card (self-reported; teacher 3.19%)
    LibriSpeech test-cleanWER1.72%Model card (self-reported; teacher 1.52%)

    Performance

    Input SizeEnglish speech audio files
    Embedding Dimn/a (outputs text with word timestamps)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The card reports an hour of audio in about 20 seconds on an M5 MacBook Air (174x real time), 143x on eight Zen 5 cores and 6,680x on one H100 in batches of 128. The download is 164 MB. We have not measured it.

    Frequently Asked Questions

    How accurate is Phonon-2 compared with Whisper?

    On the Open ASR Leaderboard's seven English sets the card reports 5.21% average word error for Phonon-2 against 6.58% for Whisper large-v3-turbo, from a 164 MB download against 1,618 MB. On AMI meeting audio it reports 9.37% against 13.88%.

    Does Phonon-2 run without a GPU?

    Yes. It runs on Apple silicon through MLX, and on Linux and Windows CPUs; the card reports 143x real time on eight Zen 5 cores. A CUDA Docker image is available for NVIDIA GPUs.

    Which languages does Phonon-2 support?

    English. Its benchmarks are the Open ASR Leaderboard's English sets. For other languages, the NVIDIA model it is based on, parakeet-tdt-0.6b-v3, covers 25 European languages at full size.

    How do I search transcripts from Phonon-2 with Mixpeek?

    Store each timestamped segment as a text object with the recording and its start time in metadata, and search it with a text retriever, as in the example on this page. To have Mixpeek transcribe recordings itself, index them with its multimodal extractor instead.

    Specification

    OrganizationFermionResearch
    Retriever-
    Parameters0.6B (about 2.1-bit encoder)
    LicenseCC-BY-4.0
    Downloads/moN/A
    Likes141

    Research Paper

    Phonon speech docs (Fermion Research)

    arxiv.org

    Build a pipeline with Phonon-2

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free