Phonon-2
by FermionResearch
English speech-to-text in a 164 MB download, about 2-bit, that keeps the accuracy of its 2.5 GB Parakeet teacher
FermionResearch/Phonon-2Deploy Phonon-2
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Phonon-2 is an English speech-to-text model that fits in a 164 MB download. Fermion Research compressed NVIDIA's parakeet-tdt-0.6b-v3 with quantization-aware training so its encoder holds each weight at one of five learned levels, about 2.1 bits, and released it on 28 September 2026 under CC-BY-4.0.
Across the Open ASR Leaderboard's seven English sets the card reports 5.21% average word error, against 4.96% for its 2.5 GB full-precision teacher and 6.58% for Whisper large-v3-turbo. It beats the teacher on AMI meetings (9.37% against 9.42%) and VoxPopuli (2.46% against 3.19%).
It runs on Apple silicon, CPUs and GPUs, with word timestamps. It is English only, and the scores are self-reported.
Architecture
A Parakeet TDT (token-and-duration transducer) model, the architecture of NVIDIA's parakeet-tdt-0.6b-v3: a FastConformer encoder feeds a transducer decoder that predicts each token together with how many frames it spans, which is where word timestamps come from. Quantization-aware training holds each encoder weight at one of five learned levels, about 2.1 bits, which shrinks the download from 2.5 GB to 164 MB. The tokenizer and output conventions (punctuation, casing, numerals) are the original's. It runs through MLX on Apple silicon and through Fermion's engines on CPUs and CUDA.
Mixpeek SDK Integration
# Transcribe with Phonon-2 (phonon transcribe recording.wav --json gives word times),
# then store each timestamped segment as a text object so a retriever can find the moment.
import requests
for seg in segments: # [{"start": 12.4, "end": 18.9, "text": "..."}]
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={
"blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
"metadata": {"recording": "s3://calls/2026-10-01-acme.wav",
"start_s": seg["start"], "end_s": seg["end"]},
},
)Capabilities
- English speech to text with punctuation, casing and numerals
- Word-level start and end times with --json
- Runs on Apple silicon (MLX), Linux and Windows CPUs, and NVIDIA GPUs via Docker
- CC-BY-4.0 weights, same as the NVIDIA original; command line and code under Apache 2.0
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Open ASR Leaderboard, 7 English sets | WER (lower is better) | 5.21% | Model card: FermionResearch/Phonon-2 (self-reported; 2.5 GB teacher 4.96%) |
| AMI meetings | WER | 9.37% | Model card (self-reported; teacher 9.42%) |
| VoxPopuli | WER | 2.46% | Model card (self-reported; teacher 3.19%) |
| LibriSpeech test-clean | WER | 1.72% | Model card (self-reported; teacher 1.52%) |
Performance
The card reports an hour of audio in about 20 seconds on an M5 MacBook Air (174x real time), 143x on eight Zen 5 cores and 6,680x on one H100 in batches of 128. The download is 164 MB. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
How accurate is Phonon-2 compared with Whisper?
On the Open ASR Leaderboard's seven English sets the card reports 5.21% average word error for Phonon-2 against 6.58% for Whisper large-v3-turbo, from a 164 MB download against 1,618 MB. On AMI meeting audio it reports 9.37% against 13.88%.
Does Phonon-2 run without a GPU?
Yes. It runs on Apple silicon through MLX, and on Linux and Windows CPUs; the card reports 143x real time on eight Zen 5 cores. A CUDA Docker image is available for NVIDIA GPUs.
Which languages does Phonon-2 support?
English. Its benchmarks are the Open ASR Leaderboard's English sets. For other languages, the NVIDIA model it is based on, parakeet-tdt-0.6b-v3, covers 25 European languages at full size.
How do I search transcripts from Phonon-2 with Mixpeek?
Store each timestamped segment as a text object with the recording and its start time in metadata, and search it with a text retriever, as in the example on this page. To have Mixpeek transcribe recordings itself, index them with its multimodal extractor instead.
Specification
Research Paper
Phonon speech docs (Fermion Research)
arxiv.orgBuild a pipeline with Phonon-2
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free