Audio8-ASR-Infinite
by Edge0
Streaming Chinese and English speech recognition that runs on unlimited-length audio
Edge0/Audio8-ASR-InfiniteDeploy Audio8-ASR-Infinite
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Audio8-ASR-Infinite transcribes speech as it happens, in Chinese and English, and keeps going indefinitely. A rolling KV cache holds memory and latency constant, so the card says it runs 24/7 without drifting, and you choose the delay, from 240 to 560 ms, to trade speed for accuracy. Edge0 released it on 21 September 2026 under Apache-2.0 as a preview; it is a 4.1B-parameter model.
At a 480 ms delay the card reports 1.75% character error on AISHELL-1 and 2.89% on AISHELL-4, far below the streaming models it compares against, and 3.04% word error on LibriSpeech test-clean, where Voxtral Realtime does better at 2.21%.
It also ships semantic voice-activity heads that tell a thinking pause from the end of a turn. The scores are self-reported and the technical report has not been published yet.
Architecture
A Voxtral-style causal audio tower (32 layers, initialized from Voxtral Realtime 4B) feeds a projector into a Qwen2.5-3B-Instruct decoder that emits one text token per audio clock step, in the DSM streaming style. A frame-length embedding lets one checkpoint run at an 80, 120 or 160 ms clock, and the transcription delay is set in multiples of the clock. The native context is 30 seconds; a rolling KV window with exact RoPE re-basing extends that to unlimited length at constant memory. Separate semantic VAD heads classify end of turn at horizons of 0.5 to 3 seconds.
Mixpeek SDK Integration
# Transcribe with the model yourself, then store each timestamped segment as a text
# object so a Mixpeek text collection can embed it and a retriever can find the moment.
import requests
for seg in segments: # [{"start": 12.4, "end": 18.9, "text": "..."}]
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={
"blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
"metadata": {"recording": "s3://calls/2026-09-29-acme.mp4",
"start_s": seg["start"], "end_s": seg["end"]},
},
)Capabilities
- Native streaming transcription with a selectable 80, 120 or 160 ms audio clock
- Configurable delay (240 to 560 ms) to trade latency for accuracy
- Unlimited-length audio with constant memory, via a rolling 30-second KV cache
- Semantic end-of-turn detection; Chinese and English; Apache-2.0
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| AISHELL-1 test | CER (lower is better) | 1.750% | Model card: Edge0/Audio8-ASR-Infinite (self-reported, 480 ms delay, 80 ms clock) |
| AISHELL-4 test | CER | 2.893% | Model card (self-reported) |
| LibriSpeech test-clean | WER | 3.042% | Model card (self-reported; Voxtral Realtime reports 2.210%) |
| LibriSpeech test-other | WER | 6.808% | Model card (self-reported; Voxtral Realtime reports 5.552%) |
Performance
Weights are about 8.2 GB in bfloat16. The card's 24/7 path runs on its adapted vLLM build via Docker compose. We have not measured latency or throughput.
Common Pipeline Companions
Frequently Asked Questions
What makes Audio8-ASR-Infinite different from other speech recognition models?
It is built to run continuously. It transcribes as audio arrives, emitting up to 12.5 decisions per second, and a rolling 30-second KV cache keeps memory and latency constant, so it can run 24/7 without the output drifting. Most recognizers process fixed-length files instead.
Which languages does Audio8-ASR-Infinite support?
Chinese and English. On the card's table it is far ahead of the compared streaming models on Chinese (1.75% character error on AISHELL-1) and slightly behind Voxtral Realtime on English LibriSpeech.
What is semantic VAD?
Voice-activity detection that uses meaning as well as sound to tell a thinking pause or a stutter from the real end of a speaker's turn. Acoustic VAD only hears silence, so it often cuts people off mid-thought. Audio8 ships semantic VAD heads that predict the end of a turn at horizons from half a second to three seconds.
How do I search what Audio8-ASR-Infinite transcribed with Mixpeek?
Store the transcript in timestamped segments as text objects, with the source and start time in metadata, and search them with a text retriever, as in the example on this page. Mixpeek can also transcribe recordings itself with its multimodal extractor.
Specification
Research Paper
Audio8-ASR-Infinite on GitHub (technical report coming)
arxiv.orgBuild a pipeline with Audio8-ASR-Infinite
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free