NEWVectors or files. Pick a path.Start →
    Models/Edge0/Audio8-ASR-Infinite
    Apache-2.0

    Audio8-ASR-Infinite

    by Edge0

    Streaming Chinese and English speech recognition that runs on unlimited-length audio

    Identifiers
    Model ID
    Edge0/Audio8-ASR-Infinite
    Feature URI

    Deploy Audio8-ASR-Infinite

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Audio8-ASR-Infinite transcribes speech as it happens, in Chinese and English, and keeps going indefinitely. A rolling KV cache holds memory and latency constant, so the card says it runs 24/7 without drifting, and you choose the delay, from 240 to 560 ms, to trade speed for accuracy. Edge0 released it on 21 September 2026 under Apache-2.0 as a preview; it is a 4.1B-parameter model.

    At a 480 ms delay the card reports 1.75% character error on AISHELL-1 and 2.89% on AISHELL-4, far below the streaming models it compares against, and 3.04% word error on LibriSpeech test-clean, where Voxtral Realtime does better at 2.21%.

    It also ships semantic voice-activity heads that tell a thinking pause from the end of a turn. The scores are self-reported and the technical report has not been published yet.

    Architecture

    A Voxtral-style causal audio tower (32 layers, initialized from Voxtral Realtime 4B) feeds a projector into a Qwen2.5-3B-Instruct decoder that emits one text token per audio clock step, in the DSM streaming style. A frame-length embedding lets one checkpoint run at an 80, 120 or 160 ms clock, and the transcription delay is set in multiples of the clock. The native context is 30 seconds; a rolling KV window with exact RoPE re-basing extends that to unlimited length at constant memory. Separate semantic VAD heads classify end of turn at horizons of 0.5 to 3 seconds.

    Mixpeek SDK Integration

    # Transcribe with the model yourself, then store each timestamped segment as a text
    # object so a Mixpeek text collection can embed it and a retriever can find the moment.
    import requests
    
    for seg in segments:  # [{"start": 12.4, "end": 18.9, "text": "..."}]
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={
                "blobs": [{"property": "transcript", "type": "text", "data": seg["text"]}],
                "metadata": {"recording": "s3://calls/2026-09-29-acme.mp4",
                             "start_s": seg["start"], "end_s": seg["end"]},
            },
        )

    Capabilities

    • Native streaming transcription with a selectable 80, 120 or 160 ms audio clock
    • Configurable delay (240 to 560 ms) to trade latency for accuracy
    • Unlimited-length audio with constant memory, via a rolling 30-second KV cache
    • Semantic end-of-turn detection; Chinese and English; Apache-2.0

    Use Cases on Mixpeek

    Live captions for calls, broadcasts and streams that run for hours
    Voice agents that need to know when a speaker has finished, not just paused
    Continuous transcription of camera or radio feeds for later search
    Chinese and English content where one streaming model covers both

    Benchmarks

    DatasetMetricScoreSource
    AISHELL-1 testCER (lower is better)1.750%Model card: Edge0/Audio8-ASR-Infinite (self-reported, 480 ms delay, 80 ms clock)
    AISHELL-4 testCER2.893%Model card (self-reported)
    LibriSpeech test-cleanWER3.042%Model card (self-reported; Voxtral Realtime reports 2.210%)
    LibriSpeech test-otherWER6.808%Model card (self-reported; Voxtral Realtime reports 5.552%)

    Performance

    Input Size16 kHz mono audio, streamed; no length limit
    Embedding Dimn/a (outputs text as it is spoken)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Weights are about 8.2 GB in bfloat16. The card's 24/7 path runs on its adapted vLLM build via Docker compose. We have not measured latency or throughput.

    Frequently Asked Questions

    What makes Audio8-ASR-Infinite different from other speech recognition models?

    It is built to run continuously. It transcribes as audio arrives, emitting up to 12.5 decisions per second, and a rolling 30-second KV cache keeps memory and latency constant, so it can run 24/7 without the output drifting. Most recognizers process fixed-length files instead.

    Which languages does Audio8-ASR-Infinite support?

    Chinese and English. On the card's table it is far ahead of the compared streaming models on Chinese (1.75% character error on AISHELL-1) and slightly behind Voxtral Realtime on English LibriSpeech.

    What is semantic VAD?

    Voice-activity detection that uses meaning as well as sound to tell a thinking pause or a stutter from the real end of a speaker's turn. Acoustic VAD only hears silence, so it often cuts people off mid-thought. Audio8 ships semantic VAD heads that predict the end of a turn at horizons from half a second to three seconds.

    How do I search what Audio8-ASR-Infinite transcribed with Mixpeek?

    Store the transcript in timestamped segments as text objects, with the source and start time in metadata, and search them with a text retriever, as in the example on this page. Mixpeek can also transcribe recordings itself with its multimodal extractor.

    Specification

    OrganizationEdge0
    Retriever-
    Parameters4.1B
    LicenseApache-2.0
    Downloads/moN/A
    Likes1,554

    Research Paper

    Audio8-ASR-Infinite on GitHub (technical report coming)

    arxiv.org

    Build a pipeline with Audio8-ASR-Infinite

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free