NEWVectors or files. Pick a path.Start →
    Models/netease-youdao/Confucius4-R2T2
    NetEase Youdao Model Use License (separate commercial license above 100M monthly active users)

    Confucius4-R2T2

    by netease-youdao

    A 1.7B streaming speech recogniser whose text never changes once emitted, 2.13% WER on LibriSpeech clean at 160 ms chunks

    540likes
    1.7B (Qwen3-ASR base)params
    Identifiers
    Model ID
    netease-youdao/Confucius4-R2T2
    Feature URI

    Deploy Confucius4-R2T2

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Confucius4-R2T2 is a real-time speech recognition model from NetEase Youdao, built on Qwen3-ASR-1.7B and released in September 2026. It streams: you choose a decoding chunk from 80 ms to 2 s, and each piece of text it emits is final, so captions and voice agents never see words rewritten.

    At 160 ms chunks the card reports 2.13% WER on LibriSpeech clean and 9.36% on Earnings-22 calls, well ahead of the base model streaming at the same chunk size, with 200 to 600 ms average latency. It is optimised for Chinese and English. The license is free below 100 million monthly active users and forbids using the model to improve other AI models.

    Architecture

    Qwen3-ASR-1.7B trained for append-only streaming with stable-prefix data, forced time-alignment data and token-level supervision, so the decoder commits text as each chunk arrives. Decoding chunks are configurable from 80 ms to 2 s; vLLM serves it for throughput.

    Mixpeek SDK Integration

    # Transcribe with the repository's example runner at 160 ms chunks, then store
    # the transcript as a text object so a Mixpeek text collection indexes it.
    #   ./run_example.sh call-0412.wav --model_path ./Confucius4-R2T2 --infer_mode stream_vllm --chunk_size_ms 160
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
        headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
        json={"blobs": [{"property": "transcript", "type": "text", "data": transcript}],
              "metadata": {"recording": "s3://calls/call-0412.wav"}},
    )

    Capabilities

    • True streaming: emitted text is committed and never revised
    • Decoding chunks configurable from 80 ms to 2 s
    • Chinese and English optimized, with other languages supported
    • Context and hotword prompts
    • vLLM backend for throughput, plus a transformers backend

    Use Cases on Mixpeek

    Live captions and meeting transcripts that do not rewrite themselves
    Voice agents that act on words as they are spoken
    Streaming transcripts of calls or broadcasts, stored for later search

    Benchmarks

    DatasetMetricScoreSource
    LibriSpeech test-clean, 160 ms chunksWER (lower is better)2.13%Model card: netease-youdao/Confucius4-R2T2 (self-reported; Qwen3-ASR base at 160 ms 22.30, Nemotron 3.71)
    LibriSpeech test-other, 160 msWER4.88%Model card (self-reported; Nemotron 8.27)
    Earnings-22, 160 msWER9.36%Model card (self-reported; Nemotron 17.22)
    AMI meetings, 160 msWER11.37%Model card (self-reported; one proprietary system 8.44)
    WenetSpeech meeting (Chinese), 160 msCER7.27%Model card (self-reported; Qwen3-ASR base 20.38)

    Performance

    Input SizeAudio at any sample rate (resampled to 16 kHz), mono or stereo, streamed in 80 ms to 2 s chunks
    Embedding Dimn/a (outputs text)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The card reports 200 to 600 ms average latency with accuracy close to offline recognition, and no loss in offline accuracy from the streaming training. The license forbids using the model to improve other AI models. We have not measured it.

    Frequently Asked Questions

    What is Confucius4-R2T2?

    A streaming speech recognition model from NetEase Youdao that transcribes audio as it arrives, in chunks as short as 80 ms, without revising text it has already emitted.

    How accurate is Confucius4-R2T2 in real time?

    At 160 ms chunks the card reports 2.13% WER on LibriSpeech clean, 4.88% on LibriSpeech other and 9.36% on Earnings-22, with 200 to 600 ms average latency.

    Can I use Confucius4-R2T2 commercially?

    Under NetEase Youdao's model license, yes, unless your products had more than 100 million monthly active users in the previous month, which needs a separate commercial license. The license also forbids using it to improve other AI models.

    How do I search R2T2 transcripts with Mixpeek?

    Store each transcript as a text object with the recording in metadata, then search it with a text retriever, as in the example on this page.

    Specification

    Organizationnetease-youdao
    Retriever-
    Parameters1.7B (Qwen3-ASR base)
    LicenseNetEase Youdao Model Use License (separate commercial license above 100M monthly active users)
    Downloads/moN/A
    Likes540

    Research Paper

    Confucius4-R2T2 on GitHub

    arxiv.org

    Build a pipeline with Confucius4-R2T2

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free