Confucius4-R2T2
by netease-youdao
A 1.7B streaming speech recogniser whose text never changes once emitted, 2.13% WER on LibriSpeech clean at 160 ms chunks
netease-youdao/Confucius4-R2T2Deploy Confucius4-R2T2
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Confucius4-R2T2 is a real-time speech recognition model from NetEase Youdao, built on Qwen3-ASR-1.7B and released in September 2026. It streams: you choose a decoding chunk from 80 ms to 2 s, and each piece of text it emits is final, so captions and voice agents never see words rewritten.
At 160 ms chunks the card reports 2.13% WER on LibriSpeech clean and 9.36% on Earnings-22 calls, well ahead of the base model streaming at the same chunk size, with 200 to 600 ms average latency. It is optimised for Chinese and English. The license is free below 100 million monthly active users and forbids using the model to improve other AI models.
Architecture
Qwen3-ASR-1.7B trained for append-only streaming with stable-prefix data, forced time-alignment data and token-level supervision, so the decoder commits text as each chunk arrives. Decoding chunks are configurable from 80 ms to 2 s; vLLM serves it for throughput.
Mixpeek SDK Integration
# Transcribe with the repository's example runner at 160 ms chunks, then store
# the transcript as a text object so a Mixpeek text collection indexes it.
# ./run_example.sh call-0412.wav --model_path ./Confucius4-R2T2 --infer_mode stream_vllm --chunk_size_ms 160
import requests
requests.post(
"https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={"blobs": [{"property": "transcript", "type": "text", "data": transcript}],
"metadata": {"recording": "s3://calls/call-0412.wav"}},
)Capabilities
- True streaming: emitted text is committed and never revised
- Decoding chunks configurable from 80 ms to 2 s
- Chinese and English optimized, with other languages supported
- Context and hotword prompts
- vLLM backend for throughput, plus a transformers backend
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| LibriSpeech test-clean, 160 ms chunks | WER (lower is better) | 2.13% | Model card: netease-youdao/Confucius4-R2T2 (self-reported; Qwen3-ASR base at 160 ms 22.30, Nemotron 3.71) |
| LibriSpeech test-other, 160 ms | WER | 4.88% | Model card (self-reported; Nemotron 8.27) |
| Earnings-22, 160 ms | WER | 9.36% | Model card (self-reported; Nemotron 17.22) |
| AMI meetings, 160 ms | WER | 11.37% | Model card (self-reported; one proprietary system 8.44) |
| WenetSpeech meeting (Chinese), 160 ms | CER | 7.27% | Model card (self-reported; Qwen3-ASR base 20.38) |
Performance
The card reports 200 to 600 ms average latency with accuracy close to offline recognition, and no loss in offline accuracy from the streaming training. The license forbids using the model to improve other AI models. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What is Confucius4-R2T2?
A streaming speech recognition model from NetEase Youdao that transcribes audio as it arrives, in chunks as short as 80 ms, without revising text it has already emitted.
How accurate is Confucius4-R2T2 in real time?
At 160 ms chunks the card reports 2.13% WER on LibriSpeech clean, 4.88% on LibriSpeech other and 9.36% on Earnings-22, with 200 to 600 ms average latency.
Can I use Confucius4-R2T2 commercially?
Under NetEase Youdao's model license, yes, unless your products had more than 100 million monthly active users in the previous month, which needs a separate commercial license. The license also forbids using it to improve other AI models.
How do I search R2T2 transcripts with Mixpeek?
Store each transcript as a text object with the recording in metadata, then search it with a text retriever, as in the example on this page.
Specification
Research Paper
Confucius4-R2T2 on GitHub
arxiv.orgBuild a pipeline with Confucius4-R2T2
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free