NEWVectors or files. Pick a path.Start →
    Models/Cactus-Compute/whistle
    Apache-2.0

    whistle

    by Cactus-Compute

    A 16.9 MB on-device speech recogniser for 7 languages with word timestamps, 4.31% WER on LibriSpeech clean

    129likes
    16.9 MB model file (2 to 4 bit)params
    Identifiers
    Model ID
    Cactus-Compute/whistle
    Feature URI

    Deploy whistle

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Whistle is a speech-to-text model from Cactus Compute, released in September 2026 under Apache-2.0, built for phones, wearables, robots, cars and microcontrollers. The whole model is one 16.9 MB file that runs on Cactus's CPU engine, Needle, with no network.

    It transcribes English, German, French, Spanish, Italian, Dutch and Polish, returns every word with its timing, and can also output a speech embedding for matching. On the card's tests it scores 4.31% WER on LibriSpeech test-clean and reaches its first token in 11.1 ms on a 10-second clip.

    Architecture

    A log-mel front end and a convolutional stem feed an audio encoder, which a Needle-style decoder reads through gated cross-attention at every layer. Word timestamps come from the decoder's own attention. Weights are quantized to 2 to 4 bits in Cactus's .cact container and run on SIMD CPU kernels.

    Mixpeek SDK Integration

    # Transcribe on the device, then store each clip's transcript as a text object
    # with its timing so a Mixpeek retriever can find the moment.
    import needle, requests
    
    out = needle.transcribe("clip.wav", word_timestamps=True)
    requests.post(
        "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
        headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
        json={"blobs": [{"property": "transcript", "type": "text", "data": out["text"]}],
              "metadata": {"recording": "s3://voice-notes/2026-10-07-0912.wav", "language": out.get("language")}},
    )

    Capabilities

    • Transcription of 16 kHz mono audio, up to 30 seconds per pass, in 7 languages
    • Word timestamps with a probability per word
    • Speech embeddings, one row per 80 ms frame, for matching without a transcript
    • Keyword biasing for names, places and product words
    • Runs on CPUs, phones, wearables and microcontrollers, and as WebAssembly

    Use Cases on Mixpeek

    Voice commands and dictation on phones, wearables and in cars, with no network
    Timestamped transcripts of short clips for search, captions or cutting on a word
    Matching spoken phrases by embedding where a transcript is not needed

    Benchmarks

    DatasetMetricScoreSource
    LibriSpeech test-cleanWER (lower is better)4.31%Model card: Cactus-Compute/whistle (self-reported, full test split, Whisper normalizers)
    LibriSpeech test-otherWER10.49%Model card (self-reported)
    TED-LIUMWER7.61%Model card (self-reported)
    FLEURS, 7 languagesWER21.4%Model card (self-reported)
    10 s clip, Apple M4 ProTime to first token11.1 msModel card (self-reported; Whisper base 73.2 ms, Moonshine tiny v2 22.8 ms)

    Performance

    Input Size16 kHz mono audio, up to 30 s per pass
    Embedding Dimn/a (text with word timestamps; optional speech embedding per 80 ms frame)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    16.9 MB against 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2, decoding 1,319 tokens a second on a 10 s clip on an M4 Pro (card figures). Longer recordings need to be cut into 30-second pieces. We have not measured it.

    Frequently Asked Questions

    Does Whistle need a GPU or an internet connection?

    No. It runs on CPUs, including phones, wearables and microcontrollers, and as a WebAssembly component, with no network calls.

    Which languages does Whistle support?

    English, German, French, Spanish, Italian, Dutch and Polish. It detects the language unless you name it.

    How does Whistle compare with Whisper?

    On the card's figures it is about a ninth of the size of Whisper base (16.9 MB against 145.3 MB) and reaches the first token about 6 times sooner on a 10-second clip, with 4.31% WER on LibriSpeech clean. The card uses Whisper's published error rates for the comparison.

    How do I search Whistle transcripts with Mixpeek?

    Store each transcript as a text object with the recording and its timing in metadata, then search it with a text retriever, as in the example on this page.

    Specification

    OrganizationCactus-Compute
    Retriever-
    Parameters16.9 MB model file (2 to 4 bit)
    LicenseApache-2.0
    Downloads/moN/A
    Likes129

    Research Paper

    Cactus-Compute/whistle model card

    arxiv.org

    Build a pipeline with whistle

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free