NEWVectors or files. Pick a path.Start →
    Models/Cloudflare/clef
    Apache-2.0

    clef

    by Cloudflare

    A 27B multimodal decision model: ask typed questions about text, images or video and get a probability for every answer

    Identifiers
    Model ID
    Cloudflare/clef
    Feature URI

    Deploy clef

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Clef turns a situation and a set of typed questions into decisions. You describe the state as text, JSON, images or video, ask yes/no, pick-one or score questions, and it returns a probability for every allowed answer of every question in a single forward pass, with no free text to parse. Cloudflare released it on 30 September 2026 under Apache-2.0, post-trained from Qwen3.8-27B.

    On Cloudflare's own Decision Index run the card reports 94.2 macro-F1 on BANKING77 intent classification, 97.4 on CLINC150 with out-of-scope queries, and 79.4 hallucination F1 on RAGTruth, at a median latency of 209 ms. It trails a general model on reasoning-heavy sets such as GPQA Diamond and MMLU-Pro.

    The scores are self-reported, and the model needs a large GPU; Clef-Flash is the faster variant.

    Architecture

    The Qwen3.8-27B backbone, including its vision encoder, reads the encoded record: the state, any images or video frames, and each question with its options. A small transformer head reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly, producing one logit per allowed option. A softmax per question gives the probabilities. Questions are typed as noul (true or false), choice (named options with descriptions) or score (ordered options).

    Mixpeek SDK Integration

    # Ask Clef typed questions about each image or video frame, then store the answers
    # in metadata so Mixpeek search can filter or group by them.
    import requests
    
    for item in items:  # [{"url": "s3://...", "answers": {"legible": True, "department": "billing"}}]
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={
                "blobs": [{"property": "image", "type": "image", "url": item["url"]}],
                "metadata": item["answers"],
            },
        )

    Capabilities

    • Typed questions: true/false, pick one of named options, or score on an ordered scale
    • Reads the situation as text, JSON, images or video frames
    • Returns a probability for every allowed option of every question in one forward pass
    • Apache-2.0; a smaller Clef-Flash variant trades accuracy for latency

    Use Cases on Mixpeek

    Routing support messages, tickets and alerts by team and urgency
    Checking images and video frames against a policy (legible, on-brand, safe) with calibrated scores
    Labeling a media library with fixed categories before it is searched
    Choosing an agent's next tool or action from a fixed set

    Benchmarks

    DatasetMetricScoreSource
    BANKING77Macro-F194.2Model card: Cloudflare/clef (self-reported, Cloudflare's Decision Index 0.2.1 run)
    CLINC150 + out-of-scopeMacro-F197.4Model card (self-reported)
    RAGTruthHallucination F179.4Model card (self-reported)
    Decision IndexMedian latency209 msModel card (self-reported; Clef-Flash 38.8 ms)

    Performance

    Input SizeText or JSON state plus optional images and video frames; up to 16,384 tokens by default
    Embedding Dimn/a (outputs a probability per allowed option of each question)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Tested by Cloudflare with torch 2.11 and transformers 5.10.2 on a single H200. It needs a large GPU; Clef-Flash is the smaller, faster variant. We have not measured it.

    Frequently Asked Questions

    What does Clef do?

    You give it a situation (text, JSON, images or video) and a set of typed questions, and it returns a probability for every allowed answer to every question in one pass. It does not write free text, so there is nothing to parse.

    Can Clef read images and video?

    Yes. It keeps the vision encoder of its Qwen3.8-27B backbone, so a record can include images or video frames alongside the text state, and text-only and multimodal records can share a batch.

    How is Clef different from Clef-Flash?

    Clef is the larger model and leads on most of the card's classification and reasoning benchmarks, such as CLINC150 (97.4 against 66.8 macro-F1). Clef-Flash is faster, 38.8 ms median latency against 209 ms, and leads on some tasks such as the home appliance simulator.

    How do I use Clef answers in Mixpeek search?

    Store each answer in the object's metadata and filter or group a retriever by it, as in the example on this page. Mixpeek taxonomies can also assign labels inside the pipeline.

    Specification

    OrganizationCloudflare
    Retriever-
    Parameters27B
    LicenseApache-2.0
    Downloads/moN/A
    Likes451

    Research Paper

    Clef decision models (Cloudflare blog)

    arxiv.org

    Build a pipeline with clef

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free