NEWVectors or files. Pick a path.Start →
    Models/convaiinnovations/laya
    Apache-2.0

    laya

    by convaiinnovations

    A 421M decision model: typed answers with calibrated probabilities in one forward pass, about 33 ms

    Identifiers
    Model ID
    convaiinnovations/laya
    Feature URI

    Deploy laya

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Laya answers typed questions about a piece of text with calibrated probabilities, in one forward pass of about 33 ms. Give it a state (a message, an email, a ticket or JSON) and questions such as which department, how urgent, or will this customer leave, and it returns a choice, a score or a yes/no probability for each. It never generates text. Convai Innovations released it in September 2026 under Apache-2.0.

    The repository holds three checkpoints: an English one on ModernBERT-large (421M), a multilingual one on mmBERT-base (322M) for 100+ languages, and one fine-tuned on four typed-decision workflows. On the card's typed-decisions benchmark the base English checkpoint scores 0.362 accuracy and the fine-tuned one 0.766, so fine-tuning matters.

    In Cloudflare's independent Decision Index run, Laya scored well below larger decision models, for example 14.3 macro-F1 on BANKING77.

    Architecture

    A bidirectional encoder (ModernBERT-large for English, mmBERT-base for the multilingual checkpoint) reads the state and the questions together, and a head scores each allowed answer without generating tokens. It is trained with reinforcement learning against strictly proper scoring rules (RLCD), which rewards honest probabilities. A Router detects the script and language and sends non-English text to the multilingual checkpoint.

    Mixpeek SDK Integration

    # Ask typed questions about each item, then store the answers in metadata so
    # Mixpeek search can filter or group by them.
    import requests
    
    for item in items:  # [{"text": "...", "answers": {"department": "billing", "urgent": True}}]
        requests.post(
            "https://api.mixpeek.com/v1/buckets/bkt_your_bucket/objects",
            headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
            json={
                "blobs": [{"property": "body", "type": "text", "data": item["text"]}],
                "metadata": item["answers"],
            },
        )

    Capabilities

    • Typed questions: choice, score on an ordered scale, or yes/no probability
    • English checkpoint plus a multilingual one for 100+ languages, picked automatically by its Router
    • No text generation; answers come back as structured probabilities
    • Apache-2.0; HTTP server, MCP server, LangChain and ONNX extras

    Use Cases on Mixpeek

    Email and ticket triage by department, urgency and churn risk
    Guardrail and moderation checks on user text at low latency
    Multilingual routing where one model must cover many languages
    Domain decisions after fine-tuning on your own labeled examples

    Benchmarks

    DatasetMetricScoreSource
    Typed-decisions benchmark (2,000 decisions)Accuracy0.362Model card: convaiinnovations/laya (self-reported, base English checkpoint)
    Typed-decisions benchmark (2,000 decisions)Accuracy, fine-tuned0.766Model card (self-reported, laya-typed-decisions checkpoint)
    One question per call, Tesla T4Latency39.5 msModel card (self-reported; multilingual checkpoint 32.8 ms)
    BANKING77Macro-F114.3Third-party: Cloudflare's Decision Index 0.2.1 run, on the Cloudflare/clef model card

    Performance

    Input SizeText or JSON state; 512 tokens for the English checkpoint, up to 8,192 for laya-multilingual with max_len=8192
    Embedding Dimn/a (outputs typed answers with probabilities)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The card reports 39.5 ms for one question and 158.6 ms for ten on a Tesla T4; the multilingual checkpoint is about twice as fast with many questions. We have not measured it.

    Frequently Asked Questions

    What does Laya do?

    You give it a state (text, an email, a ticket or JSON) and typed questions, and it returns an answer to each with a probability, in a single forward pass of about 33 ms. It never generates text.

    How accurate is Laya without fine-tuning?

    On the card's typed-decisions benchmark the base English checkpoint scores 0.362 accuracy and a fine-tuned checkpoint 0.766. In Cloudflare's independent Decision Index run it scored 14.3 macro-F1 on BANKING77. Plan to fine-tune on your own decisions; the card links a free Kaggle notebook for it.

    Which languages does Laya support?

    The English checkpoint covers English; laya-multilingual covers 100+ languages and the Router sends non-English text to it automatically.

    How do I use Laya answers in Mixpeek search?

    Store each answer in the object's metadata and filter or group a retriever by it, as in the example on this page.

    Specification

    Organizationconvaiinnovations
    Retriever-
    Parameters421M
    LicenseApache-2.0
    Downloads/moN/A
    Likes5,026

    Research Paper

    Laya documentation

    arxiv.org

    Build a pipeline with laya

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free