NEWVectors or files. Pick a path.Start →
    Models/Embeddings/laion/clap-htsat-fused
    HFAudio Embeddingsapache-2.0

    clap-htsat-fused

    by laion

    Contrastive Language-Audio Pretraining for audio-text retrieval

    8.3Mdl/month
    126likes
    154Mparams
    Identifiers
    Model ID
    laion/clap-htsat-fused
    Feature URI
    mixpeek://audio_extractor@v1/laion_clap_fused_v1

    Deploy clap-htsat-fused

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    CLAP learns aligned audio and text representations through contrastive learning, similar to how CLIP works for images and text. The HTSAT-fused variant uses the HTS-AT audio transformer fused with RoBERTa text embeddings.

    On Mixpeek, CLAP enables semantic audio search, find audio segments matching natural language descriptions like "crowd cheering" or "rain on a roof."

    Architecture

    HTS-AT (Hierarchical Token-Semantic Audio Transformer) as audio encoder, RoBERTa as text encoder. Trained on AudioSet, Clotho, and other audio-text pair datasets with contrastive loss. Outputs 512-dim joint embedding space.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so clap-htsat-fused runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The vector name has to match a vector index on the collection.
              vectors: { "audio-embedding": yourVector },
              payload: { source_key: "archive/2026/asset-00412" },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // audio_fingerprint_extractor@v1 runs laion/clap-htsat-tiny
    // (512-d) over a bucket, with no inference of your own.

    Capabilities

    • Audio-text cross-modal retrieval
    • 512-dimensional audio embeddings
    • Zero-shot audio classification
    • Environmental sound recognition

    Use Cases on Mixpeek

    Sound effect search, find audio by description
    Music discovery, semantic similarity across audio tracks
    Environmental monitoring, classify ambient sounds

    Benchmarks

    DatasetMetricScoreSource
    ESC-50Accuracy (zero-shot)93.7%Wu et al., 2023: Table 2
    AudioCaps (text→audio)Recall@136.7%Wu et al., 2023: Table 3

    Performance

    Input Sizevariable audio (10s chunks typical)
    Embedding Dim512
    GPU Latency~6ms / chunk (A100)
    GPU Throughput~165 chunks/sec (A100)
    GPU Memory~0.5 GB

    Specification

    FrameworkHF
    Organizationlaion
    FeatureAudio Embeddings
    Output512-dim vector
    Modalitiesvideo, audio
    RetrieverAudio Similarity
    Parameters154M
    Licenseapache-2.0
    Downloads/mo8.3M
    Likes126

    Research Paper

    Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation

    arxiv.org

    Build a pipeline with clap-htsat-fused

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free