NEWVectors or files. Pick a path.Start →
    Models/Embeddings/facebook/pe-av-large
    HFAudio Embeddingsapache-2.0

    pe-av-large

    by facebook

    Joint audio-video-text embeddings from Meta's Perception Encoder family

    71Kdl/month
    67likes
    2.2Bparams
    Identifiers
    Model ID
    facebook/pe-av-large
    Feature URI
    mixpeek://audio_extractor@v1/facebook_pe_av_large_v1

    Deploy pe-av-large

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    PE-AV Large embeds audio, video, synchronized audio-video, and text into one shared retrieval space. It is useful when the same event is expressed through motion, sound, or language, such as a siren, a crowd reaction, a machine failure, or a tennis serve.

    On Mixpeek, PE-AV Large gives agents a single evidence channel for audiovisual retrieval. Instead of searching transcripts, frames, and audio fingerprints separately, an agent can retrieve clips where the sound and visual motion jointly match the query, then pass the top results to a reasoning model.

    Architecture

    Perception Encoder audio-video model with roughly 2.2B parameters. The model aligns raw audio, video frames, audio-video pairs, and text through contrastive training so cross-modal retrieval works across all supported input combinations.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so pe-av-large runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The vector name has to match a vector index on the collection.
              vectors: { "audio-embedding": yourVector },
              payload: { source_key: "archive/2026/asset-00412" },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // audio_fingerprint_extractor@v1 runs laion/clap-htsat-tiny
    // (512-d) over a bucket, with no inference of your own.

    Capabilities

    • Text-to-video, text-to-audio, and text-to-audio-video retrieval
    • Joint embeddings for synchronized sound and motion
    • Useful for clips where audio carries the key signal
    • Apache 2.0 license

    Use Cases on Mixpeek

    Find video moments by sound events, visual motion, or both
    Retrieve security, sports, or broadcast clips where audio changes the meaning
    Build agent memory over camera footage with synchronized audio
    Use one embedding family before transcript, object, or VLM reranking

    Performance

    Input SizeAudio, video, audio-video, or text input
    Embedding DimModel dependent
    GPU LatencyInput dependent
    GPU ThroughputBatch by clip for best throughput
    GPU Memory~5 GB plus serving overhead

    Specification

    FrameworkHF
    Organizationfacebook
    FeatureAudio Embeddings
    Output512-dim vector
    Modalitiesvideo, audio
    RetrieverAudio Similarity
    Parameters2.2B
    Licenseapache-2.0
    Downloads/mo71K
    Likes67

    Research Paper

    PE Audio Video

    arxiv.org

    Build a pipeline with pe-av-large

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free