NEWVectors or files. Pick a path.Start →
    Models/m-a-p/MERT-v2-FullSong
    CC-BY-NC-4.0

    MERT-v2-FullSong

    by m-a-p

    Music embeddings from whole songs, for similarity search, tagging and analysis (non-commercial license)

    Identifiers
    Model ID
    m-a-p/MERT-v2-FullSong
    Feature URI

    Deploy MERT-v2-FullSong

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    MERT-v2-FullSong turns a complete song into a vector that captures how it sounds: genre, mood, instrumentation, key and rhythm. That supports music similarity search, recommendation and auto-tagging from the audio itself, without relying on titles or metadata. It is a 632M-parameter bidirectional encoder that reads 24 kHz mono audio at 25 frames per second and outputs 1024-dimension vectors per frame, which can be averaged into one vector per song.

    The license comes first: the weights are CC BY-NC 4.0, so they can be used for research and other non-commercial work only. M-A-P released it in September 2026 as part of the YuE2 family. It continues training from MERT-v2-30s on songs of 30 to 360 seconds.

    On the MARBLE music benchmark the card reports frozen-encoder results ahead of the strongest external baseline on 13 of 15 metrics, including 90.69 genre accuracy on GTZAN. The baseline figures are copied from another paper rather than rerun, and the card says evaluation protocols may differ.

    Architecture

    A 24-layer bidirectional Transformer encoder over audio, 632M parameters, taking 24 kHz mono input and emitting 1024-dimension hidden states at 25 Hz. It continues pretraining from MERT-v2-30s on full-length songs, keeping the same feature interface, and loads through Hugging Face Transformers with trust_remote_code. Every layer's hidden state is exposed; the card's layer guide recommends a different layer per task, for example layer 23 for tagging and layer 12 for instrument recognition, with the encoder frozen and a small probe trained on top.

    Mixpeek SDK Integration

    # CC BY-NC 4.0: check the license before any commercial use. No Mixpeek extractor
    # runs these weights. Mean-pool the last hidden state over the song's frames, as the
    # card's example does (1024 floats), then upsert into a namespace whose vector index
    # is declared at 1024 dimensions.
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
        headers={"Authorization": "Bearer API_KEY"},
        json={
            "collection_id": "col_your_collection",
            "documents": [{
                "document_id": "track-0042",
                "vectors": {"mert_v2": vector},  # the name of your 1024-d index
                "payload": {"source_key": "s3://music/track-0042.wav"},
            }],
        },
    )

    Capabilities

    • Music representations from complete songs of 30 to 360 seconds
    • Frame-level (25 Hz) and whole-recording embeddings, 1024 dimensions
    • Strong frozen-encoder results on tagging, genre, key, beat and emotion tasks
    • CC BY-NC 4.0: research and other non-commercial use only

    Use Cases on Mixpeek

    Research on music similarity and recommendation over a catalog
    Prototyping sound-alike search before choosing a commercially licensed model
    Auto-tagging experiments: genre, mood, instrument and key
    Comparing full-song representations with 30-second clip models on your own music

    Benchmarks

    DatasetMetricScoreSource
    GTZAN genreAccuracy90.69Model card: m-a-p/MERT-v2-FullSong, MARBLE frozen-encoder results (self-reported)
    MagnaTagATuneROC-AUC91.74Model card (self-reported)
    GiantSteps keyAccuracy67.05Model card (self-reported)
    MTG-Jamendo top-50 tagsROC-AUC84.13Model card (self-reported)

    Performance

    Input SizeMono audio at 24 kHz, complete songs of 30 to 360 seconds
    Embedding Dim1024 per frame at 25 Hz; mean-pool for one vector per song
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    We have not measured encoding latency. A 1024-dimension float32 song vector is 4 KiB; keeping frame-level vectors costs 25 of those per second of audio.

    Frequently Asked Questions

    Can I use MERT-v2-FullSong commercially?

    Not under its published license. The weights are released under CC BY-NC 4.0, which allows research and other non-commercial use and rules out commercial use unless the rights holders agree otherwise. For a commercial product, use a model with a commercial license or ask the authors.

    How do I get one vector per song from MERT-v2-FullSong?

    Run the song through the model and average the last hidden state over its frames, using the feature attention mask so padding is ignored. That gives one 1024-dimension vector per recording. The card's layer guide notes that some tasks, such as instrument or mood tagging, work better from middle layers than from the last one.

    What is the difference between MERT-v2-FullSong and MERT-v2-30s?

    Same size and feature interface. The 30s model is trained on 30-second clips; FullSong continues training from it on complete songs of 30 to 360 seconds. On the card's MARBLE results they are close, with the 30s model ahead on most tagging metrics and FullSong ahead on key and valence.

    Does Mixpeek run MERT-v2-FullSong?

    Not as a built-in extractor, and its non-commercial license limits where it can be used. For research, encode songs yourself and store one vector per song in a Mixpeek namespace, then search with a retriever that takes a query vector. For commercial sound search, Mixpeek's audio fingerprint extractor uses CLAP.

    Specification

    Organizationm-a-p
    Retriever-
    Parameters632M
    LicenseCC-BY-NC-4.0
    Downloads/moN/A
    Likes41

    Research Paper

    MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training (ICLR 2024)

    arxiv.org

    Build a pipeline with MERT-v2-FullSong

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free