MERT-v2-FullSong
by m-a-p
Music embeddings from whole songs, for similarity search, tagging and analysis (non-commercial license)
m-a-p/MERT-v2-FullSongDeploy MERT-v2-FullSong
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
MERT-v2-FullSong turns a complete song into a vector that captures how it sounds: genre, mood, instrumentation, key and rhythm. That supports music similarity search, recommendation and auto-tagging from the audio itself, without relying on titles or metadata. It is a 632M-parameter bidirectional encoder that reads 24 kHz mono audio at 25 frames per second and outputs 1024-dimension vectors per frame, which can be averaged into one vector per song.
The license comes first: the weights are CC BY-NC 4.0, so they can be used for research and other non-commercial work only. M-A-P released it in September 2026 as part of the YuE2 family. It continues training from MERT-v2-30s on songs of 30 to 360 seconds.
On the MARBLE music benchmark the card reports frozen-encoder results ahead of the strongest external baseline on 13 of 15 metrics, including 90.69 genre accuracy on GTZAN. The baseline figures are copied from another paper rather than rerun, and the card says evaluation protocols may differ.
Architecture
A 24-layer bidirectional Transformer encoder over audio, 632M parameters, taking 24 kHz mono input and emitting 1024-dimension hidden states at 25 Hz. It continues pretraining from MERT-v2-30s on full-length songs, keeping the same feature interface, and loads through Hugging Face Transformers with trust_remote_code. Every layer's hidden state is exposed; the card's layer guide recommends a different layer per task, for example layer 23 for tagging and layer 12 for instrument recognition, with the encoder frozen and a small probe trained on top.
Mixpeek SDK Integration
# CC BY-NC 4.0: check the license before any commercial use. No Mixpeek extractor
# runs these weights. Mean-pool the last hidden state over the song's frames, as the
# card's example does (1024 floats), then upsert into a namespace whose vector index
# is declared at 1024 dimensions.
import requests
requests.post(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
headers={"Authorization": "Bearer API_KEY"},
json={
"collection_id": "col_your_collection",
"documents": [{
"document_id": "track-0042",
"vectors": {"mert_v2": vector}, # the name of your 1024-d index
"payload": {"source_key": "s3://music/track-0042.wav"},
}],
},
)Capabilities
- Music representations from complete songs of 30 to 360 seconds
- Frame-level (25 Hz) and whole-recording embeddings, 1024 dimensions
- Strong frozen-encoder results on tagging, genre, key, beat and emotion tasks
- CC BY-NC 4.0: research and other non-commercial use only
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| GTZAN genre | Accuracy | 90.69 | Model card: m-a-p/MERT-v2-FullSong, MARBLE frozen-encoder results (self-reported) |
| MagnaTagATune | ROC-AUC | 91.74 | Model card (self-reported) |
| GiantSteps key | Accuracy | 67.05 | Model card (self-reported) |
| MTG-Jamendo top-50 tags | ROC-AUC | 84.13 | Model card (self-reported) |
Performance
We have not measured encoding latency. A 1024-dimension float32 song vector is 4 KiB; keeping frame-level vectors costs 25 of those per second of audio.
Common Pipeline Companions
Frequently Asked Questions
Can I use MERT-v2-FullSong commercially?
Not under its published license. The weights are released under CC BY-NC 4.0, which allows research and other non-commercial use and rules out commercial use unless the rights holders agree otherwise. For a commercial product, use a model with a commercial license or ask the authors.
How do I get one vector per song from MERT-v2-FullSong?
Run the song through the model and average the last hidden state over its frames, using the feature attention mask so padding is ignored. That gives one 1024-dimension vector per recording. The card's layer guide notes that some tasks, such as instrument or mood tagging, work better from middle layers than from the last one.
What is the difference between MERT-v2-FullSong and MERT-v2-30s?
Same size and feature interface. The 30s model is trained on 30-second clips; FullSong continues training from it on complete songs of 30 to 360 seconds. On the card's MARBLE results they are close, with the 30s model ahead on most tagging metrics and FullSong ahead on key and valence.
Does Mixpeek run MERT-v2-FullSong?
Not as a built-in extractor, and its non-commercial license limits where it can be used. For research, encode songs yourself and store one vector per song in a Mixpeek namespace, then search with a retriever that takes a query vector. For commercial sound search, Mixpeek's audio fingerprint extractor uses CLAP.
Specification
Research Paper
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training (ICLR 2024)
arxiv.orgBuild a pipeline with MERT-v2-FullSong
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free