speaker-diarization-3.1
by pyannote
Who spoke when, end-to-end neural speaker diarization
pyannote/speaker-diarization-3.1mixpeek://transcription@v1/pyannote_diarization_v3Deploy speaker-diarization-3.1
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Pyannote's speaker diarization pipeline segments audio into speaker-homogeneous regions, determining "who spoke when" without requiring prior knowledge of the number or identity of speakers.
On Mixpeek, speaker diarization enriches transcription data with speaker labels, enabling queries like "find all segments where Speaker A talks about budgets."
Architecture
End-to-end pipeline: (1) segmentation model based on PyanNet (SincNet + LSTM + feedforward), (2) embedding extraction using ECAPA-TDNN, (3) agglomerative clustering for speaker assignment. Supports overlapping speech detection.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so speaker-diarization-3.1 runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The model produces text, so it lands in payload. Give the
// collection a text vector index and embed that text to make it
// searchable rather than only filterable.
payload: { extracted_text: modelOutput, source_key: "archive/2026/asset-00412" },
vectors: { "text-embedding": embeddingOfModelOutput },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// universal_extractor@v1 runs google/gemini-embedding-2
// (3072-d) over a bucket, with no inference of your own.Capabilities
- Automatic speaker count estimation
- Overlapping speech detection
- Speaker embedding extraction
- Fine-tunable on custom speaker data
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| AMI (headset) | DER | 18.2% | Plaquet & Bredin, 2023: Table 2 |
| DIHARD III | DER | 20.5% | Plaquet & Bredin, 2023: Table 2 |
| VoxConverse | DER | 11.2% | Plaquet & Bredin, 2023: Table 2 |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Powerset multi-class cross entropy loss for neural speaker diarization
arxiv.orgBuild a pipeline with speaker-diarization-3.1
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free