NEWVectors or files. Pick a path.Start →
    data

    Audio
    Sentiment
    Converter

    Analyze the sentiment and emotional tone of audio recordings by combining speech transcription with acoustic feature analysis. Detects positive, negative, and neutral sentiment at utterance and segment levels, with additional emotion classification for anger, joy, frustration, and more.

    Max file size: 2 GB
    Estimated: 2-6 min per hour of audio
    5 input formats

    How It Works

    1

    Upload an audio file or provide a URL to the Mixpeek API.

    2

    The audio is transcribed with speaker diarization and utterance-level segmentation.

    3

    Lexical sentiment is analyzed from the transcript text using an NLP model.

    4

    Acoustic sentiment is analyzed from vocal features including pitch, energy, speaking rate, and tone.

    5

    Lexical and acoustic scores are fused into a combined sentiment and emotion profile per segment.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "audio-inputs",
        "bucket_schema": {"properties": {"audio": {"type": "audio"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "audio", "type": "audio",
                   "data": "https://example.com/call.mp3"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "audio-to-sentiment",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "multimodal_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Analyze customer satisfaction trends across call center recordings
    Monitor agent tone and empathy during support interactions
    Detect escalation points in recorded disputes and complaints
    Measure audience engagement and emotional response in focus group recordings

    Supported Input Formats

    MP3
    WAV
    FLAC
    OGG
    AAC

    Quick Info

    Categorydata
    Max File Size2 GB
    Est. Time2-6 min per hour of audio

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert audio to sentiment?

    Start using the Mixpeek Audio to Sentiment in minutes. Sign up for a free API key and follow the documentation to get started.