NEWVectors or files. Pick a path.Start →
    media

    Audio
    Text
    Converter

    Transcribe audio files into text with high accuracy. Supports speaker diarization, punctuation restoration, timestamps, and over 50 languages. Handles podcasts, calls, meetings, and broadcast audio.

    Max file size: 2 GB
    Estimated: 1-5 min per hour of audio
    7 input formats

    How It Works

    1

    Upload an audio file or provide a URL.

    2

    The audio is preprocessed (noise reduction, normalization).

    3

    Speech is transcribed using a large speech model.

    4

    Speaker diarization assigns text segments to individual speakers.

    5

    Timestamps, punctuation, and formatting are applied.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "audio-inputs",
        "bucket_schema": {"properties": {"audio": {"type": "audio"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "audio", "type": "audio",
                   "data": "https://example.com/call.mp3"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "audio-to-text",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "multimodal_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Transcribe podcast episodes for show notes and SEO
    Convert call center recordings to searchable text
    Generate meeting minutes from recorded calls
    Create text datasets from audio archives

    Supported Input Formats

    MP3
    WAV
    FLAC
    OGG
    AAC
    M4A
    WMA

    Quick Info

    Categorymedia
    Max File Size2 GB
    Est. Time1-5 min per hour of audio

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert audio to text?

    Start using the Mixpeek Audio to Text in minutes. Sign up for a free API key and follow the documentation to get started.