NEWVectors or files. Pick a path.Start →
    document

    Image
    Searchable Text
    Converter

    Convert images containing text into fully searchable, indexed content using advanced OCR combined with layout understanding. Preserves document structure including paragraphs, columns, headers, and reading order for downstream search and retrieval.

    Max file size: 50 MB
    Estimated: 1-5 sec per image
    5 input formats

    How It Works

    1

    Upload an image or provide a URL to the Mixpeek API.

    2

    The image is preprocessed with deskew, binarization, and contrast enhancement.

    3

    Layout analysis detects columns, paragraphs, headers, and reading order.

    4

    OCR extracts character-level text with confidence scores per word.

    5

    Structured, search-ready text is returned with preserved reading order and optional positional metadata.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "image-inputs",
        "bucket_schema": {"properties": {"image": {"type": "image"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "image", "type": "image",
                   "data": "https://example.com/photo.jpg"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "image-to-searchable-text",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "image_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Digitize scanned documents and make them full-text searchable
    Extract text from photographed whiteboards and handwritten notes
    Process scanned receipts and invoices for accounting systems
    Convert historical document archives into searchable digital collections

    Supported Input Formats

    JPEG
    PNG
    WebP
    TIFF
    BMP

    Quick Info

    Categorydocument
    Max File Size50 MB
    Est. Time1-5 sec per image

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert image to searchable text?

    Start using the Mixpeek Image to Searchable Text in minutes. Sign up for a free API key and follow the documentation to get started.