NEWVectors or files. Pick a path.Start →
    document

    PDF
    Table Data
    Converter

    Extract tables from PDF documents and convert them into structured formats like JSON arrays, CSV, or Excel. Handles complex table layouts with merged cells, nested headers, multi-page tables, and borderless tables using AI-powered layout detection.

    Max file size: 200 MB
    Estimated: 2-10 sec per page
    1 input formats

    How It Works

    1

    Upload a PDF file or provide a URL to the Mixpeek API.

    2

    AI-powered layout analysis detects all table regions on each page.

    3

    Cell boundaries are identified using a combination of rule detection and machine learning.

    4

    Merged cells, nested headers, and multi-page continuation tables are resolved into clean row/column structures.

    5

    Tables are returned as structured arrays with headers, rows, and optional type inference per column.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "pdf-inputs",
        "bucket_schema": {"properties": {"pdf": {"type": "pdf"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "pdf", "type": "pdf",
                   "data": "https://example.com/report.pdf"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "pdf-to-table-data",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "document_graph_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Extract financial tables from SEC filings and annual reports
    Pull pricing tables from vendor quotes and proposals
    Digitize scientific data tables from research papers
    Convert regulatory compliance tables into spreadsheet-ready formats

    Supported Input Formats

    PDF

    Quick Info

    Categorydocument
    Max File Size200 MB
    Est. Time2-10 sec per page

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert pdf to table data?

    Start using the Mixpeek PDF to Table Data in minutes. Sign up for a free API key and follow the documentation to get started.