PDFJSONConverter
Extract structured key-value pairs, tables, and form fields from PDF documents. Uses layout analysis and LLM extraction to produce clean JSON output, even from complex forms and invoices.
How It Works
Upload a PDF or provide a URL.
Layout analysis identifies form fields, tables, and key-value regions.
An LLM extracts values and maps them to a structured schema.
Tables are converted to row/column JSON arrays.
The complete structured output is returned as JSON.
Code Examples
import os, requests
API = "https://api.mixpeek.com"
H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
"X-Namespace": os.environ["NAMESPACE_ID"]}
# 1. a bucket, with a schema that declares the field you will send
bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
"bucket_name": "pdf-inputs",
"bucket_schema": {"properties": {"pdf": {"type": "pdf"}}},
}).json()
# 2. land the file as an object. the URL goes in data, on the blob
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
"key_prefix": "run-1",
"blobs": [{"property": "pdf", "type": "pdf",
"data": "https://example.com/report.pdf"}],
})
# 3. a collection over that bucket, running the extractor
collection = requests.post(f"{API}/v1/collections", headers=H, json={
"collection_name": "pdf-to-structured-data",
"source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
"feature_extractor": {"feature_extractor_name": "document_graph_extractor", "version": "v1"},
}).json()
# 4. run extraction over the bucket
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
"collection_ids": [collection["collection_id"]],
"auto_submit": True,
})
# 5. read the output
docs = requests.get(
f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
).json()
print(docs)Use Cases
Supported Input Formats
Quick Info
Run it over a library
Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.
Frequently Asked Questions
Related Converters
PDF to Text
Extract clean, structured text from PDF documents including scanned pages, multi-column layouts, headers/footers, and tables. Combines traditional parsing with OCR and layout analysis for maximum accuracy.
PDF to Embeddings
Convert PDF documents into semantic vector embeddings for search, retrieval, and RAG applications. Pages are chunked intelligently by sections and paragraphs, then embedded using text or multimodal models.
PDF to Markdown
Convert PDF documents to clean Markdown format, preserving headings, lists, tables, links, and emphasis. Ideal for migrating content into wikis, CMS platforms, and documentation systems.
HTML to Structured Data
Extract structured data from web pages using a combination of CSS/XPath selectors and LLM-based extraction. Captures product details, article metadata, contact information, and custom schemas from any website.
Ready to convert pdf to json?
Start using the Mixpeek PDF to Structured Data in minutes. Sign up for a free API key and follow the documentation to get started.