HTMLJSONConverter
Extract structured data from web pages using a combination of CSS/XPath selectors and LLM-based extraction. Captures product details, article metadata, contact information, and custom schemas from any website.
How It Works
Provide a URL or upload an HTML file.
Existing structured data (JSON-LD, microdata, RDFa) is extracted first.
An LLM analyzes the page to extract additional structured fields.
Results are merged and validated against your target schema.
Clean JSON output is returned with confidence scores.
Code Examples
import os, requests
API = "https://api.mixpeek.com"
H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
"X-Namespace": os.environ["NAMESPACE_ID"]}
# 1. a bucket, with a schema that declares the field you will send
bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
"bucket_name": "html-inputs",
"bucket_schema": {"properties": {"html": {"type": "text"}}},
}).json()
# 2. land the file as an object. the URL goes in data, on the blob
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
"key_prefix": "run-1",
"blobs": [{"property": "html", "type": "text",
"data": "https://example.com/page.html"}],
})
# 3. a collection over that bucket, running the extractor
collection = requests.post(f"{API}/v1/collections", headers=H, json={
"collection_name": "html-to-structured-data",
"source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
"feature_extractor": {"feature_extractor_name": "web_scraper", "version": "v1"},
}).json()
# 4. run extraction over the bucket
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
"collection_ids": [collection["collection_id"]],
"auto_submit": True,
})
# 5. read the output
docs = requests.get(
f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
).json()
print(docs)Use Cases
Supported Input Formats
Quick Info
Run it over a library
Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.
Frequently Asked Questions
Related Converters
PDF to Structured Data
Extract structured key-value pairs, tables, and form fields from PDF documents. Uses layout analysis and LLM extraction to produce clean JSON output, even from complex forms and invoices.
JSON to Embeddings
Convert JSON objects and arrays into semantic vector embeddings. Supports nested structures, field selection, and configurable serialization strategies for optimal embedding quality.
HTML to Text
Extract clean, readable text from HTML pages by stripping tags, scripts, and styles while preserving semantic structure. Handles navigation removal, boilerplate detection, and main content extraction.
Ready to convert html to json?
Start using the Mixpeek HTML to Structured Data in minutes. Sign up for a free API key and follow the documentation to get started.