HTMLJSONConverter
Pull the structured data out of an HTML page right on this page: title, meta description, canonical URL, Open Graph and Twitter tags, headings, links, images, JSON-LD and microdata, as JSON. Upload a saved .html file or paste the source. It is read in your browser and nothing is uploaded.
Pull the structured data from an HTML page here
Runs in your browser. The file stays on your device and nothing is uploaded.
Drop a saved .html file here, choose one from your device, or paste the page's source below.
Loading the reader.
How It Works
Choose a saved .html file, drop it on the reader at the top of this page, or paste a page's source. It stays in your browser.
Your browser's own parser reads the HTML without running its scripts or loading its images, stylesheets or frames.
The reader collects the title, language, meta tags, canonical URL, Open Graph and Twitter card tags, and every heading, link and image, resolving relative links against the page's base or canonical URL.
It parses each JSON-LD block, flagging any that is not valid JSON, and reads microdata items with their nested properties. RDFa is not read.
Download everything as JSON. For structured extraction across many pages, the API code below runs an extractor on Mixpeek's servers.
Code Examples
import os, requests
API = "https://api.mixpeek.com"
H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
"X-Namespace": os.environ["NAMESPACE_ID"]}
# 1. a bucket, with a schema that declares the field you will send
bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
"bucket_name": "html-inputs",
"bucket_schema": {"properties": {"html": {"type": "text"}}},
}).json()
# 2. land the file as an object. the URL goes in data, on the blob
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
"key_prefix": "run-1",
"blobs": [{"property": "html", "type": "text",
"data": "https://example.com/page.html"}],
})
# 3. a collection over that bucket, running the extractor
collection = requests.post(f"{API}/v1/collections", headers=H, json={
"collection_name": "html-to-structured-data",
"source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
"feature_extractor": {"feature_extractor_name": "web_scraper", "version": "v1"},
}).json()
# 4. run extraction over the bucket
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
"collection_ids": [collection["collection_id"]],
"auto_submit": True,
})
# 5. read the output
docs = requests.get(
f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
).json()
print(docs)Use Cases
Supported Input Formats
Quick Info
Run it over a library
Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. The reader on this page handles one file at a time.
Frequently Asked Questions
Related Converters
PDF to Markdown
Convert a PDF into Markdown right on this page. pdf.js reads the text in your browser, and the reader turns larger type into headings, bullet lines into lists and wrapped lines into paragraphs. Nothing is uploaded.
HTML to Text
Turn an HTML page into clean text right on this page. Upload a saved .html file or paste the source, and the reader strips scripts, styles and hidden elements in your browser while keeping headings, paragraphs, lists and tables as plain text. Nothing is uploaded.
Ready to convert html to json?
Start using the Mixpeek HTML to Structured Data in minutes. Sign up for a free API key and follow the documentation to get started.