NEWVectors or files. Pick a path.Start →
    data

    HTML
    JSON
    Converter

    Pull the structured data out of an HTML page right on this page: title, meta description, canonical URL, Open Graph and Twitter tags, headings, links, images, JSON-LD and microdata, as JSON. Upload a saved .html file or paste the source. It is read in your browser and nothing is uploaded.

    Max file size: 50 MB, read in your browser
    Estimated: Under a second, in your browser
    2 input formats

    Pull the structured data from an HTML page here

    Runs in your browser. The file stays on your device and nothing is uploaded.

    Drop a saved .html file here, choose one from your device, or paste the page's source below.

    Loading the reader.

    How It Works

    1

    Choose a saved .html file, drop it on the reader at the top of this page, or paste a page's source. It stays in your browser.

    2

    Your browser's own parser reads the HTML without running its scripts or loading its images, stylesheets or frames.

    3

    The reader collects the title, language, meta tags, canonical URL, Open Graph and Twitter card tags, and every heading, link and image, resolving relative links against the page's base or canonical URL.

    4

    It parses each JSON-LD block, flagging any that is not valid JSON, and reads microdata items with their nested properties. RDFa is not read.

    5

    Download everything as JSON. For structured extraction across many pages, the API code below runs an extractor on Mixpeek's servers.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "html-inputs",
        "bucket_schema": {"properties": {"html": {"type": "text"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "html", "type": "text",
                   "data": "https://example.com/page.html"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "html-to-structured-data",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "web_scraper", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Check that a page's JSON-LD and Open Graph tags are valid before it goes live
    Scrape product details from e-commerce pages
    Extract article metadata from news sites
    Build structured datasets from web sources

    Supported Input Formats

    HTML
    XHTML

    Quick Info

    Categorydata
    Max File Size50 MB, read in your browser
    Est. TimeUnder a second, in your browser
    Extractorweb-scraper

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. The reader on this page handles one file at a time.

    Frequently Asked Questions

    Ready to convert html to json?

    Start using the Mixpeek HTML to Structured Data in minutes. Sign up for a free API key and follow the documentation to get started.