NEWVectors or files. Pick a path.Start →
    data

    HTML
    JSON
    Converter

    Extract structured data from web pages using a combination of CSS/XPath selectors and LLM-based extraction. Captures product details, article metadata, contact information, and custom schemas from any website.

    Max file size: 50 MB
    Estimated: 2-8 sec per page
    2 input formats

    How It Works

    1

    Provide a URL or upload an HTML file.

    2

    Existing structured data (JSON-LD, microdata, RDFa) is extracted first.

    3

    An LLM analyzes the page to extract additional structured fields.

    4

    Results are merged and validated against your target schema.

    5

    Clean JSON output is returned with confidence scores.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "html-inputs",
        "bucket_schema": {"properties": {"html": {"type": "text"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "html", "type": "text",
                   "data": "https://example.com/page.html"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "html-to-structured-data",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "web_scraper", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Scrape product details from e-commerce pages
    Extract article metadata from news sites
    Capture business listings from directory pages
    Build structured datasets from web sources

    Supported Input Formats

    HTML
    XHTML

    Quick Info

    Categorydata
    Max File Size50 MB
    Est. Time2-8 sec per page
    Extractorweb-scraper

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert html to json?

    Start using the Mixpeek HTML to Structured Data in minutes. Sign up for a free API key and follow the documentation to get started.