NEWVectors or files. Pick a path.Start →
    data

    Document
    Knowledge Graph
    Converter

    Transform documents into structured knowledge graphs by extracting entities, relationships, and concepts. Produces nodes and edges suitable for graph databases, enabling complex queries, reasoning, and visualization over document content.

    Max file size: 200 MB
    Estimated: 5-30 sec per page
    5 input formats

    How It Works

    1

    Upload a document or provide a URL to the Mixpeek API.

    2

    Text is extracted and segmented into paragraphs and sections.

    3

    Named entity recognition identifies people, organizations, locations, concepts, and domain-specific terms.

    4

    Relationship extraction identifies connections between entities (e.g., 'works at', 'located in', 'causes').

    5

    A knowledge graph is returned as nodes and edges in JSON-LD, RDF, or a custom graph format.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "document-inputs",
        "bucket_schema": {"properties": {"document": {"type": "pdf"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "document", "type": "pdf",
                   "data": "https://example.com/contract.pdf"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "document-to-knowledge-graph",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "document_graph_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the graph the extractor produced
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Build knowledge bases from legal contracts and regulatory documents
    Map relationships between entities in research paper collections
    Create interactive knowledge graphs for corporate intelligence platforms
    Power question-answering systems that reason over document relationships

    Supported Input Formats

    PDF
    DOCX
    TXT
    HTML
    Markdown

    Quick Info

    Categorydata
    Max File Size200 MB
    Est. Time5-30 sec per page

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert document to knowledge graph?

    Start using the Mixpeek Document to Knowledge Graph in minutes. Sign up for a free API key and follow the documentation to get started.