DocumentKnowledge GraphConverter
Transform documents into structured knowledge graphs by extracting entities, relationships, and concepts. Produces nodes and edges suitable for graph databases, enabling complex queries, reasoning, and visualization over document content.
How It Works
Upload a document or provide a URL to the Mixpeek API.
Text is extracted and segmented into paragraphs and sections.
Named entity recognition identifies people, organizations, locations, concepts, and domain-specific terms.
Relationship extraction identifies connections between entities (e.g., 'works at', 'located in', 'causes').
A knowledge graph is returned as nodes and edges in JSON-LD, RDF, or a custom graph format.
Code Examples
import os, requests
API = "https://api.mixpeek.com"
H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
"X-Namespace": os.environ["NAMESPACE_ID"]}
# 1. a bucket, with a schema that declares the field you will send
bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
"bucket_name": "document-inputs",
"bucket_schema": {"properties": {"document": {"type": "pdf"}}},
}).json()
# 2. land the file as an object. the URL goes in data, on the blob
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
"key_prefix": "run-1",
"blobs": [{"property": "document", "type": "pdf",
"data": "https://example.com/contract.pdf"}],
})
# 3. a collection over that bucket, running the extractor
collection = requests.post(f"{API}/v1/collections", headers=H, json={
"collection_name": "document-to-knowledge-graph",
"source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
"feature_extractor": {"feature_extractor_name": "document_graph_extractor", "version": "v1"},
}).json()
# 4. run extraction over the bucket
requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
"collection_ids": [collection["collection_id"]],
"auto_submit": True,
})
# 5. read the graph the extractor produced
docs = requests.get(
f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
).json()
print(docs)Use Cases
Supported Input Formats
Quick Info
Run it over a library
Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.
Frequently Asked Questions
Related Converters
PDF to Text
Extract clean, structured text from PDF documents including scanned pages, multi-column layouts, headers/footers, and tables. Combines traditional parsing with OCR and layout analysis for maximum accuracy.
PDF to Structured Data
Extract structured key-value pairs, tables, and form fields from PDF documents. Uses layout analysis and LLM extraction to produce clean JSON output, even from complex forms and invoices.
HTML to Structured Data
Extract structured data from web pages using a combination of CSS/XPath selectors and LLM-based extraction. Captures product details, article metadata, contact information, and custom schemas from any website.
Text to Embeddings
How to turn text into embeddings for semantic search and RAG: which model to pick, how many dimensions you need, whether to chunk first, and where to store the vectors. Converts sentences, paragraphs and documents into dense vectors with the same model for queries and passages.
PDF to JSON
Convert PDF documents into clean, structured JSON output. Extracts text, tables, form fields, metadata, and document structure into a machine-readable JSON format suitable for API ingestion, database storage, and programmatic processing.
Ready to convert document to knowledge graph?
Start using the Mixpeek Document to Knowledge Graph in minutes. Sign up for a free API key and follow the documentation to get started.