NEWVectors or files. Pick a path.Start →
    document

    PDF
    Markdown
    Converter

    Convert a PDF into Markdown right on this page. pdf.js reads the text in your browser, and the reader turns larger type into headings, bullet lines into lists and wrapped lines into paragraphs. Nothing is uploaded.

    Max file size: 200 MB, read in your browser
    Estimated: Seconds for most documents, in your browser
    1 input formats

    Convert a PDF to Markdown here

    Runs in your browser. The file stays on your device and nothing is uploaded.

    Drop a PDF here, or choose one from your device.

    Loading the reader.

    How It Works

    1

    Choose a PDF or drop it on the reader at the top of this page. It stays in your browser.

    2

    pdf.js reads the text layer of every page in a worker thread.

    3

    The text size covering the most characters is taken as body text. A short line set at least 1.15 times larger becomes a heading: 1.8 times or more is #, 1.3 times ##, and smaller than that ###.

    4

    Lines starting with a bullet or a number become list items, and lines close together join into paragraphs. A larger gap above a line starts a new paragraph.

    5

    Copy the Markdown or download it as a .md file. For a document library, the API code below runs an extractor over every PDF in a bucket on Mixpeek's servers.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "pdf-inputs",
        "bucket_schema": {"properties": {"pdf": {"type": "pdf"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "pdf", "type": "pdf",
                   "data": "https://example.com/report.pdf"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "pdf-to-markdown",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "document_graph_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Migrate documentation from PDF to GitBook or Notion
    Convert research papers to blog-ready Markdown
    Import PDF content into static site generators (Hugo, Jekyll)
    Create editable drafts from finalized PDF reports

    Supported Input Formats

    PDF

    Quick Info

    Categorydocument
    Max File Size200 MB, read in your browser
    Est. TimeSeconds for most documents, in your browser

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. The reader on this page handles one file at a time.

    Frequently Asked Questions

    Ready to convert pdf to markdown?

    Start using the Mixpeek PDF to Markdown in minutes. Sign up for a free API key and follow the documentation to get started.