NEWVectors or files. Pick a path.Start →
    Document AI
    9 min read
    Updated 2026-09-28

    How Do I Make Hundreds of Scanned PDFs Searchable?

    A scanned PDF is a stack of page pictures, so searching it for a word you can plainly see returns nothing. This explains how to tell, the ways to add a text layer from one file to thousands, how to get OCR accurate on real scans, and how to search the result by meaning as well as by exact word.

    OCR
    PDF
    Scanned Documents
    Document Search
    Document AI

    How do I make scanned PDFs searchable?



    Run OCR (optical character recognition) on them, which reads the text in each page image and saves it as an invisible text layer behind the picture. After that, search inside the PDF works, desktop search can find the files, and any search index can read the text. For a handful of files, the OCR command in a PDF editor is enough. For hundreds or thousands, use a batch OCR tool or an OCR service, then load the text into a search index so you can find a passage across every file at once.

    Why can't I search my scanned PDFs?



    A PDF made by a scanner or a phone camera stores each page as an image. The file holds pixels shaped like letters, and search has nothing to match against them. PDFs exported from Word or a website carry real text, which is why some of your PDFs search fine and others do not.

    The quick test: open the file and try to select a word with the cursor, or search for a word you can see on the page. If you cannot select it, or search finds nothing, the page is an image.

    What are my options, from one file to thousands?



    OptionBest forWhat it producesWhere it falls short
    The OCR command in a PDF editor (for example Adobe Acrobat)A few files, done by handA searchable PDF with a text layerOne file at a time; the batch features are usually in paid tiers
    Opening the PDF in Google Docs from Google DriveOccasional single filesAn editable document with the recognised textFormatting and page layout are often lost; not built for batches
    An open-source batch tool such as OCRmyPDF (built on Tesseract)Hundreds to thousands of files on your own machineSearchable PDFs that keep the original page imagesNeeds a command line; tables, handwriting and poor scans need care
    A cloud document OCR service (Google Document AI, Amazon Textract, Azure AI Document Intelligence)Large volumes, tables and formsText plus layout: blocks, tables, key-value fieldsPriced per page; you build the pipeline around it
    A search platform that runs OCR at ingestSearching the whole collection, not just making files searchableAn index you can query across every pageYou adopt a platform instead of producing files
    A text layer makes each file searchable on its own. Searching every file at once, ranking the results, and finding a passage worded differently from your query needs a search index built on top of the extracted text.

    How do I get OCR right on real scans?



    OCR accuracy depends mostly on the page image. The fixes that matter most:

  1. Scan at 300 DPI or more. Tesseract's own documentation recommends at least 300 DPI; below that,
  2. small print loses the shapes OCR relies on.
  3. Straighten and clean the page. Deskewing a tilted scan and removing dark borders and speckle
  4. noticeably improves recognition. Most batch tools have options for both.
  5. Tell it the language. OCR engines use a language model to decide between look-alike characters,
  6. so the wrong language setting produces plausible but wrong words.
  7. Check the reading order on multi-column pages and tables. Plain OCR reads left to right across
  8. the page, which can interleave two columns or flatten a table into a run of numbers. Layout-aware OCR detects blocks, columns and tables first, then reads each in order.
  9. Treat handwriting separately. Printed-text OCR handles handwriting poorly. Cloud services and
  10. vision-language models do better, and it is worth checking a sample before trusting the output.

    To measure it, pick twenty pages that represent your collection, correct their text by hand once, and compare the OCR output against that. It tells you whether the problem is the scans or the tool.

    How do I search across all of them, and by meaning?



    Once the text exists, put it in a search index with a reference back to the file and the page. Two kinds of search are worth having together:

  11. Keyword search finds exact words, names and numbers, such as an invoice number or a part code.
  12. Semantic search finds passages by meaning, so "termination notice period" finds a clause that
  13. says "either party may end this agreement with 90 days' notice". It works by turning each passage and each query into an embedding, a vector where similar meanings sit close together.

    Running both and merging the results, called hybrid search, covers the exact lookups and the questions asked in the reader's own words. Some answers are not in the text at all, such as a value read off a chart; for those, see why document search misses the answer in a chart or table.

    How do I do this with Mixpeek?



    Put the PDFs in object storage you already use, such as S3 or GCS, and connect the bucket to Mixpeek. The document layout extractor runs OCR with layout detection on every page, splitting it into paragraphs, tables, forms and headers, and sends low-confidence blocks from degraded scans to a vision-language model for correction. Each block is indexed with its page and position, so a retriever returns the passage and the page it came from, with keyword and semantic search in one query. At the published rate of $1.50 per thousand document pages, indexing 10,000 scanned pages costs about $15 in processing, before storage and queries.

    Related: how OCR works, from detection to reading order, the best OCR APIs, the best PDF extraction tools and why search can't find exact part numbers, codes or names.

    Frequently Asked Questions



    How can I tell if a PDF is scanned or has real text?



    Try to select a word with the cursor, or search for a word you can see on the page. If you cannot select it and search finds nothing, the page is an image and the PDF needs OCR before it can be searched.

    What is the best free way to OCR a lot of PDFs?



    An open-source batch tool such as OCRmyPDF, which uses the Tesseract OCR engine, adds a text layer to each PDF while keeping the original page images. It runs on your own machine, handles whole folders, and has options to straighten pages and set the language.

    Does OCR work on handwriting?



    Poorly, with classic printed-text OCR. Cloud document services and vision-language models read handwriting better, but accuracy varies with the writer, so check a sample of your own pages before relying on it.

    Why does my OCR mix up columns or tables?



    Plain OCR reads straight across the page, so two columns interleave and a table turns into a line of numbers. Use layout-aware OCR, which detects columns, tables and blocks first and reads each one in its own order.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs