How do I make scanned PDFs searchable?
Run OCR (optical character recognition) on them, which reads the text in each page image and saves it as an invisible text layer behind the picture. After that, search inside the PDF works, desktop search can find the files, and any search index can read the text. For a handful of files, the OCR command in a PDF editor is enough. For hundreds or thousands, use a batch OCR tool or an OCR service, then load the text into a search index so you can find a passage across every file at once.
Why can't I search my scanned PDFs?
A PDF made by a scanner or a phone camera stores each page as an image. The file holds pixels shaped like letters, and search has nothing to match against them. PDFs exported from Word or a website carry real text, which is why some of your PDFs search fine and others do not.
The quick test: open the file and try to select a word with the cursor, or search for a word you can see on the page. If you cannot select it, or search finds nothing, the page is an image.
What are my options, from one file to thousands?
| Option | Best for | What it produces | Where it falls short |
| The OCR command in a PDF editor (for example Adobe Acrobat) | A few files, done by hand | A searchable PDF with a text layer | One file at a time; the batch features are usually in paid tiers |
| Opening the PDF in Google Docs from Google Drive | Occasional single files | An editable document with the recognised text | Formatting and page layout are often lost; not built for batches |
| An open-source batch tool such as OCRmyPDF (built on Tesseract) | Hundreds to thousands of files on your own machine | Searchable PDFs that keep the original page images | Needs a command line; tables, handwriting and poor scans need care |
| A cloud document OCR service (Google Document AI, Amazon Textract, Azure AI Document Intelligence) | Large volumes, tables and forms | Text plus layout: blocks, tables, key-value fields | Priced per page; you build the pipeline around it |
| A search platform that runs OCR at ingest | Searching the whole collection, not just making files searchable | An index you can query across every page | You adopt a platform instead of producing files |
How do I get OCR right on real scans?
OCR accuracy depends mostly on the page image. The fixes that matter most:
To measure it, pick twenty pages that represent your collection, correct their text by hand once, and compare the OCR output against that. It tells you whether the problem is the scans or the tool.
How do I search across all of them, and by meaning?
Once the text exists, put it in a search index with a reference back to the file and the page. Two kinds of search are worth having together:
Running both and merging the results, called hybrid search, covers the exact lookups and the questions asked in the reader's own words. Some answers are not in the text at all, such as a value read off a chart; for those, see why document search misses the answer in a chart or table.
How do I do this with Mixpeek?
Put the PDFs in object storage you already use, such as S3 or GCS, and connect the bucket to Mixpeek. The document layout extractor runs OCR with layout detection on every page, splitting it into paragraphs, tables, forms and headers, and sends low-confidence blocks from degraded scans to a vision-language model for correction. Each block is indexed with its page and position, so a retriever returns the passage and the page it came from, with keyword and semantic search in one query. At the published rate of $1.50 per thousand document pages, indexing 10,000 scanned pages costs about $15 in processing, before storage and queries.
Related: how OCR works, from detection to reading order, the best OCR APIs, the best PDF extraction tools and why search can't find exact part numbers, codes or names.
Frequently Asked Questions
How can I tell if a PDF is scanned or has real text?
Try to select a word with the cursor, or search for a word you can see on the page. If you cannot select it and search finds nothing, the page is an image and the PDF needs OCR before it can be searched.
What is the best free way to OCR a lot of PDFs?
An open-source batch tool such as OCRmyPDF, which uses the Tesseract OCR engine, adds a text layer to each PDF while keeping the original page images. It runs on your own machine, handles whole folders, and has options to straighten pages and set the language.
Does OCR work on handwriting?
Poorly, with classic printed-text OCR. Cloud document services and vision-language models read handwriting better, but accuracy varies with the writer, so check a sample of your own pages before relying on it.
Why does my OCR mix up columns or tables?
Plain OCR reads straight across the page, so two columns interleave and a table turns into a line of numbers. Use layout-aware OCR, which detects columns, tables and blocks first and reads each one in its own order.