About OCR PDF
A scanner or a phone scanning app saves each sheet as a picture. A PDF viewer sees only pixels, so a search for a name finds nothing and dragging the cursor grabs the whole page instead of a sentence. Optical character recognition (OCR) studies those pixels and works out which letters they show.
On this page, every page of your file is drawn at 300 DPI with pdf.js and passed to Tesseract, the open-source OCR engine, in its browser build tesseract.js (Apache-2.0 licence). The words it recognises are written back into the PDF as an invisible text layer, positioned over the matching spots of the original page. The visible page is left alone: the scan, its colors, stamps and signatures look just as they did. What changes is that the document now answers a search, and you can highlight and copy from it like any PDF made on a computer. If you only want the words, everything recognised is also offered as a plain .txt file.
Recognition languages
Tesseract reads more accurately when it knows which language to expect. English is selected by default. Spanish, French, German, Portuguese, Italian, Chinese (Simplified), Chinese (Traditional) and Japanese are available as well. Each language model is a download of roughly 0.7 to 3 MB from pd00.com the first time you choose it; your browser caches it for later runs.
How to use OCR PDF
- Choose the scanned PDFSelect the file or drop it onto the tool. A password-protected scan has to go through Unlock PDF beforehand, and that step needs the password itself.
- Tell it which languages to readKeep English or select the languages the document is written in, up to three for pages that mix them.
- Recognise, then downloadStart OCR and keep the tab open while it works through the pages. Then save the searchable PDF, plus the .txt file if you want just the words.
When people use it
Finding a clause in a signed contract
A signed agreement came back as a scan. After OCR, searching for a party’s name or the word termination jumps to the right paragraph instead of paging through dozens of images.
Paper records you want to find years later
Invoices, warranty cards and letters scanned for safekeeping can be found by what they say, not only by file name, once each carries a text layer your computer’s file search can index.
Quoting from a photocopied chapter
A scanned book chapter or journal article in a course reader can’t be quoted without retyping. With OCR you select the passage, copy it, then compare it against the page image to catch misread characters.
What decides how accurate the text is
OCR can only read what the scan shows, so the source matters more than any setting. The best results come from pages captured at 300 DPI, lying straight, with dark print on a light background. Accuracy drops with:
- Low resolution or blur. Rendering at 300 DPI can’t restore detail a low-resolution scan never captured, and small print in a shaky phone photo has too few pixels per letter.
- Crooked or sideways pages. Tilted lines and the curve near a book’s spine pull words apart. Turn pages that lie on their side upright with Rotate PDF before you start.
- Handwriting. Tesseract is built for printed and typed text. Handwritten notes, signatures and filled-in boxes are not recognised reliably.
- A missing language. A German letter read with English alone tends to lose its umlauts and misjudge words, so select every language the pages use.
Even a clean scan will contain the odd mistake, such as a 1 read as an l or rn read as m. Search still finds most words, but proofread anything you copy into a document that matters.
How long recognition takes
Expect several seconds per page on a laptop and longer on a phone; a two-page letter is done quickly, while a long report can take minutes, and the tab has to stay open until it finishes. Long files are better handled on a computer. When you only need a section of a big scan, cut those pages out with Extract PDF Pages and run OCR on the shorter file.
Skipping pages that already contain text is on by default and saves time on mixed documents, such as a typed report with scanned signature sheets at the end: only pages without text are recognised. Switch it off to have every page read.
What the searchable copy keeps
Nothing is cleaned up, straightened or added to the visible page, so the result looks the same on screen and on paper as the scan you added. The added text is tiny next to those images, so the file usually grows only a little. If you also want the PDF smaller, run it through Compress PDF as a separate step; to give it a clear title for your archive, set one with Edit PDF Metadata.
Files that don’t need OCR
- PDFs that already have selectable text. If you can highlight words in your viewer, skip OCR and copy the words with PDF to Text, or keep headings and lists with PDF to Markdown.
- Documents you want to edit. OCR makes a scan searchable, not editable. Each page stays a picture with text behind it, so you can’t retype a sentence in place.
Pages are rendered and read inside this browser tab. The only extra downloads are the language models, fetched from pd00.com, and neither your scan nor its text is sent anywhere; here is how to check.