Searchable PDF (OCR)

Add an invisible text layer to a scanned PDF so it's searchable and selectable.

About this tool

Drop a scanned PDF, choose the language, run: every page is read by the open-source Tesseract engine in your browser and rebuilt with an invisible text layer over the words. The pages look identical; Ctrl+F, selection and screen readers now work.

What a text layer is, and what OCR gets right

Office scanners with an โ€œOCRโ€ option produce PDFs that look like scans but let you search them; the trick is a second, transparent layer of text positioned over the picture of each word. This tool does the same thing: each page is rendered at 150 DPI, read by Tesseract with word positions, and re-saved as the page image plus invisible text sized to each word's box โ€” so a search hit highlights the right place on the scan. The engine and language data (about 9 MB) download from the jsDelivr CDN on first use and cache; the PDF itself never leaves your device, unlike the OCR sites that upload every page. Honest limits: accuracy follows scan quality (clean 300-DPI office scans read near-perfectly, faxes and phone photos worse, handwriting rarely), pages are re-saved as images, and the built-in font can position Latin-script text only โ€” for other scripts, the PDF OCR tool gets the text out as a file instead. Check first whether the PDF is actually a scan: if text selects, it already has a layer and extract text is instant. Shrink big results with the compressor.

Frequently asked questions

How do I know if my PDF needs this?

Try selecting text in it. If nothing highlights, it's a scan and this tool applies. If text selects, it already has a text layer.

Will the pages look different?

No โ€” the scan is kept as the page image; the text layer is invisible. Search hits highlight over the right words.

Is the document uploaded?

No. Only the OCR engine and language data are downloaded (once, from a CDN). Pages are rendered, read and rebuilt in your browser.

How accurate is the search?

As accurate as the OCR: near-perfect on clean printed scans, weaker on faxes, skewed phone photos and unusual fonts. A missed word is a word search won't find โ€” rescan at 300 DPI if it matters.