PDF OCR
Turn a scanned PDF into text, page by page, with OCR that runs in your browser.
About this tool
For PDFs that are pictures of pages โ scans, faxes, photographed documents. Each page is rendered in your browser and read by the open-source Tesseract engine; the text comes out page by page with markers, ready to copy or download. Nothing is uploaded.
When to OCR, and when not to
First, check whether you need this at all: if you can select text in the PDF, it already has a text layer and the extract-text tool gets it instantly and exactly. OCR is for the rest โ the PDF a scanner or a phone camera made, where selecting text grabs nothing. Each page renders at 200 DPI on your device and goes through the engine; expect a few seconds per page on a laptop, longer on a phone, and accuracy that follows scan quality: clean 300-DPI office scans read almost perfectly, faxes and skewed phone photos less so, handwriting rarely. The engine and language data (~9 MB) download from the jsDelivr CDN on first use and stay cached; your document never leaves the device โ a real difference from the OCR sites that upload every page. Once you have text, the word counter and text to PDF take it onward.
Frequently asked questions
How do I know if my PDF needs OCR?
Try selecting text in any PDF viewer. If nothing highlights, the pages are images and need OCR. If text selects, use the plain extract-text tool โ it's instant and exact.
Is the PDF uploaded?
No โ pages are rendered and read entirely in your browser. Only the engine's code and language data are downloaded (once, from a CDN), never your file.
How accurate is it?
On clean scans of printed text, near-perfect. Faxes, low-resolution scans and phone photos at an angle lose accuracy; handwriting mostly fails. Rescan at 300 DPI if you can โ it's the biggest lever.
How long does it take?
A few seconds per page on a laptop, longer on phones โ a 30-page scan is a coffee break. The text appears page by page as it's read, so you can start copying early.