What OCR does (and does not) do
When you scan a document or photograph a page, the scanner records pixels — light and dark shapes — not letters. OCR runs a recognition engine over those pixels, identifies the patterns that make up characters, and writes the recognized words into a hidden text layer positioned over the page. The result looks identical to the original scan, but now you can press Ctrl+F to find a word, select a sentence to copy it, and let search engines index the document.
PDFZento uses tesseract.js, a full OCR engine running as WebAssembly directly in your browser. Because recognition happens on your device, your document never leaves it — important for scans of contracts, invoices, and personal records.
Image-only vs. searchable PDFs
- Image-only PDF — created by scanners, faxes, or "print to PDF" from a photo viewer. Contains no text layer: nothing is searchable or selectable. This is the file type OCR is built for.
- Born-digital PDF — exported from Word, Excel, or a design tool. Already contains a real text layer; OCR has nothing to add.
- Partially searchable PDF — sometimes the body text is real but inserted images of text are not. OCR the whole document once to unify it.
Which languages are supported
PDFZento ships OCR language models for English, German, French, and Spanish. If a document mixes languages, run it once with the dominant language, then re-run if a second language is mixed in.
How to OCR a PDF step by step
- Open the OCR tool. Go to OCR PDF in your browser.
- Upload the scanned PDF. Drag and drop the image-only document or choose it from your device. Files are processed in browser memory, not on a remote server.
- Select the language. Pick English, German, French, or Spanish to match the document. The engine downloads the matching language pack on demand.
- Run recognition. Start the process and wait while the engine reads every page. Larger documents take longer, but your device does all the work.
- Verify and download. Search for a term or select some text to confirm the layer is present, then download the searchable PDF. Your original file is untouched.
Maximizing OCR accuracy
The recognition engine can only read what the scan shows clearly. Follow these rules before scanning to keep accuracy high:
- Scan at a solid resolution — text-heavy pages at 300 DPI and above generally recognize far better than lower resolutions.
- Keep pages flat and straight. Skewed pages make the engine misread lowercase letters like
rnvsmorclvsd. If the tilt came from the scanner or copier rather than the original, deskew scanned pages first and then run OCR. - Crop out margins and shadows, and brighten the scan so letterforms are crisp and high-contrast.
- Prefer clean sans or serif body fonts over dense handwritten text; handwriting recognition is dramatically less reliable than printed text.
Why "no text found" can still mean OCR worked
If a page contains only a photo, a logo, or decorative elements with no letters, the engine legitimately finds nothing to recognize — that is expected, not a failure. Try searching for a word you can clearly read in the body text rather than a graphic heading.