Free plan: 1 conversion/hour, 1 file at a time
Go Unlimited →

OCR PDF

Read the Words Inside PDF documents

Choose your files

*Files deleted after 24 hours

Transform files free, Pro users can convert much larger files; Sign up now

Uploading

0%

How to OCR PDF

1 Upload the scanned PDF documents; a clean 300 DPI grayscale scan gives the recogniser the most to work with.
2 Tell it which language to expect — naming the language beats letting it guess by a wide margin.
3 Run the recognition pass; the words are written back as an invisible layer sitting over the original page image.
4 Download the searchable PDF file — it looks identical, but you can now search and select the text in it.

OCR PDF FAQ

Does the source format affect recognition?
+
It does. PDF fixes the page: fonts, vectors, raster images and text coordinates are frozen so every reader sees identical layout. How the page is stored decides what resolution and colour information the recogniser has to work with.
Yes — a PDF is an object graph, not a page image, so text stays selectable and vectors stay sharp no matter what happens to the raster content inside it. It affects what the recogniser can see.
On a clean 300 DPI scan of printed text, high enough that errors are rare and mostly in unusual proper nouns. Accuracy falls off sharply with low resolution, skew, shadows from a phone camera, unusual typefaces and handwriting — handwriting in particular is not what this is for.
Concretely, each page is rendered, run through a Tesseract text-recognition pass in over 100 languages, and the recognised words are written back as an invisible text layer positioned over the original image — so the page looks unchanged but is searchable and selectable. The page still looks exactly as it did — the recognised text sits invisibly behind the image so search and selection work without changing the appearance.
Yes — upload the set and they process in parallel under one set of settings, which is the point of doing a document workflow here rather than clicking through a desktop reader.
Outlines, internal links and annotations are preserved wherever the operation allows it. Digital signatures are the exception: any change to the file necessarily invalidates a signature, because that is precisely what a signature is for.
Yes: free accounts process documents up to 5 MB each, which covers most reports and contracts but not a long scanned document at 600 DPI; Ghostscript and qpdf do the work. Scanned PDFs are the usual thing that exceeds it, and compressing the document first is normally enough to bring it back under.
Text and vector artwork are objects rather than pixels, so they stay perfectly sharp at any zoom no matter what happens. Only the embedded raster images can degrade, and only if the operation you chose resamples them.
WORD.to is built around the editable end of a document's life — the DOCX that is still being written, still being styled and still being argued over, before anyone flattens it for sending. Office files are containers full of other people's media — images, embedded audio, fonts — so the work people need on them is usually the work they would need on those contents anyway. OCR PDF shares the upload, the caps and the account with the conversions for that reason.
The converter on this site takes documents out to PDF for sending, to images for embedding and to plain text for anything that has to read them programmatically, and back the other way. Doing that afterwards keeps the editable original around, which is the part you cannot get back once it has been flattened.
The engines are shared — the same document toolchain, the same workers, the same limits. What a Word site adds is a view on what survives leaving the Office format and what does not, which is the question every one of these jobs actually turns on. It also starts from one fact about the format this site is named after: the document is a zip of XML, so text edits are cheap and the weight is almost always the embedded images.
No account, and nothing is kept: uploads are deleted from the workers shortly after the job finishes, nothing is read and nothing is indexed. Free accounts exist for history and batch size, not for access.

Rate this utility
5.0/5 - 0 votes
ns6.com — Buy your domain in minutes. Search, register, launch.
Or release your files here