PDF to Text
Pull the words out as a plain .txt
Turn flat scans into selectable, searchable text, with each page straightened and cleaned up before it is read.
Gone within the hour.
Files are uploaded over HTTPS, used only to run this tool, and deleted from our server automatically about an hour later. Never sold, never used to train anything.
Add the scan
Drop the PDF onto the page, or click to browse for it. One file at a time, up to 50 MB and up to 60 pages. If your scan is longer than that, split it first and run the parts separately.
Say what language the scan is in
Pick from the 13 in the list. English is already selected, so this is one click if that is what you have. It is the only choice that changes the result: the recogniser uses that language's own letter shapes and vocabulary to decide between characters that look alike.
Press Make it searchable
The file uploads and every page is read. A long scan takes a few minutes, and the page says where it has got to rather than leaving you with a spinner.
Download the searchable copy
It looks exactly like the scan you started with. The recognised text sits in an invisible layer underneath the image, so Find works, copy and paste works, and nothing about the page's appearance changes. Any page that already contained real text is left untouched.
It reads the scanned pages and adds an invisible text layer underneath, so the words become selectable, searchable, and copy-pasteable. Your picture of the page is kept, and the file comes back as an ordinary PDF rather than being converted to PDF/A, so it opens and behaves the way it did before. It is not left untouched, though: crooked pages are straightened and sideways ones turned upright, so a scan can come back tidier than it went in.
Usually, and all three fixes are gated on a measurement rather than applied blindly. Straightening: the common approach is to ask the recognition engine itself for the skew angle, but it works that angle out with the same page analysis it uses to read, so it gets least reliable on badly skewed pages, which are exactly the ones that needed it. Pages here are measured geometrically instead, by searching for the rotation that packs the ink into the sharpest rows, to the nearest quarter degree and out to 15 degrees either way. A page sitting less than half a degree off is left alone, because the resampling would cost more than the tilt does. Speckle: salt and pepper noise blinds the engine's own thresholding, so it is cleaned first, but only on a page where isolated single pixels are more than 3 percent of the ink. A clean page measures zero there and is passed straight through. Uneven light: a phone photo, or a scan made with the lid up, carries a broad shadow, and one global cutoff loses the text inside it. The paper colour is estimated across the page and subtracted, again only when that estimate varies enough for it to matter. The order matters too, because speckle and shadow both read as ink: each is taken off the copy the angle is measured on, so neither drags the straightening to a false answer. All three change only the image the reader sees, never the page in the PDF you get back.
Because straightening it would have made it harder to read, not easier. Rotating a page resamples its specks into soft blobs that the noise filter can no longer clear, so on a badly speckled scan the correction can cost more than it wins. That is checked rather than assumed: on a page found to be speckled, the angle is applied only if the cleaned-up image that would come out of the rotation still comes back clean, and otherwise the page is deliberately left crooked and read as it is.
You are told, rather than handed an empty result. A scan that could not be read still produces a perfectly valid PDF, so the easy thing is to return it and report success, and what you get is a file called searchable with nothing in it to find. When no text at all comes back from a file that did contain scanned pages, the job fails and says so, and suggests straightening or cleaning the scan first. A blank or born-digital PDF is not treated as a failure: it comes back unchanged.
Thirteen: English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Arabic, Hindi, Urdu, Simplified Chinese, and Japanese.
Up to 60 pages and 50 MB per job, and a scan that takes more than about eight minutes to read is stopped rather than left running. Split bigger PDFs first, then OCR the parts. The same 60 page cap reaches PDF to Word whenever the file is mostly scanned, because that conversion runs this OCR first.