OCR a PDF
Make scanned PDFs searchable
Extract the text a PDF already stores and download it as a plain text file, with a marker at the start of every page.
Gone within the hour.
Files are uploaded over HTTPS, used only to run this tool, and deleted from our server automatically about an hour later. Never sold, never used to train anything.
Add the PDF
One file, up to 50 MB. The pages render so you can confirm you have the right document.
Choose the pages, or leave it blank
Blank means the whole file. Otherwise the box takes ranges like 1-3, single pages like 5, and open-ended ones like 8-. Up to 500 pages in one job.
Press Extract text
This reads the text the PDF already stores, so it is fast and exact: what comes out is what is in the file.
If it comes back empty, it is a scan
A scanned page is a picture and stores no text at all. Rather than hand you an empty file, the tool says so and points you at OCR PDF, which adds the text layer first.
Because the pages are pictures rather than text, which is what a scan or a photographed document is. Run OCR PDF on it first to add a text layer, then extract the text from that result.
No. A plain text file has no columns, tables, fonts or images. If you need the layout, convert to Word instead.
Yes. Type ranges like 1-3, 7, 12- in the pages field. Leaving it blank extracts the whole document, up to 500 pages a job.
For ordinary documents, yes: the text is sorted into reading order rather than the order the file happens to store it in. Heavily designed pages with side panels can still interleave.