Working with scans
How do I copy text from a PDF when copying doesn't work?
Open the file in PDF to Text instead of selecting. Pages with real text are read directly, and pages that are only pictures of text are recognised with OCR. You can then copy one page, copy everything, or download a .txt file.
You need three paragraphs out of a PDF. You drag across them and nothing highlights. Or the whole page turns blue as one block. Or it copies, and what you paste is a line of boxes and symbols.
Each of those is a different problem, and knowing which one you have tells you what will work.
What the symptom tells you
| What happens | What it means | What works |
|---|---|---|
| Nothing highlights, or the whole page selects as one block | The page is a picture of text, usually a scan | PDF to Text, which recognises it with OCR |
| It pastes as boxes, symbols or wrong letters | The text is there, but its font has no usable character map | Read the picture instead: PDF to Image, then Image to Text |
| Your viewer says copying is not allowed | The file carries a copying restriction | See below |
| It pastes with a line break after every line | Normal: a PDF stores lines, not paragraphs | Join the lines in a text editor |
Getting the text out, step by step
- Open PDF to Text.
- If the document is scanned, set Document language. The default, Detect automatically, works for most files. You can also choose Bangla, English, Hindi, Arabic, Chinese (Simplified), Russian, Spanish or Dutch; each is read together with English. The setting only affects scanned pages.
- Choose the PDF. Every page is read; the status line shows how many are ready and how many needed OCR.
- Take the text out with one of the three controls below.
- Copy, in the top corner of each page, copies just that page.
- Copy All Pages copies the whole document.
- Download as .txt saves everything to a text file, with each page under a "Page N" heading.
On a phone, Copy All Pages and the .txt download sit at the bottom of the screen.
There is no selecting involved, so it does not matter whether your viewer can highlight the text or not.
Where the work happens
Pages that already contain real text are read by your browser, and for those the file never leaves your device. Pages that turn out to be pictures are drawn as images and sent together to our server to be recognised, and the text comes back.
Recognised text is an estimate. On a clean scan it is usually very good; on a faint or crooked page it will have mistakes. Check names and numbers before relying on them.
When the text pastes as gibberish
This happens with some PDFs made by older software or unusual fonts. The letters look right on screen, but the file does not record which character each shape is. Copying, and any tool that reads the text layer, including PDF to Text, gets the same meaningless codes.
The way round it is to ignore the text layer and read the picture:
- Convert the pages you need to images with PDF to Image, choosing PNG.
- Put those images through Image to Text, which reads up to five images at a time.
When a scan has a little real text on it
PDF to Text decides page by page: a page with any real text is read as text; a page with none is sent for OCR. A scan with a small typed stamp on it, such as a page number or a date added by the scanner, counts as text, and you get just the stamp.
If a scanned page comes back almost empty, use the same PDF to Image and Image to Text route for that page.
When copying is "not allowed"
Some PDFs carry a flag asking viewers not to allow copying. That is a restriction on the file, not a technical barrier to reading it, and it is covered in detail in My PDF opens but won't let me print or copy.
If it is your document, or you have the right to use its text, Unlock PDF removes the protection from your copy. A PDF that needs a password just to open has to be unlocked first; PDF to Text does not ask for passwords.
About the line breaks
The text comes out one line per line, as it was laid out on the page. That is how PDFs store text. A paragraph pasted into an email arrives with a break at the end of every line.
To rejoin it, paste into a word processor and replace single line breaks with spaces. On pages with two or more columns, check the order too. The lines are read in the order the file stores them, which can run across columns rather than down them.
Common questions
Why can't I highlight any text in my PDF?
Because the page is a picture of text, usually a scan or a photo, with no text layer. PDF to Text recognises such pages with OCR automatically.
Why does copied PDF text paste as symbols or boxes?
The font in the PDF has no usable character map, so the copied codes mean nothing. Convert the page to PNG with PDF to Image and read it with Image to Text instead.
Is the extracted text exact?
For pages with real text, yes: the characters are read directly. For scanned pages it is an OCR estimate, so check names and numbers.
Can I copy just one page?
Yes. Every page has its own Copy button. Copy All Pages and Download as .txt give you the whole document.
Is my PDF uploaded?
Pages with real text are read in your browser and never leave your device. Only pages that need OCR are sent, as images, to be recognised.
Which languages can it recognise in scans?
Bangla, English, Hindi, Arabic, Chinese (Simplified), Russian, Spanish and Dutch, or it can detect the language automatically. Each is read together with English.