Converting files
How to convert a PDF to Word (and what usually goes wrong)
Upload the PDF to a converter, let it read the text and rebuild the layout, then download the .docx, but the result is only editable if the PDF contains real text rather than a scanned picture of text.
Almost everyone who converts a PDF to Word is doing it for one of two reasons. Either somebody sent a document and did not send the original, or the original is long gone and a small change is needed. Both are ordinary problems. The frustrating part is that the same conversion can produce something perfect one day and something unusable the next, with no obvious explanation.
The explanation is usually the file itself, not the converter. It is worth understanding, because it tells you what to expect before you spend any time on it.
What a PDF actually contains
A PDF is not a document in the sense a Word file is. It has no paragraphs, no headings and no concept of "the next line". What it has is a list of instructions: put this glyph at this coordinate, in this font, at this size. Draw a line from here to here. Paint this image inside this rectangle.
That is why PDFs look identical everywhere (the positions are absolute) and it is also why converting one back is genuinely hard. A converter has to look at thousands of individually placed characters and work out, after the fact, that these ones were a paragraph, those ones were a table, and that gap over there was a column break rather than a wide space.
Word, meanwhile, wants flowing content. It wants to know that this is a paragraph so it can rewrap it when you type. The whole conversion is an act of reconstruction.
The two kinds of PDF, and why it matters so much
Before converting anything, it helps to know which kind you have.
A text PDF was made by a program: exported from Word, printed to PDF from a browser, generated by an accounting system. The characters are stored as characters. You can select a sentence with your mouse and copy it. These convert well.
A scanned PDF is a photograph. Someone put paper on a scanner, and the result is an image of a page wrapped in a PDF container. There is no text in it at all, only pixels that happen to look like text. Try to select a word and you will select the whole page as a picture.
The quickest check takes three seconds: open the PDF and try to drag-select a line of text. If a blue highlight follows your cursor along the words, it is a text PDF. If the entire page highlights as one block, it is a scan.
Converting a text PDF
For a text PDF the process is short:
- Open PDF to Word and drop the file onto the panel.
- Wait while the page is read. The converter is measuring where every character sits, which font it uses and how the lines group into paragraphs and tables.
- Download the .docx and open it.
What you should get is a document where the words are the right words, in the right order, in something close to the right place, with the fonts that the PDF actually used rather than a generic substitute.
What you should not expect is a file identical to whatever produced the PDF originally. That file had structure the PDF never recorded. If the original had a numbered list with automatic numbering, the PDF stored the numbers as plain characters, and no converter can know they were once automatic.
Why your fonts sometimes change
This one causes more complaints than anything else, and the cause is worth knowing.
PDFs usually embed a subset of each font, only the characters the document actually uses, renamed with a random prefix so it does not clash with anything installed. A font that started life as Times New Roman can arrive inside the PDF called something like ABCDEF+TimesNewRomanPSMT. A converter that reads that name naively sees an unknown font and falls back to a default, which is how a serif document becomes an Arial document.
A better approach is to look inside the embedded font file itself, where the real family name is recorded, and use that. It is the difference between a converted document that looks like the original and one that looks vaguely related to it.
What happens to tables
Tables are the other place conversions go wrong, and the reason is the same as before: PDFs do not have tables. They have text and they have lines. A converter has to notice that these lines form a grid, and that these blocks of text sit inside its cells.
When the table has visible borders this works well. When it does not (when the "table" was just text aligned with tab stops) the converter has to guess from the alignment alone, and sometimes it guesses wrong. If a table comes across as loose text, that is usually why.
If your PDF is a scan
You have two honest options.
Accept it as an image. The Word file contains a picture of each page. You cannot edit the words, but you can add your own text around them, and the document prints exactly as before. For a signed contract you only need to countersign, this is often all you wanted.
Run OCR on it. Optical character recognition looks at the pixels and works out what letters they represent. PDF to Text does this and gives you the words as real, editable text. What it cannot give you back is the layout. OCR recovers the writing, not the design. Expect to do the formatting yourself.
Be realistic about accuracy. A clean 300 DPI scan of printed text is recognised very well. A phone photo of a page at an angle, in poor light, with a coffee stain, is not. Handwriting is harder still.
Practical advice before you convert
A few things genuinely improve the result:
- Find the original. If the PDF came from a colleague, ask whether they still have the Word file. Thirty seconds of asking beats an hour of repair work.
- Check the page count. Long documents take longer and give the layout engine more chances to make a decision you disagree with. If you only need three pages, extract those three pages first.
- Convert once and check the whole thing. People check page one, decide it worked, and discover the problem on page eleven after an hour of editing.
- Keep the PDF. The converted file is a copy, not a replacement. If something is wrong, you want the source to check against.
When a converter is the wrong tool entirely
Sometimes you do not actually want Word. You want to change one number, cross out a clause, add a signature or fill in a form. Converting to Word, editing and converting back will change the layout twice and introduce errors both ways.
For those jobs, editing the PDF directly is both faster and safer. The PDF Editor lets you type on the page, add images and sign, then save a real PDF with everything else untouched. If the file is a scan, the OCR PDF is built for exactly that case.
The rule of thumb: if you are rewriting the document, convert it. If you are annotating it, do not.
Common questions
Will the converted Word file look exactly like the PDF?
Close, but rarely identical. The text, fonts and positions come across; things the PDF never recorded (automatic numbering, styles, the original table structure) cannot be recovered because they were not stored. Expect to tidy a few things.
Why is my converted document just a picture?
Because the PDF was a scan. There was no text to extract, so the honest result is the page as an image. Run OCR on it if you need the words as editable text.
Can I convert a password-protected PDF?
Not while it is locked. Remove the password first with Unlock PDF (you will need the password it was locked with) then convert.
Is there a page limit?
Files up to 200 pages and 80 MB convert. Longer documents are best split first.