Converting files
How to turn a web page into a PDF that actually looks right
Use a converter that renders with a real browser engine and turn on background graphics. The two reasons web pages usually print wrong are a weak rendering engine and the browser's default of dropping backgrounds.
Saving a web page as a PDF sounds like it should be trivial. The page already exists, it already looks a particular way, and all anyone wants is that same thing in a file. Yet the result is so often disappointing (colours gone, boxes stacked in the wrong order, half the design missing) that people assume PDFs are just bad at this.
They are not. There are two specific causes, both fixable.
Cause one: the thing doing the rendering
A web page is a program as much as a document. Modern layout uses CSS grid and flexbox to position things relative to each other, media queries to change the design at different widths, and web fonts loaded separately from the page itself.
Rendering that correctly means running a browser engine. Not something that reads HTML and approximates it. An actual engine that does layout the way Chrome or Firefox does.
Plenty of converters do not. Some pass the HTML to an office suite's import filter, which understands roughly the CSS of two decades ago. Feed a modern page into one of those and the grid collapses to a single column, the flex rows stack vertically, and anything positioned lands somewhere unrelated. The output is not a bad rendering of the page; it is a rendering of a different page.
HTML to PDF uses a real browser engine, which is why grid, flexbox, web fonts and media queries all survive.
Cause two: browsers throw away your backgrounds
This one surprises people who have already got a good engine.
Every browser, by default, does not print background colours and background images. Open a page with a dark theme, press Ctrl+P, and the preview is white with black text.
This default is old and was well-intentioned: in 2003, printing a page with a black background wasted an entire ink cartridge. It is now the single most common reason a saved page looks wrong, because modern design uses background colour structurally. Cards, banners, table headers, coloured sections. All of it is background, and all of it vanishes.
The setting exists in every browser's print dialog, usually under "More settings" as Background graphics. Ticking it fixes the problem instantly.
Because "I want this page as it looks" is what people actually mean, HTML to PDF turns backgrounds on by default and lets you switch them off if you specifically want the ink-saving version.
The setting almost nobody thinks about: screen width
Modern sites are responsive. The same page has several different layouts and picks one based on how wide the window is. A three-column dashboard at 1440 pixels becomes a single stacked column at 375.
So "convert this page to PDF" is an incomplete instruction. Which layout?
This matters more than people expect. Convert at a narrow width and you get the mobile design, usually taller, simpler, with navigation collapsed into a menu button that is useless in a PDF. Convert at a wide width and you get the full desktop layout, which is normally what someone means when they say they want the page.
Being able to choose is genuinely useful. Capturing a site's mobile layout to show a client is a real task, and so is capturing the desktop one for a report.
One long page, or proper pages?
There are two sensible outputs and they suit different jobs.
Paginated cuts the content into A4 or Letter pages. This is what you want if the PDF will be printed, emailed as a document, or filed alongside other documents.
One long page produces a single sheet as tall as the content. Nothing is cut across a page break, which matters when the page has a long table or a diagram that a break would ruin. It is the closest thing to a screenshot that still has selectable text.
Neither is more correct. Ask what happens to the file next.
What cannot be captured
Being straightforward about the limits saves time.
Pages behind a login. A converter fetches the page from its own server, which has no access to your browser session or cookies. It sees what a logged-out visitor sees. If you need a page from inside an account, save it from your own browser (Ctrl+S, "Webpage, Complete") and upload the resulting .html file, or the folder zipped up.
Anything that needs interaction. Content that only appears after clicking a tab, expanding an accordion or scrolling to trigger loading will not be there. The converter loads the page and renders it; it does not use it.
Addresses on a private network. A URL pointing at localhost or an internal IP is refused, deliberately. Allowing a server to fetch arbitrary internal addresses on a stranger's request is a well-known security hole, and the check is applied again on every redirect.
Four ways in
Different situations call for different inputs, and all four are worth knowing:
- A web address for any public page.
- An .html file you have saved or been sent.
- A .zip of a whole site: the page plus its stylesheets, images and fonts, so everything resolves properly. This is the right choice for a page saved from a browser, which produces exactly that structure.
- Pasted markup for a snippet, an email template or something you are building and want to see on paper.
Getting a cleaner result
A few habits improve the output noticeably:
- Use the site's own print view if it has one. Many publications offer a print or reader mode that strips navigation and ads. That version usually converts beautifully.
- Convert the article, not the homepage. Homepages are full of carousels and dynamically loaded panels that make poor documents.
- Check the margins. "None" is right for a design you want edge to edge. "Normal" is right for anything that will be read as a document.
- Try landscape for wide content. A data table or a dashboard often fits far better rotated.
Common questions
Why did my dark web page come out white?
Because browsers drop background colours when printing. Turn on background graphics. This tool does it by default, so if you see a white page, check that setting has not been switched off.
Can I convert a page that needs me to log in?
No. The page is fetched by a server with no access to your session. Save the page from your own browser and upload the .html file or a .zip of it instead.
Why was my address rejected?
Addresses on private networks (localhost, internal IP ranges and cloud metadata endpoints) are blocked as a safety measure, along with anything that is not http or https.
Will the links in the page still work in the PDF?
Links to other pages are preserved as clickable links in most cases. Anything that depended on JavaScript running will not, because a PDF cannot run scripts.