How to Convert a PDF File to HTML (Clean Web Pages)
Turn a PDF into a clean HTML page: extract the text, use pdftohtml or Calibre, or embed instead. An honest look at what survives the conversion.

If you want to convert a PDF file to HTML, the honest news up front: no tool turns an arbitrary PDF into clean, responsive HTML automatically. PDF is a print format — it stores "put this glyph at x=142, y=380," not "this is a heading." HTML is the opposite. Every conversion is a translation between two philosophies, and something always gets lost.
That doesn't mean it's hard. It means you should pick your method based on what you actually need: clean markup, fast one-off output, or just getting the document onto a web page at all. Here are all three, plus the case for not converting.
Why turn a PDF into a web page at all?
Three reasons come up constantly:
- SEO. Search engines index PDFs, but rank real HTML pages far better: HTML has structure (headings, links, alt text), loads faster, and works on mobile. Content buried in a PDF is content most searchers never find.
- Accessibility. Screen readers handle semantic HTML far better than the average untagged PDF. If the document matters to everyone, HTML is the accessible format.
- User experience. A PDF link on mobile means download → open in another app → pinch-zoom a fixed A4 layout. An HTML page just reads.
Method 1 — Extract and rebuild (cleanest markup)
For content you actually care about — a report going on your site, a guide, documentation — this beats every automatic converter. You take the raw material out of the PDF and write the simple HTML around it yourself.
Pull out the text
Drop the PDF into Extract Text. You get the full text content, free up to 20 MB and 50 pages, no sign-up.
Pull out the images
Run the same PDF through Extract Images — every embedded image comes out as a separate file at original resolution, ready for img tags.
Mark up the structure
Paste the text into your editor and add the semantics the PDF never had: h2 for section titles, p for paragraphs, ul for lists, table where needed. This is the manual part — typically 15-30 minutes for a 10-page document.
Drop the images back in
Add img tags with real alt text where the figures belong. Compress images for the web while you're at it.
Publish
The result is hand-quality HTML: responsive, accessible, indexable — none of which an automatic converter gives you.
Yes, it's manual work. But the output is HTML you'd actually want to maintain, and for anything destined for real readers, that trade is worth it.
Method 2 — pdftohtml and Calibre (for the command line crowd)
If you want automation and can live with messier output:
pdftohtml (part of the Poppler utilities, preinstalled or one package away on Linux/Mac):
pdftohtml -s -noframes document.pdf output.html
The -s flag gives one continuous HTML document instead of one file
per page. The -c flag adds "complex mode," which reproduces layout
using absolutely-positioned divs — visually closer to the PDF, but
the markup is unusable for anything except display and breaks
completely on mobile.
Calibre (free, cross-platform) treats the PDF as an ebook: open Calibre → Add the PDF → Convert → choose HTMLZ output. It works best on single-column, text-heavy PDFs and struggles with multi-column layouts, which it often reads in the wrong order.
Paid converters (Adobe Acrobat Pro's Export-to-HTML, various online services at $6-20/month) do better at tables and preserve more formatting, but their output is still generated markup — hundreds of inline styles and nested spans. Fine if you just need the content online today; painful if you ever edit it.
Method 3 — Don't convert: embed or link the PDF
Sometimes the right answer to "how do I put this PDF on my website" is to put the PDF on your website:
<embed src="/files/report.pdf" type="application/pdf" width="100%" height="800px" />
Or just a plain link with a size hint:
<a href="/files/report.pdf">Annual report (PDF, 2.4 MB)</a>.
The downsides: embedded PDFs are still weak for SEO and awkward on phones. A common middle path is both — an HTML summary of the key content on the page, with the full PDF linked below it.
Which method fits your case?
| Situation | Best method | Why |
|---|---|---|
| Content going on your site long-term | Extract + rebuild | Clean, responsive, editable, indexable |
| One-off, internal, nobody edits it | pdftohtml / Calibre | Fast and free; messiness doesn't matter |
| Batch of 200 PDFs to get online | pdftohtml scripted | Only automatable option at that scale |
| Fixed layout is the point (forms, certificates) | Embed or link the PDF | Conversion can only make these worse |
| Scanned PDF (images of pages) | OCR first, then rebuild | There's no text layer to extract until OCR runs |
| SEO is the goal | Extract + rebuild | Search engines rank structured HTML, not positioned divs |
Honest expectations: what survives and what doesn't
Whatever method you choose, calibrate expectations:
- Plain text: survives everywhere. This is the easy 90%.
- Fonts: usually replaced with web-safe or system fonts unless you manually add webfonts.
- Multi-column layouts: the classic failure. Converters either flatten columns into one confusing stream or freeze them with absolute positioning.
- Tables: hit-and-miss. Simple grids often convert; merged cells and nested tables rarely do.
- Headers, footers, page numbers: come along as junk text repeated every "page" — you'll be deleting them by hand.
- Interactive form fields: don't survive any converter. Rebuild them as an HTML form.
If you're going the other direction — you have a web page and want a PDF of it — that's a much better-solved problem; see how to convert HTML to PDF.
Frequently asked questions
Is there a truly free PDF to HTML converter online?
Free automatic converters exist, but the output is generated markup — positioned divs and inline styles. The genuinely free route to clean HTML is extracting text and images with pdfty (free up to 20 MB, 10 operations/day, no sign-up) and writing the markup yourself.
My PDF is a scan — extraction returns nothing. Why?
A scanned PDF is pictures of pages, with no text layer inside. Run OCR on it first; after that, text extraction works normally. Without OCR, no converter on earth can find text that isn't there.
Will hyperlinks in the PDF survive conversion?
pdftohtml preserves most link annotations as a tags. In the
extract-and-rebuild workflow, plain text extraction drops them — copy
important URLs from the PDF manually as you mark up.
How do I get the images out at full quality?
Use Extract Images — it pulls the embedded originals rather than screenshotting pages, so you get the resolution that was actually stored in the PDF. Then compress them for the web before publishing.
Does Google index PDFs at all?
Yes — Google indexes and ranks PDFs. But HTML pages with the same content consistently outperform them: better mobile experience, faster loads, real heading structure, and internal links all feed ranking. If the content matters for search, publish it as HTML.
Can Word do PDF to HTML?
Sort of. Word opens many PDFs (File → Open), reflows them into an editable document, and can Save As → Web Page. The HTML it emits is notoriously bloated — usable as a starting point, not as production markup.
What about PDF to Markdown instead of HTML?
Often smarter. Extract the text, add Markdown heading and list syntax (faster to type than tags), and let your site generator produce the HTML. Same workflow as Method 1 with less typing.
Is an embedded PDF bad for accessibility?
Usually, yes. Screen reader support inside browser PDF viewers is inconsistent, and untagged PDFs read poorly everywhere. If accessibility matters, provide the content as real HTML — even a simplified version alongside the embed.
The pdfty team builds privacy-first online PDF tools — compress, convert, OCR, sign and protect. Files are deleted within 1 hour. About us →


