How HTML and PDF File Sizes Compare
HTML is not always smaller than PDF. Learn how to compare equivalent content, transfer sizes, assets, compression, and pagination accurately.
There is no universal size winner. An HTML document can be smaller than a PDF, but the result depends on what you count, which assets are included, compression, image quality, fonts, and whether the PDF paginates the content. Compare equivalent information and report the measurement method before drawing a conclusion.
What exactly are you comparing?
“HTML size” can mean three different measurements:
| Measurement | What it includes | Why it matters |
|---|---|---|
| HTML document response | The bytes returned for the document request | Useful for measuring the initial text document, but excludes separately fetched resources |
| Whole rendered page | HTML plus CSS, JavaScript, fonts, images, video, embedded content, and other requests | Closest to the bytes a browser needs to render the page |
| Transfer size | Compressed bytes sent over HTTP | Measures network cost; it can be much smaller than the decompressed resource size |
MDN describes HTML as mostly text and therefore generally small, while images, video, embedded content, and resource loading order can dominate page performance. See MDN’s HTML performance guide and Chrome’s resource summary documentation.
A PDF is one file, but its size depends on embedded images, fonts, metadata, page count, compression settings, and the PDF generator. If the HTML references external images while the PDF embeds those images, you are comparing different asset scopes.
A real example, with limits
The UK Government Communication Service reported one matched example in which an HTML page was 1.4 MB and the PDF was nearly 2 MB, making the PDF 42% larger. The article estimated 0.395g CO2e per HTML view and 0.561g per PDF view. These are estimates for that example, not a general HTML-to-PDF ratio or universal emissions factors. Read the Government Communication Service article for its methods and caveats.
PDF conversion can also change the representation. Adobe explains that a continuous web page may be divided into multiple standard-size PDF pages during conversion. A page break, repeated header, or embedded font can add bytes even when the information is unchanged. See Adobe’s web-page-to-PDF documentation.
How to measure HTML and PDF fairly
1. Define the comparison boundary
- Use equivalent content and the same revision.
- Record whether you are measuring only the HTML document or the complete rendered page.
- Record transfer size and resource size separately.
- State whether images, fonts, stylesheets, scripts, video, and embeds are external or included.
- Record PDF page count, image quality, font embedding, and whether the PDF is linearized or otherwise optimized.
2. Measure the HTML document from the command line
This downloads the compressed response and counts the bytes saved locally:
curl --compressed -L https://example.com/page -o page.html
wc -c < page.html
To inspect HTTP headers and compression:
curl -sSIL --compressed -L https://example.com/page
Look for Content-Encoding, Content-Length, redirects, and the final URL. A server may omit Content-Length when it streams or uses chunked transfer.
3. Measure the PDF file
curl -L https://example.com/document.pdf -o document.pdf
wc -c < document.pdf
Use the actual file delivered to users. Do not compare a source document or an unoptimized intermediate PDF with a compressed production HTML response.
4. Measure the complete rendered page in browser tools
- Open the page in a Chromium-based browser.
- Open Developer Tools and select the Network panel.
- Enable the option that disables cache, then reload.
- Inspect the document request for the HTML response size.
- Use the summary totals to record transferred and resource sizes for documents, stylesheets, scripts, fonts, images, media, and other requests.
- Record whether the page loads lazy images only after scrolling or interaction.
Chrome’s resource summary explains how Lighthouse aggregates request counts and transfer sizes. Google’s older measurement guidance also describes using the Network panel and cURL; it remains useful for measurement even though its historical crawler limit is outdated.
5. Automate a repeatable comparison with Python
from pathlib import Path
import requests
urls = {
"html": "https://example.com/page",
"pdf": "https://example.com/document.pdf",
}
for name, url in urls.items():
response = requests.get(url, allow_redirects=True, timeout=60)
response.raise_for_status()
data = response.content
Path(f"{name}.bin").write_bytes(data)
print({
"name": name,
"final_url": response.url,
"bytes_downloaded": len(data),
"content_type": response.headers.get("content-type"),
"content_encoding": response.headers.get("content-encoding"),
"content_length_header": response.headers.get("content-length"),
})
This reports bytes received by the client. For HTML, it measures the document response only; it does not crawl the page’s subresources.
6. Automate a PDF and HTML download with Node.js
import { writeFile } from 'node:fs/promises';
const targets = [
['html', 'https://example.com/page'],
['pdf', 'https://example.com/document.pdf'],
];
for (const [name, url] of targets) {
const response = await fetch(url, { redirect: 'follow' });
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
const buffer = Buffer.from(await response.arrayBuffer());
await writeFile(`${name}.bin`, buffer);
console.log({
name,
finalUrl: response.url,
bytesDownloaded: buffer.length,
contentType: response.headers.get('content-type'),
contentEncoding: response.headers.get('content-encoding'),
contentLengthHeader: response.headers.get('content-length'),
});
}
Compression changes the answer
HTML, CSS, JavaScript, and other text resources are commonly compressed with gzip or Brotli. The transfer size is the compressed network cost; the resource size is the decompressed amount used by the browser. web.dev recommends compressing text-based resources and gives general guidance that Brotli can improve text compression by about 15% to 20% over gzip. That is guidance, not a guaranteed reduction for every file.
Binary assets such as JPEG, WebP, PNG, and many PDFs may already be compressed. Recompressing them can have little effect or reduce quality. Always label which size you report.
Why a PDF can be larger or smaller
- Images: A PDF may embed full-resolution images while HTML serves responsive, modern formats. The reverse can happen when a page loads many large assets.
- Fonts: Embedded font subsets add PDF bytes; HTML may fetch the same fonts once and reuse them across pages.
- Repeated content: Headers, footers, and page backgrounds can be stored repeatedly or shared internally depending on the PDF generator.
- Pagination: Page breaks, print margins, and repeated headings change the file representation.
- Scripts and embeds: A minimal HTML document can trigger substantial JavaScript, analytics, video, or third-party requests.
- Lazy loading: The initial HTML view may not download images until they enter the viewport, while a full-page capture or PDF export may load every image.
- Compression settings: Different PDF producers use different image quality, object streams, and font compression settings.
Google crawler limits are not format recommendations
Google Search Central’s March 2026 guidance says Googlebot stops an HTML fetch at 2 MB, including HTTP request headers, and gives PDF files a 64 MB limit. These are Googlebot processing ceilings that can change; they are not typical file sizes or publishing targets. The older 2022 guidance discussed a 15 MB limit, so use the newer dated guidance when discussing current crawler behavior. See the March 2026 Google Search Central article and its older measurement explanation.
Choosing HTML or PDF after measuring
| Question | HTML often fits when… | PDF often fits when… |
|---|---|---|
| Bytes | You can serve optimized, shared external assets and compressed text | The document is mostly text or already-optimized graphics in one downloadable artifact |
| Reading format | Users need responsive reflow and links | Users need stable pagination and print fidelity |
| Offline use | A service worker or saved page is acceptable | A single downloadable file is required |
| Updates | Content changes frequently and should update centrally | A versioned, immutable document is useful |
Measure the bytes and then consider pagination, print behavior, accessibility requirements, update workflow, and offline needs. File size alone does not decide the format.
Or skip the browser setup
ScreenshotNeo captures a URL as PNG, JPEG, WebP, or PDF through one request. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
See the ScreenshotNeo API documentation for all options, including full-page capture with lazy images loaded, CSS-selector element capture, PDF paper size and margins, custom CSS and JavaScript, waits, request blocking, headers and cookies, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/page"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/page' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Troubleshooting
The HTML number is tiny, but the page is slow
You measured only the document response. Inspect the Network panel totals and identify images, scripts, fonts, video, and embeds.
Content-Length does not match the downloaded bytes
Compression, chunked transfer, redirects, or an intermediary may explain the difference. Count the bytes saved by the client and record the encoding headers.
The same content produces different PDF sizes
Check PDF generator settings, image resolution, font embedding, metadata, page count, and whether the output is optimized or linearized.
A full-page export is much larger than the first browser view
Full-page capture may load lazy images and below-the-fold resources. Compare the same viewport and resource scope before judging the formats.
The request returns an HTML error page instead of a PDF
Check the final URL, status code, authentication, redirects, and Content-Type. Save response headers and inspect the first bytes of the file.
ScreenshotNeo reports an unexpected result
Inspect the X-Page-Verdict and X-Billed response headers. A bot check, blank page, timeout, failed load, or cache hit has its own verdict and is not billed when it is not a clean shot.
Performance, reliability, and cost notes
- Use Brotli or gzip for text and report transfer size separately from resource size.
- Cache immutable PDFs and static assets; set cache headers deliberately for changing HTML.
- Repeat measurements across cold and warm cache conditions.
- Use a fixed viewport, user agent, locale, and authentication state when comparing rendered output.
- For automated screenshots, wait for a selector, a delay, or network idle so dynamic content is settled.
- With ScreenshotNeo, choose a cache TTL, block unnecessary requests, and use asynchronous jobs or bulk capture for larger batches. Clean shots are billed; bot checks, blank pages, timeouts, failed loads, and cache hits are not.
FAQ
Is an HTML page smaller than a PDF?
Sometimes. There is no dependable universal ratio because asset scope, compression, fonts, images, and pagination vary.
Should I compare the HTML source file with the PDF?
Only if you clearly state that you are comparing the document response with a complete PDF. For user experience, also measure the full rendered page.
Does Brotli guarantee a smaller result than gzip?
No. web.dev gives about 15% to 20% as general guidance for text, but the result depends on content and settings.
Does Google’s 2 MB HTML limit mean every page must be under 2 MB?
No. It is a dated Googlebot processing ceiling and can change; it is not a general publishing target.
Can one continuous web page and one PDF page count be compared directly?
Not always. PDF conversion can paginate continuous content, add repeated elements, and change the representation.
