Convert HTML to PDF or PNG
Convert HTML to PDF or PNG with Puppeteer, Playwright, wkhtmltopdf, or one ScreenshotNeo request. Includes code, options, troubleshooting, and layout tips.
To convert HTML to PDF or PNG, render it in a browser or HTML rendering tool, wait for the content to be ready, then export the result. Use Puppeteer or Playwright when you need JavaScript, browser interactions, modern CSS, and precise control. Use wkhtmltopdf or wkhtmltoimage for a command-line workflow. If you only need a hosted screenshot or PDF endpoint, ScreenshotNeo can do it with one GET request.
Choose the right conversion method
| Need | Start with | Why |
|---|---|---|
| PDF from a JavaScript application | Puppeteer | Chrome or Firefox automation with page.pdf() and screenshots. |
| PDF using the Playwright Page API | Playwright | Browser-controlled rendering with print or screen media options. |
| Command-line PDF conversion | wkhtmltopdf | HTML to PDF from a shell, using Qt WebKit. |
| Command-line image conversion | wkhtmltoimage |
HTML to an image file from a shell. |
| Hosted screenshots or PDFs | ScreenshotNeo | One API request, with page cleanup, PDF options, and usage controls. |
These tools have different renderers and deployment requirements. The sources do not establish a universal winner for speed, quality, privacy, accessibility, or compatibility, so choose based on your page features and output requirements.
Convert HTML to PDF with Puppeteer
Puppeteer is a JavaScript library for automating Chrome and Firefox. Its PDF API uses print CSS media by default. The basic sequence is: launch a browser, create a page, load HTML or a URL, wait for required content, and call page.pdf().
Install Puppeteer
mkdir html-export
cd html-export
npm init -y
npm install puppeteer
Complete HTML-to-PDF script
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({
headless: true
});
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.goto('file:///absolute/path/to/input.html', {
waitUntil: 'networkidle0'
});
await page.emulateMediaType('print');
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'output.pdf',
format: 'A4',
printBackground: true,
margin: {
top: '16mm',
right: '16mm',
bottom: '16mm',
left: '16mm'
},
preferCSSPageSize: true,
displayHeaderFooter: false
});
} finally {
await browser.close();
}
})();
Replace the file URL with an absolute path, or use an HTTP(S) URL. Puppeteer waits for fonts during PDF generation by default, but that does not guarantee that every image, script, or third-party request has finished. Add an explicit readiness signal for dynamic pages.
Render an HTML string instead of a URL
const fs = require('node:fs/promises');
const puppeteer = require('puppeteer');
(async () => {
const html = await fs.readFile('input.html', 'utf8');
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.setContent(html, { waitUntil: 'networkidle0' });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'output.pdf',
format: 'Letter',
printBackground: true
});
} finally {
await browser.close();
}
})();
Convert HTML to PNG with Puppeteer
Use page.screenshot() when the output should be an image of the rendered page. A screenshot can capture the viewport, the full page, or a selected element.
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 768, deviceScaleFactor: 2 });
await page.goto('https://example.com', { waitUntil: 'networkidle0' });
await page.evaluate(() => document.fonts.ready);
await page.screenshot({
path: 'page.png',
type: 'png',
fullPage: true
});
} finally {
await browser.close();
}
})();
Set fullPage: true for the entire document. Omit it for the current viewport. To capture one element, locate it and pass its bounding box to clip:
const element = await page.$('.invoice');
if (!element) throw new Error('Missing .invoice element');
const box = await element.boundingBox();
if (!box) throw new Error('Element is not visible');
await page.screenshot({ path: 'invoice.png', clip: box });
PDF layout, media, and page options
Print CSS versus screen CSS
Puppeteer and Playwright generate PDFs with the print CSS media type by default. If your page is designed for the screen, switch media before exporting:
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-styled.pdf', printBackground: true });
Print rendering can adjust colors. When exact colors matter, add this rule to the page stylesheet:
@media print {
* {
-webkit-print-color-adjust: exact;
print-color-adjust: exact;
}
}
Use this carefully: forcing every color can increase ink usage on physical prints.
Paper size, dimensions, and margins
await page.pdf({
path: 'report.pdf',
format: 'A4',
landscape: false,
margin: {
top: '20mm',
right: '15mm',
bottom: '20mm',
left: '15mm'
},
printBackground: true,
preferCSSPageSize: true
});
Puppeteer supports named formats plus explicit width and height. Its documentation states that format takes priority over width and height. Playwright accepts units such as pixels, inches, centimeters, and millimeters.
await page.pdf({
path: 'custom.pdf',
width: '210mm',
height: '297mm',
margin: { top: '10mm', bottom: '10mm', left: '10mm', right: '10mm' }
});
Control page breaks with CSS
.avoid-break {
break-inside: avoid;
page-break-inside: avoid;
}
.page-break {
break-before: page;
page-break-before: always;
}
@page {
size: A4;
margin: 16mm;
}
Headers and footers
Puppeteer can add header and footer templates. Enable them with displayHeaderFooter; otherwise the templates are ignored.
await page.pdf({
path: 'numbered.pdf',
format: 'A4',
displayHeaderFooter: true,
headerTemplate: '<div></div>',
footerTemplate: `
<div style="font-size:9px;width:100%;text-align:center">
Page <span class="pageNumber"></span> of <span class="totalPages"></span>
</div>`,
margin: { top: '20mm', bottom: '20mm' }
});
Convert HTML to PDF with Playwright
Playwright exposes the same core page workflow. Install it with:
npm init -y
npm install -D playwright
npx playwright install chromium
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto('https://example.com', { waitUntil: 'networkidle' });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'playwright-output.pdf',
format: 'A4',
printBackground: true,
margin: { top: '15mm', right: '15mm', bottom: '15mm', left: '15mm' }
});
} finally {
await browser.close();
}
})();
To use screen styles with Playwright, call await page.emulateMedia({ media: 'screen' }) before page.pdf(). Refer to the Playwright Page.pdf documentation for the current option list.
Convert HTML with wkhtmltopdf and wkhtmltoimage
The wkhtmltopdf project documents two open-source LGPLv3 command-line tools: wkhtmltopdf for PDF and wkhtmltoimage for images, using Qt WebKit. Verify current maintenance and compatibility before choosing it for a new production system.
wkhtmltopdf input.html output.pdf
wkhtmltoimage --format png input.html output.png
You can convert a URL directly:
wkhtmltopdf https://example.com report.pdf
wkhtmltoimage --width 1440 https://example.com page.png
Use a browser API instead when your page depends on modern browser behavior, complex JavaScript interactions, or explicit readiness checks.
Reliable rendering for dynamic HTML
- Wait for navigation: use
networkidle0in Puppeteer ornetworkidlein Playwright when appropriate. - Wait for fonts: call
document.fonts.readybefore export. - Wait for an application signal: have your page set
window.renderReady = true, then wait for it.
await page.waitForFunction(() => window.renderReady === true, null, { timeout: 30000 });
For images, wait for the page’s image elements:
await page.waitForFunction(() => {
return [...document.images].every(img => img.complete && img.naturalWidth > 0);
}, null, { timeout: 30000 });
Use deterministic data, freeze animations, and provide stable fonts when output is used for tests, invoices, or other repeatable artifacts.
Or skip the browser setup
ScreenshotNeo converts a URL to a PNG, JPEG, WebP, or PDF through one request. Its cleanup step accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. See the ScreenshotNeo documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', buffer);
Set the output format and PDF settings through the documented query options. ScreenshotNeo also supports full-page capture with lazy images loaded, CSS selector capture, dark mode, device presets, custom viewports, retina scale, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and usage reporting. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF uses the wrong colors | Print media adjusts colors. | Use emulateMediaType('screen') when appropriate and add print-color-adjust: exact for required colors. |
| Backgrounds are missing | Background printing is disabled. | Set printBackground: true in Puppeteer or Playwright. |
| Content is cut off | Viewport capture or unsuitable page size. | Use fullPage: true for screenshots, set PDF margins and size, and add page-break CSS. |
| Fonts are wrong or fall back | Fonts have not loaded or are inaccessible. | Check font URLs, wait for document.fonts.ready, and ensure the runtime can reach the font host. |
| Charts or images are blank | Asynchronous rendering has not completed. | Wait for a selector, an application-ready flag, or complete image elements before export. |
| Navigation times out | The page or a dependency is slow or blocked. | Inspect the URL, network access, redirects, and resource requests; increase the timeout only after fixing the dependency. |
| Element screenshot fails | The selector is missing, hidden, or has no box. | Verify the selector, wait for it, scroll it into view, and check boundingBox(). |
| wkhtml output differs from the browser | Qt WebKit does not match a current Chromium renderer. | Use Puppeteer, Playwright, or ScreenshotNeo when modern CSS and JavaScript compatibility matter. |
Performance, reliability, and cost considerations
- Reuse one browser process for multiple pages instead of launching a new browser for every document.
- Limit concurrency so CPU, memory, and file descriptors remain available for other work.
- Use a readiness condition instead of an unnecessarily long fixed delay.
- Cache stable output and include the input version, CSS version, and renderer version in your cache key.
- Keep external resources predictable; third-party fonts, analytics, ads, and embeds can change output or prevent readiness.
- For PDF output, verify page breaks, font embedding, links, and image resolution with representative documents.
- For PNG output, choose viewport dimensions and device scale deliberately. A larger scale increases pixel dimensions and memory use.
- With ScreenshotNeo, cache TTL is configurable, bulk capture supports up to 100 URLs per call, and failed loads, blank pages, bot checks, timeouts, and cache hits are not billed.
Conversion checklist
- Choose PDF for paginated, printable documents; choose PNG for a visual snapshot.
- Decide whether the output needs print CSS or screen CSS.
- Set paper size, margins, viewport, and device scale.
- Wait for fonts, images, charts, and application data.
- Enable backgrounds when the design depends on them.
- Add page-break rules for long documents.
- Test remote assets and authenticated pages in the deployment environment.
- Record failures and inspect response headers or exit codes.
FAQ
Can HTML be converted to both PDF and PNG?
Yes. Use page.pdf() for PDF and page.screenshot() for PNG in Puppeteer or Playwright. The two outputs have different layout models: PDFs paginate, while screenshots represent pixels in a viewport or full document.
Why does my PDF look different from the webpage?
PDF generation uses print media by default, and print color handling can alter appearance. Emulate screen media when that is the intended design, then verify page size and margins.
Should I use wkhtmltopdf for a new project?
It can be useful for a simple command-line workflow, but its Qt WebKit renderer may not match modern browser behavior. Check current project status and compatibility before committing to it.
How do I convert a private or authenticated page?
In Puppeteer or Playwright, establish the session before export and set cookies or headers on the page. ScreenshotNeo supports custom headers, cookies, user agents, and Authorization parameters.
How do I avoid paying for failed screenshots?
ScreenshotNeo does not bill bot checks or CAPTCHAs, blank pages, timeouts, failed loads, or cache hits. Inspect the X-Page-Verdict and X-Billed response headers for each request.


