ScreenshotNeo

BlogHow-to

How to replace wkhtmltopdf for full-page website screenshots

Replace wkhtmltopdf with Playwright or Puppeteer for full-page website images. Choose the right capture mode and handle dynamic pages, lazy loading, and common migration issues.

By the ScreenshotNeo team4 October 20268 min read

Short answer: for a full-page image of a rendered website, replace wkhtmltopdf with browser automation that explicitly supports full-page screenshots. Playwright and Puppeteer both provide this option. Use PDF generation only when the required output is a paginated document; a PDF is not interchangeable with a screenshot image.

wkhtmltopdf is designed to render HTML as PDF. If your current workflow uses it to produce website screenshots, first confirm whether the desired output is one tall image, a viewport-sized image, or a PDF. The right replacement depends on that output.

1. Choose the capture mode

Required output Recommended approach What to know
One image of the whole scrollable page Playwright or Puppeteer full-page screenshot Both expose an explicit full-page screenshot option. Validate rendering against your pages and runtime.
Image of the visible viewport Browser screenshot or Chrome Headless screenshot The screenshot covers the viewport; it does not necessarily include content below the fold.
Paginated document Chrome Headless PDF output or Puppeteer PDF PDF output follows print behavior and can differ from the screen layout.

Playwright describes full-page screenshots as images of the entire scrollable page, as if the page fit on a very tall screen. Its API is a direct match for a single tall image. Puppeteer also exposes a fullPage screenshot option. Chrome Headless documents command-line screenshots and PDF output, but its cited screenshot example is viewport-sized; use an automation API when you need an explicit full-page image option. Playwright screenshot documentation, Puppeteer ScreenshotOptions, Chrome Headless documentation.

2. Replace wkhtmltopdf with Playwright

Install Playwright and its Chromium browser in your project, then capture the page with fullPage: true. The following is a runnable Node.js example for a public URL:

import { chromium } from 'playwright';

const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });

try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto(url, { waitUntil: 'networkidle', timeout: 60_000 });
  await page.screenshot({ path: 'screenshot.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with node capture.mjs https://example.com. The example waits for network activity to settle before capturing. Some sites keep connections open or continually fetch data, so network idle may never arrive; in that case wait for a meaningful selector or a bounded delay instead. Playwright documents full-page capture and page screenshots in its screenshot guide.

Wait for the content you need

For a client-rendered page, navigation completion does not guarantee that the content you care about has appeared. Prefer a page-specific readiness signal:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.locator('main article').waitFor({ state: 'visible', timeout: 20_000 });
await page.screenshot({ path: 'screenshot.png', fullPage: true });

For pages with lazy-loaded images, scroll through the page before the final capture so image elements have a chance to load. This changes the page state, so validate the result on representative targets. For example:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.locator('main').waitFor({ state: 'visible', timeout: 20_000 });
await page.evaluate(async () => {
  const step = Math.max(300, window.innerHeight * 0.8);
  for (let y = 0; y < document.body.scrollHeight; y += step) {
    window.scrollTo(0, y);
    await new Promise(resolve => setTimeout(resolve, 100));
  }
  window.scrollTo(0, 0);
});
await page.screenshot({ path: 'screenshot.png', fullPage: true });

This simple scroll loop is not a guarantee for every site: some pages append content as you scroll, and their height can keep increasing. For those targets, use a site-specific readiness condition or a bounded scroll strategy and check that the final page height stabilizes.

3. Replace wkhtmltopdf with Puppeteer

Puppeteer is another browser automation option with an explicit full-page screenshot setting. Install the package and run this Node.js example:

import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch({ headless: true });

try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1440, height: 900 });
  await page.goto(url, { waitUntil: 'networkidle2', timeout: 60_000 });
  await page.screenshot({ path: 'screenshot.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with node capture.mjs https://example.com. Puppeteer’s fullPage option captures the full page. See the official ScreenshotOptions reference.

4. When Chrome Headless CLI is enough

Chrome Headless is useful for straightforward viewport screenshots and PDF output. Its documented command-line interface includes --screenshot, --window-size, a timeout flag, and --print-to-pdf. The screenshot example documents a browser window size, not an explicit full-scrollable-page capture option, so choose Playwright or Puppeteer when the requirement is a single image of the entire scrollable page.

chrome --headless --no-sandbox --window-size=1440,900 --screenshot=screenshot.png https://example.com

To generate a PDF instead:

chrome --headless --no-sandbox --print-to-pdf=page.pdf https://example.com

Use the flags appropriate to your installed Chrome and environment; headless command-line behavior can depend on the runtime. Consult the official Chrome Headless guide for the documented options.

5. Keep PDF output separate from image capture

If you still need the artifact wkhtmltopdf originally produced, a PDF, use PDF generation and inspect the print layout. Puppeteer documents that Page.pdf() uses print CSS by default. That can alter colors, visibility, dimensions, and pagination compared with a screen screenshot. Puppeteer Page.pdf().

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle2' });
  await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
  await browser.close();
}

Use the screenshot APIs for raster images and PDF APIs for paginated output. If matching the on-screen appearance in a PDF matters, review print styles and compare the generated file with the page shown in a browser.

6. Migration checklist

  1. Confirm the artifact. Record whether callers expect a viewport image, a full-page image, or a PDF, along with image format, dimensions, and naming conventions.
  2. Select the capture API. Start with Playwright or Puppeteer for a full-page image; use PDF output for a document.
  3. Define readiness. Decide which selector, application signal, or bounded wait means the page is ready to capture.
  4. Handle dynamic content. Test lazy images, client-rendered content, fonts, animations, sticky headers, and pages that add content during scrolling.
  5. Test access requirements. Check authenticated pages, redirects, cookies, and any headers the target requires.
  6. Check output limits. Verify actual page height, file size, memory use, and whether downstream tools accept a tall image.
  7. Validate in the production runtime. Browser version, installed fonts, operating system, and launch configuration can affect rendering.
  8. Keep failure handling explicit. Set navigation and selector timeouts, close the browser in a finally block, and log the target URL and failure stage.

7. Troubleshooting

Symptom Likely cause Fix
The image ends at the viewport bottom The capture used a viewport screenshot or CLI mode without full-page support. Use Playwright or Puppeteer and set fullPage: true.
Content is missing or still loading Navigation finished before the application rendered the target content, or lazy loading has not been triggered. Wait for a meaningful selector or app readiness condition; scroll to load lazy content, then capture.
Navigation times out on an otherwise usable page The site keeps network requests active or never reaches the selected load condition. Use a less restrictive navigation condition such as DOM content loaded, then wait for the specific content needed.
Images or fonts look different in deployment The runtime may have different fonts, browser binaries, or network access than development. Use a consistent browser runtime, ensure required fonts and assets are available, and validate in the deployment environment.
A PDF looks unlike the screenshot PDF generation uses print media behavior and pagination. Review print CSS and PDF settings; use screenshot capture if the requested output is an image of the screen layout.
Fixed headers repeat or overlap in a tall image Full-page capture interacts with fixed-position elements differently from a normal viewport. Inspect the output on the target page. Adjust the page’s capture styling or use a page-specific capture approach where needed.
The screenshot is huge or capture runs out of memory Very tall pages create large raster images and consume browser memory. Measure page height, limit capture to the needed content when possible, or capture sections separately if one single image is not required.
The site shows a CAPTCHA or access-denied page The target is gating automated traffic or requires a valid session. Use an authorized session and follow the site’s access rules. Do not assume a screenshot API can bypass bot checks.

8. Performance, reliability, and cost

Browser automation gives direct control over the browser session, but each capture requires browser work and the target page’s loading behavior can dominate the time. There is no universal speed or cost comparison in the cited documentation. Measure your own representative pages, runtime, concurrency, and output sizes before estimating capacity.

  • Reuse carefully: Reusing a browser process can avoid repeated startup overhead, while using a fresh page or context per task helps isolate cookies and page state. Set concurrency to fit available memory.
  • Bound waits: Give navigation, selector waits, and any scroll-loading loop a deadline. A page with persistent network traffic should not hold a job indefinitely.
  • Make retries selective: Retry transient navigation failures with a limit. A deterministic selector timeout or persistent access denial usually needs diagnosis rather than repeated attempts.
  • Track the output: Record capture duration, page height, file size, and failure category. These measurements help identify slow targets and resource-heavy pages.
  • Consider operational cost: Self-hosting means managing browser installation, runtime compatibility, memory, and concurrency. A hosted API trades that setup for a per-plan service cost; compare against your workload and required controls.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. For an overview of supported request options, see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status} ${res.statusText}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed; response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Create a free account and get 1,000 screenshots a month with no card.

10. Frequently asked questions

Can wkhtmltopdf produce a full-page screenshot image?

wkhtmltopdf’s intended output is PDF. For a full-page raster image of a rendered site, use an API with an explicit full-page screenshot option, such as Playwright or Puppeteer.

Is a full-page screenshot the same as stitching viewport screenshots?

The APIs expose a full-page capture mode; the resulting image represents the full scrollable page. For pages with fixed elements or content that loads as you scroll, inspect the output for artifacts.

Should I choose Playwright or Puppeteer?

Both document full-page screenshot support. Choose based on the browser automation library that best fits your existing project and validate it with your actual target pages and deployment environment.

Can I use this migration for PDFs too?

Yes, but choose a PDF API or command-line PDF mode and review print styles and pagination. PDF output follows print behavior and may not match a screen screenshot.