ScreenshotNeo

BlogHTML to image & PDF

How to Stop HTML-to-PDF Conversion When Page Loading Fails

Stop stuck HTML-to-PDF jobs by separating navigation timeouts, readiness checks, cancellation, and load-error handling in Playwright, Puppeteer, and wkhtmltopdf.

By the ScreenshotNeo team29 September 20268 min read

How to Stop HTML-to-PDF Conversion When Page Loading Fails

When HTML-to-PDF conversion hangs or fails during page loading, stop treating PDF generation as one indivisible operation. First control navigation: choose what “ready” means, set a finite timeout, and cancel the navigation when it exceeds that limit. Then call the PDF writer only after navigation succeeds. The exact setting depends on whether you use Playwright, Puppeteer, or wkhtmltopdf.

Diagnose the converter before changing settings

Ask three questions before editing code:

  1. Which converter is running? Playwright and Puppeteer expose browser navigation controls. wkhtmltopdf has command-line load-error and script options.
  2. What failed? A DNS error, refused connection, redirect loop, blocked resource, JavaScript exception, or a page that keeps making requests require different fixes.
  3. What does ready mean for this page? The document may be usable at commit, after the DOM is parsed, after the load event, or only after an application-specific selector appears.

Capture the exact exception, URL, elapsed time, redirect chain, and converter version in your job logs. The reviewed documentation does not establish one universal “page loading failed” error, so version-specific advice should begin with those facts.

Playwright: use a finite timeout and intentional readiness condition

Playwright navigation supports commit, domcontentloaded, load, and networkidle. The documentation discourages using networkidle as a general readiness test because modern pages can continue making requests indefinitely; prefer web assertions or a page-specific signal instead. Navigation also accepts an AbortSignal. If the signal is aborted, navigation throws, allowing your job runner to stop work and report the failure. The documented default navigation timeout is zero unless you configure one, so set a limit for every unattended conversion. See the Playwright page.goto API.

Separate navigation readiness and cancellation from the PDF-writing step.
Separate navigation readiness and cancellation from the PDF-writing step.

Runnable Node.js example

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();
page.setDefaultNavigationTimeout(30_000);

const controller = new AbortController();
const stopTimer = setTimeout(() => controller.abort(), 35_000);

try {
  await page.goto('https://example.com/report', {
    waitUntil: 'domcontentloaded',
    timeout: 30_000,
    signal: controller.signal
  });

  // Replace this with a domain-specific readiness check when possible.
  await page.locator('main').waitFor({ state: 'visible', timeout: 5_000 });
  await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });
} catch (error) {
  console.error('HTML-to-PDF conversion stopped:', error.message);
} finally {
  clearTimeout(stopTimer);
  await browser.close();
}

Use commit when you only need the initial response, domcontentloaded when the DOM is enough, and load when the page’s load event matters. A selector, assertion, or application-ready flag is usually more reliable than waiting for all network activity. Keep the outer abort deadline slightly longer than the navigation timeout so cleanup runs predictably.

Handling a page that never becomes ready

  • Use a short navigation timeout and a separate timeout for the selector or assertion that proves the content exists.
  • Log the URL and the last navigation error. Do not continue to page.pdf() after navigation failed.
  • For pages with long polling, analytics, or streaming requests, avoid networkidle. Wait for the content your PDF actually needs.
  • If a page is optional in a batch, catch the thrown error, mark that item failed, and continue with the remaining URLs.

Puppeteer: separate goto from pdf

Puppeteer’s official example navigates with a waitUntil condition and then calls page.pdf(). Its PDF guide states: “By default, the Page.pdf() waits for fonts to be loaded.” That means a successful navigation can still spend time preparing fonts before the file is written. See the Puppeteer PDF generation guide and page.goto API.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
const page = await browser.newPage();
page.setDefaultNavigationTimeout(30_000);

try {
  await page.goto('https://example.com/report', {
    waitUntil: 'domcontentloaded',
    timeout: 30_000
  });
  await page.waitForSelector('main', { visible: true, timeout: 5_000 });
  await page.pdf({
    path: 'report.pdf',
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true
  });
} catch (error) {
  console.error('Conversion failed:', error.message);
} finally {
  await browser.close();
}

Do not confuse a PDF timeout with a navigation timeout. If goto fails, fix or classify the page-loading problem first. If navigation succeeds but PDF creation is slow, investigate fonts, print styles, very large images, and page size separately.

wkhtmltopdf: choose abort, ignore, or skip

wkhtmltopdf exposes explicit page-load behavior through --load-error-handling. It accepts abort, ignore, or skip, and abort is the documented default. The separate --load-media-error-handling option controls failed media resources and defaults to ignore. These controls are different: a missing stylesheet or image is not necessarily the same as a page-level navigation failure. Read the wkhtmltopdf usage documentation.

# Fail the conversion when the page itself cannot load.
wkhtmltopdf --load-error-handling abort https://example.com/report report.pdf

# Continue when a page-load error is acceptable for your workflow.
wkhtmltopdf --load-error-handling ignore https://example.com/report report.pdf

# Skip a failing page in a multi-page input.
wkhtmltopdf --load-error-handling skip input.html report.pdf

# Choose a separate policy for images, stylesheets, and other media.
wkhtmltopdf --load-media-error-handling ignore https://example.com/report report.pdf

Use abort when an incomplete PDF is worse than a failed job. Use ignore only when the remaining document is useful without the failed resource. Use skip when processing multiple inputs and a missing page should not block the whole output. wkhtmltopdf also documents a JavaScript delay and an option to stop slow scripts. Those settings affect script execution after loading; they do not replace diagnosing DNS, HTTP, or navigation failures.

Readiness strategies that work in production

Page behavior Recommended readiness Why
Static HTML domcontentloaded Content is available without waiting for every asset.
CSS and images required load plus a bounded timeout The load event covers the page’s normal resources.
Single-page app DOM event plus a selector or assertion Application content may appear after the initial document.
Long polling or analytics Selector, assertion, or app-ready flag Network idle may never occur.
Untrusted or unreliable URL Finite timeout plus cancellation The job cannot wait indefinitely.

Make readiness page-specific. For an invoice, wait for the invoice number and totals. For a dashboard, wait for the chart container and a loaded state. Keep the signal deterministic and bounded; a selector that never appears must produce a clear failure rather than an infinite wait.

Common errors and fixes

Cause: the selected readiness event never occurred, the host is slow, or a request is hanging. Fix: set a finite timeout, switch from networkidle to a selector or assertion, inspect the URL outside the converter, and classify the page as failed when the deadline expires.

ERR_NAME_NOT_RESOLVED, connection refused, or TLS errors

Cause: DNS, firewall, proxy, certificate, or private-network access. Fix: verify DNS and HTTPS from the same runtime, configure the required proxy or trust store, and do not hide the error by increasing the timeout.

Redirect loop or unexpected login page

Cause: missing cookies, authentication headers, geolocation, or a redirect policy. Fix: inspect the final URL and response chain, provide the required session state, and set an explicit maximum redirect policy in the surrounding HTTP or browser infrastructure.

Blank or partially rendered PDF

Cause: PDF generation started before application content was ready, or a critical stylesheet/font/image failed. Fix: wait for a content selector, check console and request failures, use print CSS deliberately, and decide whether media failures should abort.

Script never finishes

Cause: an application loop, third-party widget, or long-running script. Fix: remove or block unnecessary resources, wait for a domain-specific signal, and use wkhtmltopdf’s script controls only for script behavior rather than as a network fix.

PDF creation hangs after navigation succeeds

Cause: fonts, enormous images, complex layout, or a huge document. Fix: measure navigation and PDF-writing phases separately, reduce asset size, limit page ranges, and enforce an outer job deadline.

Reliability, performance, and cost design

  • Use layered deadlines. Set separate limits for DNS/HTTP navigation, readiness checks, PDF writing, and the complete job. A single large timeout makes failures hard to classify.
  • Retry selectively. A transient connection reset may merit one retry with backoff. A deterministic 404, authentication failure, or missing selector should fail fast.
  • Bound concurrency. Multiple browser pages can exhaust CPU, memory, file descriptors, or bandwidth. Use a queue and a fixed worker count.
  • Cache stable inputs. If the source and print state are unchanged, reuse the PDF or intermediate HTML. Add cache keys that include URL, authentication context, viewport, and print options.
  • Record phase timings. Store navigation, readiness, PDF generation, bytes written, and final verdict. This shows whether the bottleneck is the source page or the converter.
  • Choose failure semantics explicitly. A partial PDF may be cheaper operationally but dangerous for invoices, legal records, and reports. Decide whether each input is all-or-nothing.

Or skip the browser setup

ScreenshotNeo provides a website capture API and MCP server. For PDF capture, use its API options for paper size, margins, landscape mode, and page ranges. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Clean page capture removes consent banners and distracting overlays before rendering.
Clean page capture removes consent banners and distracting overlays before rendering.

See the ScreenshotNeo API documentation for the complete option list and authentication details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo has 1,000 shots per month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Should I always wait for network idle?

No. Playwright specifically discourages it as a universal readiness check. Use a selector, assertion, or application-ready signal when the page keeps background requests open.

What is the safest wkhtmltopdf setting?

There is no universal safest choice. abort protects against silently incomplete pages; ignore or skip can be appropriate when missing content is acceptable and documented.

Does increasing a timeout fix a failed DNS lookup?

No. A longer wait does not repair DNS, TLS, authentication, or firewall problems. Verify connectivity from the converter’s runtime.

Why can PDF generation still be slow after navigation?

Puppeteer waits for fonts by default, and layout, images, print styles, and document size can add work after navigation. Measure the PDF phase separately.

How do I stop a batch when one URL fails?

For all-or-nothing output, propagate the first classified failure and cancel remaining work. For independent pages, catch each error, record it, and continue with bounded concurrency.