How to Resolve Relative URLs in Puppeteer PDFs
Fix broken images, stylesheets, and links in Puppeteer PDFs with base URLs, absolute paths, navigation, validation, and troubleshooting.
Relative URLs in a Puppeteer PDF resolve against the document’s base URL. When you pass an HTML string to page.setContent(), add a suitable <base href="..."> element, rewrite references as absolute URLs, or navigate to the real page with page.goto() before calling page.pdf(). Puppeteer’s current Page.setContent() API does not provide a baseURL option.
This guide shows each fix, complete Node.js examples, asset-waiting patterns, PDF validation steps, common failure causes, and an API option when you do not want to maintain browser infrastructure.
How URL resolution works
A relative reference such as images/chart.png or details.html needs a document base URL. The browser resolves it according to standard URL parsing rules. A document injected with page.setContent() may not have the origin and path you assumed, so references can point at the wrong location or fail to load.
The path matters. With https://example.com/reports/, details.html becomes https://example.com/reports/details.html. With https://example.com/reports, the last path segment is treated as a file-like component and the result can instead be https://example.com/details.html. Use a trailing slash when the base represents a directory.
Fix 1: add a base element to injected HTML
Put <base> early in the document head, before elements that use relative href or src values.
<!doctype html>
<html>
<head>
<base href="https://example.com/reports/">
<link rel="stylesheet" href="styles/print.css">
</head>
<body>
<a href="details.html">Details</a>
<img src="images/chart.png" alt="Monthly chart">
</body>
</html>
Here, the stylesheet resolves to https://example.com/reports/styles/print.css, the image to https://example.com/reports/images/chart.png, and the anchor to https://example.com/reports/details.html. If links and assets come from different roots, give those references explicit absolute URLs instead of forcing one base to serve both.
Complete Node.js example with setContent()
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
const html = `
<!doctype html>
<html>
<head>
<meta charset="utf-8">
<base href="https://example.com/reports/">
<link rel="stylesheet" href="styles/print.css">
<style>@page { margin: 18mm; }</style>
</head>
<body>
<h1>Monthly report</h1>
<img src="images/chart.png" width="800" alt="Chart">
<a href="details.html">View details</a>
</body>
</html>`;
await page.setContent(html, { waitUntil: 'networkidle0' });
// Wait for images and fonts used by this document.
await page.evaluate(async () => {
await Promise.all([
...Array.from(document.images, image => {
if (image.complete) return Promise.resolve();
return new Promise(resolve => {
image.addEventListener('load', resolve, { once: true });
image.addEventListener('error', resolve, { once: true });
});
}),
document.fonts ? document.fonts.ready : Promise.resolve()
]);
});
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
} finally {
await browser.close();
}
Puppeteer’s setContent() API reference documents the HTML-and-options interface; it does not expose a baseURL parameter.
Fix 2: rewrite references as absolute URLs
Absolute URLs make the intended destination explicit and are useful when a document contains resources from several origins. Resolve paths before injecting the markup with JavaScript’s standard URL constructor.
const reportBase = 'https://example.com/reports/';
const imageUrl = new URL('images/chart.png', reportBase).href;
const detailsUrl = new URL('details.html', reportBase).href;
const html = `
<h1>Monthly report</h1>
<img src="${imageUrl}" alt="Chart">
<a href="${detailsUrl}">Details</a>`;
For transformed templates, resolve every URL that can be relative, including srcset candidates, CSS url(...) values, Open Graph images, favicon links, and downloadable assets. Do not assume that changing an HTML src also rewrites URLs embedded inside a stylesheet.
Fix 3: navigate to the real source URL
If the document already exists at a URL, let Chromium load that URL and then print the page. Navigation supplies the page’s real origin and path, so relative references resolve naturally.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/reports/monthly', {
waitUntil: 'networkidle2'
});
await page.pdf({
path: 'monthly.pdf',
format: 'A4',
printBackground: true
});
} finally {
await browser.close();
}
The official Puppeteer PDF guide demonstrates the navigation-then-print workflow. Use this approach for a public or authenticated page that can be reached by the browser. Use a base element or absolute URLs when the source is only an HTML string.
Choosing the right fix
| Input | Recommended fix | Reason |
|---|---|---|
| Generated HTML string | <base href> |
One declaration establishes the intended origin and directory. |
| Template with multiple asset roots | Absolute URLs | Each reference can point to its own origin or path. |
| Existing page at a URL | page.goto() |
The browser receives the real document URL automatically. |
| Private page requiring headers or login | Navigate with authentication, then print | Relative resources inherit the authenticated page context. |
Wait for resources before printing
waitUntil: 'networkidle0' or 'networkidle2' only describes network activity observed during the operation. Fonts, images, client-side rendering, and third-party resources can still need an application-specific readiness check. Wait for a known selector, image completion, a font promise, or a page-defined flag before calling page.pdf().
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#report-ready');
await page.evaluate(() => document.fonts?.ready);
await page.pdf({ path: 'report.pdf', printBackground: true });
Keep the wait bounded in production. A page that never creates the selector or keeps a connection open should fail with a useful timeout rather than hold a browser indefinitely.
PDF options that affect the result
- Print CSS:
page.pdf()uses the print media type by default. Add print-specific rules with@media print. - Backgrounds: set
printBackground: truewhen colors and background images are part of the design. - Page size: use
formatsuch asA4orLetter, or define@pageand usepreferCSSPageSize: true. - Margins: configure
marginexplicitly when headers, footers, or edge-to-edge content matter. - Headers and footers: enable
displayHeaderFooterand provide templates when page numbers or dates are required. - Landscape: set
landscape: truefor wide tables or charts.
Validate links and assets in the generated PDF
- Inspect the final HTML in DevTools or log resolved values with
document.baseURI. - Check failed requests with
page.on('requestfailed', ...)and response status codes withpage.on('response', ...). - Open the generated PDF in the viewer your users rely on and click representative links.
- Verify images, fonts, page breaks, and print colors in the actual file.
console.log(await page.evaluate(() => ({
baseURI: document.baseURI,
links: [...document.links].map(link => link.href),
images: [...document.images].map(image => ({ src: image.src, complete: image.complete }))
})));
page.on('requestfailed', request => {
console.error('Request failed:', request.url(), request.failure()?.errorText);
});
Do not promise that every relative anchor becomes a clickable PDF annotation in every Puppeteer, Chromium, and viewer combination. The official API references document PDF printing but do not specify a universal hyperlink guarantee. Test the exact versions and viewers used by your workflow.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Images are blank | No base URL, wrong directory, blocked request, or print occurred before loading finished. | Add <base> or absolute URLs, inspect failed requests, and wait for image completion. |
| Stylesheet is missing | The relative href resolved from an unexpected path. |
Log document.baseURI, correct the trailing slash, or use an absolute stylesheet URL. |
| Links point to the wrong directory | A base such as /reports was used instead of /reports/. |
Use the directory form with a trailing slash or resolve each link explicitly. |
setContent() rejects a baseURL option |
The option is not part of Puppeteer’s current setContent() API. |
Put <base href> in the HTML, rewrite URLs, or navigate with goto(). |
| PDF has old or missing client-rendered data | Printing started before the application finished rendering. | Wait for a readiness selector, a known delay, or an application flag. |
| Fonts differ from the browser view | Fonts were not loaded or print CSS selects different rules. | Await document.fonts.ready, verify font responses, and review @media print. |
| Private assets return 401 or 403 | Requests made by the page lack authentication. | Authenticate the browser context and ensure subresource requests receive the required cookies or headers. |
| PDF generation times out | A page keeps connections open or a readiness condition never occurs. | Use a bounded timeout, a precise selector, and request logging to identify the stuck resource. |
Performance, reliability, and cost considerations
- Reuse browsers carefully: launching Chromium is expensive, but sharing a browser across jobs requires isolated pages and contexts so cookies and state do not leak.
- Reduce work before printing: block unnecessary analytics and media only when the document does not depend on them; otherwise you can create incomplete PDFs.
- Use deterministic readiness: a specific selector or application flag is generally easier to operate than an arbitrary long delay.
- Bound every external dependency: set navigation and PDF timeouts, log failed requests, and retry only failures that are safe to retry.
- Control versions: Puppeteer and Chromium changes can affect layout, fonts, and link behavior. Pin versions and inspect representative output after upgrades.
- Estimate infrastructure cost: account for Chromium memory, concurrency limits, storage, and time spent debugging blocked or incomplete pages. A self-hosted browser is appropriate when you need full control; an API can be simpler for occasional or variable workloads.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. It can return PNG, JPEG, WebP, or PDF from one GET request, with options for PDF paper size, margins, landscape mode, and page ranges. Before capture, it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/reports/monthly \
-o report.pdf
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://example.com/reports/monthly",
},
timeout=90,
)
r.raise_for_status()
open("report.pdf", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/reports/monthly'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('report.pdf', data));
ScreenshotNeo also supports custom headers, cookies, user agents, authorization, waiting for selectors or network idle, custom JavaScript and CSS, request blocking, caching with a chosen TTL, asynchronous jobs with signed webhooks, bulk capture, and an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents. A free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no cost.
FAQ
Can I pass baseURL to page.setContent()?
No. Add a <base href> element, rewrite references as absolute URLs, or navigate to the source URL.
Should the base URL end with a slash?
Use a trailing slash when it represents a directory. Without it, the final path segment can be replaced during resolution.
Does networkidle0 guarantee that every image is ready?
No. Wait for the specific images, fonts, or application state your document needs, then inspect the output.
Why do links work in HTML but not in my PDF viewer?
PDF link annotations can vary with Puppeteer, Chromium, and viewer versions. Validate the generated file with the exact versions and viewers in your deployment.
When should I use navigation instead of a base element?
Use navigation when the page exists at a real URL. Use a base element or absolute references for generated HTML strings.


