How to Generate a PDF Report from Website Screenshots with Node.js
Capture website screenshots with Node.js, place them in an HTML report, and render the report to PDF with Playwright. Learn when to print a page directly instead.
To generate a PDF report that visibly contains website screenshots, capture each page or element with a Node.js browser automation library, place the resulting images in an HTML report, and render that report to PDF. The example below uses Playwright. If you want a printable version of the live page instead of screenshot evidence, call page.pdf() on that page directly.
These are two different outputs: a screenshot is an image of a rendered page, while a browser-generated PDF lays out page content for printing. Playwright and Puppeteer use print styling for PDF generation by default; emulate screen media first if you need screen CSS. See the official Playwright Page API, Puppeteer PDF guide, and Puppeteer Page.pdf() reference.
Choose the output you need
| Need | Approach | What goes in the PDF |
|---|---|---|
| Evidence of exactly what appeared on screen | Capture screenshots, add them to report HTML, then print the report | Screenshot images, captions, timestamps, and any metadata you include |
| A printable page or collection of page content | Navigate to the page and call page.pdf() |
Browser-laid-out content using print CSS by default |
| Screen-styled content in a PDF | Emulate screen media, then call page.pdf() |
PDF layout using screen media styles, subject to PDF pagination |
Choose Playwright or Puppeteer based on the tooling already in your project and the browser setup available in deployment. The APIs cited here support screenshots and PDF output; they do not establish a universal winner or performance ranking.
Build a screenshot report with Playwright
1. Install the package and browser
mkdir screenshot-report
cd screenshot-report
npm init -y
npm install playwright
npx playwright install chromium
Save the following as report.mjs. Run it with node report.mjs. It takes screenshots of two example pages, embeds them in a report layout, and writes website-report.pdf to the current directory. Replace the URLs and labels with your own targets.
import { chromium } from 'playwright';
const targets = [
{ url: 'https://example.com/', title: 'Example home page' },
{ url: 'https://www.iana.org/help/example-domains', title: 'Example domains reference' },
];
const browser = await chromium.launch({ headless: true });
try {
const captures = [];
for (const target of targets) {
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
const response = await page.goto(target.url, {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
if (!response || !response.ok()) {
const status = response ? response.status() : 'no HTTP response';
throw new Error(`Could not load ${target.url}: ${status}`);
}
// Use a site-specific readiness condition when the page renders asynchronously.
await page.locator('body').waitFor({ state: 'visible', timeout: 15_000 });
const image = await page.screenshot({
type: 'png',
fullPage: true,
animations: 'disabled',
timeout: 30_000,
});
captures.push({
url: target.url,
title: target.title,
capturedAt: new Date().toISOString(),
dataUrl: `data:image/png;base64,${image.toString('base64')}`,
});
await page.close();
}
const escapeHtml = (value) => value.replace(/[&<>"']/g, (char) => ({
'&': '&', '<': '<', '>': '>',
'"': '"', ''': ''',
})[char]);
const sections = captures.map((capture, index) => `
<section class="capture">
<h2>${index + 1}. ${escapeHtml(capture.title)}</h2>
<p class="meta">${escapeHtml(capture.url)}<br>Captured: ${escapeHtml(capture.capturedAt)}</p>
<img src="${capture.dataUrl}" alt="Screenshot of ${escapeHtml(capture.title)}">
</section>
`).join('\n');
const reportHtml = `
<!doctype html>
<html>
<head>
<meta charset="utf-8">
<title>Website screenshot report</title>
<style>
@page { size: A4; margin: 16mm; }
body { font: 12px/1.45 Arial, sans-serif; color: #17212b; }
h1 { font-size: 24px; margin: 0 0 20px; }
h2 { font-size: 17px; margin: 0 0 8px; overflow-wrap: anywhere; }
.meta { color: #52606d; overflow-wrap: anywhere; }
.capture { break-after: page; }
.capture:last-child { break-after: auto; }
img { display: block; width: 100%; height: auto; object-fit: contain; }
</style>
</head>
<body>
<h1>Website screenshot report</h1>
${sections}
</body>
</html>
`;
const reportPage = await browser.newPage();
await reportPage.setContent(reportHtml, { waitUntil: 'load', timeout: 45_000 });
await reportPage.pdf({
path: 'website-report.pdf',
format: 'A4',
printBackground: true,
margin: { top: '16mm', right: '16mm', bottom: '16mm', left: '16mm' },
});
await reportPage.close();
console.log('Created website-report.pdf');
} finally {
await browser.close();
}
The HTML escaping helper protects titles and URLs used as text. The screenshot data URLs are generated from the captured bytes and embedded as image sources. For very large reports, consider writing images to files or another accessible store and referencing those files while rendering, rather than holding every image and the full report HTML in memory.
2. Tune capture readiness and report layout
- Wait for the right state:
domcontentloadedgets the initial document parsed; it does not guarantee that a client-rendered chart or image is ready. Add a site-specific locator wait, or choose another navigation condition appropriate to the page. Avoid assuming that a fixed delay works for every site. - Full page or viewport: The example uses
fullPage: true. Remove it for a viewport screenshot. Long pages can produce very tall images and PDFs; capture a specific element or split the report when that better fits the reader’s needs. - Paper and margins:
format: 'A4'selects a page format, and the margin object reserves space around the printed content. Puppeteer documents configurable dimensions and margins; consult the installed library version’s reference for exact option support. - Page breaks: The example starts each capture on a new page. Remove or adjust
break-afterif you prefer multiple smaller captures per page. A single tall image may span pages according to print layout. - Backgrounds:
printBackground: trueincludes background graphics in the report PDF.
Print a live page directly instead
If screenshot images are not required, navigate to the target and generate a PDF from its page content. This is shorter and keeps content in the browser’s print layout rather than embedding a screenshot.
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/', {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
margin: { top: '16mm', right: '16mm', bottom: '16mm', left: '16mm' },
});
} finally {
await browser.close();
}
PDF generation uses print media by default. If the page’s screen styles are the intended appearance, emulate screen media before generating the PDF:
await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'screen-styled-page.pdf', printBackground: true });
This still produces a paginated PDF; screen styling does not turn the PDF into a screenshot. For screenshot evidence, capture with page.screenshot() and include the image in report HTML as in the first example.
cURL, Python, and Node.js options
For a DIY Node.js workflow, use Playwright’s page.screenshot() and page.pdf() methods as shown above. Puppeteer offers its own Page.pdf() API; check its official guide and the documentation for the version installed in your project.
cURL and Python do not call Playwright’s Node.js API directly. They can call a screenshot service’s HTTP API, or invoke a separate service you build around browser automation. For a single website capture, ScreenshotNeo exposes a GET endpoint that returns an image or PDF. The examples below use the documented endpoint and parameters; see the ScreenshotNeo API documentation for request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', bytes));
These service examples save a screenshot response. For a multi-image PDF report, fetch the required images and place them into your own HTML report before rendering it, or use the service’s PDF output for a page capture when that matches your report needs. Do not assume a single screenshot response automatically contains a custom report layout.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its HTTP endpoint can return a screenshot or PDF from one GET request, while the DIY Playwright examples above give you control over a multi-page report layout.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Read the API docs or sign up for 1,000 free screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser launch fails in deployment | The Chromium browser binary is missing or the environment lacks required system dependencies. | Install the browser for the Playwright version in use with npx playwright install chromium; follow the runtime’s browser dependency instructions. |
| Navigation times out | The page is slow, keeps connections open, or never reaches the chosen wait condition. | Use a navigation condition suited to the site, set a considered timeout, then wait for a specific element that means the content you need is ready. |
| Screenshot misses charts or client-rendered content | The page’s initial document loaded before those elements finished rendering. | Wait for a locator, chart-ready signal, or other page-specific state before capturing. |
| Images are missing in the report | Image resources have not loaded, a URL is inaccessible from the PDF renderer, or the report references the wrong path. | Wait for the needed images, inspect the generated HTML and resource paths, or embed the captured screenshot bytes as data URLs. |
| PDF looks different from the browser | page.pdf() uses print media by default, or print CSS changes layout. |
Inspect print styles; call page.emulateMedia({ media: 'screen' }) when screen media is desired. |
| Content is clipped or split awkwardly | The report image or section exceeds the printable area, or page-break rules conflict with the layout. | Adjust image width, paper format, margins, or break rules; for long pages consider separate captures or a deliberate report layout. |
| PDF is unexpectedly large | Many high-resolution or full-page PNG screenshots are embedded. | Use fewer or smaller captures, choose JPEG where suitable, or resize images before adding them to the report. |
| Service request returns an error | The key, URL encoding, access, or request parameters may be incorrect. | Check the key and encoded target URL, inspect the HTTP response and service headers, and confirm parameter names in the API docs. |
Performance, reliability, and cost
- Browser resources: Launching Chromium and keeping pages open consumes memory and CPU. Reuse a browser process for a batch, close pages after capture, and always close the browser in a
finallyblock. - Report memory: Base64 images increase their representation size, and embedding many full-page captures keeps both image data and HTML in memory. For large batches, write images to disk or object storage and render incrementally.
- Reliability: Check navigation responses, use explicit timeouts, wait for content-specific readiness, and make retries bounded. A capture can succeed technically while still showing an application error or incomplete page, so inspect output when correctness matters.
- Consistency: Fix viewport, locale, timezone, authentication state, and capture timing when reports need comparable screenshots. Dynamic content can still change between runs.
- Cost: DIY has no per-screenshot API fee, but you operate the browser runtime and its compute. ScreenshotNeo has a free tier of 1,000 shots/month without a card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.
FAQ
Does page.pdf() put a screenshot into the PDF?
No. It generates a PDF from page content and print layout. Capture an image and include it in report HTML when the PDF must show screenshot evidence.
Can a report include both a screenshot and notes?
Yes. Build an HTML report with images, captions, URLs, timestamps, and other metadata, then render that HTML to PDF.
Should I use Playwright or Puppeteer?
Use the one already supported by your project and deployment. Both document screenshot or PDF workflows; select based on your browser and runtime needs.


