ScreenshotNeo

BlogHTML to image & PDF

How to Convert a List of URLs to PDF with Puppeteer

Use Puppeteer to turn each URL into its own PDF, with complete Node.js code, print settings, error handling, and an optional screenshot API shortcut.

By the ScreenshotNeo team4 October 20269 min read

To convert a list of URLs into PDFs with Puppeteer, loop through the URLs, navigate a page to each one with page.goto(), then call page.pdf() with a distinct output path. Puppeteer exports the current page; the loop, naming, and per-URL error handling are application code. The example below creates one PDF per URL.

1. Install Puppeteer and prepare a URL list

Use a Node.js project and install Puppeteer, which downloads a compatible browser as part of its normal installation flow:

npm init -y
npm install puppeteer

Save the following as convert-urls.mjs. The sample uses absolute HTTPS URLs. Ensure every input includes a scheme such as https://; a hostname alone is not a complete navigation URL.

2. Convert each URL into its own PDF

import puppeteer from 'puppeteer';

const urls = [
  'https://example.com',
  'https://example.org',
];

const browser = await puppeteer.launch();
const results = [];

try {
  const page = await browser.newPage();

  for (const [index, url] of urls.entries()) {
    const outputPath = `output-${String(index + 1).padStart(3, '0')}.pdf`;

    try {
      const response = await page.goto(url, {
        waitUntil: 'networkidle2',
        timeout: 45_000,
      });

      // goto can return null for special navigations; when there is a
      // response, optionally treat HTTP error status codes as failures.
      if (response && !response.ok()) {
        throw new Error(`HTTP ${response.status()}`);
      }

      await page.pdf({
        path: outputPath,
        format: 'A4',
        printBackground: true,
        margin: {
          top: '12mm',
          right: '12mm',
          bottom: '12mm',
          left: '12mm',
        },
      });

      results.push({ url, outputPath, ok: true });
      console.log(`Saved ${url} to ${outputPath}`);
    } catch (error) {
      results.push({ url, ok: false, error: error.message });
      console.error(`Could not convert ${url}: ${error.message}`);
    }
  }
} finally {
  await browser.close();
}

const failures = results.filter((result) => !result.ok);
console.log(`Finished: ${results.length - failures.length} succeeded, ${failures.length} failed.`);
if (failures.length > 0) process.exitCode = 1;

Run it with node convert-urls.mjs. The script reuses one tab and continues to the next input if navigation or PDF generation fails for one URL. Keep the URL in each error record so a failed item can be retried or inspected. The HTTP status check is a policy choice: remove it if you intentionally need PDFs of error pages such as a 404.

The official Puppeteer PDF guide shows navigation followed by page.pdf(). The Page.pdf() API documents its output as a Promise<Uint8Array>; when path is supplied, the file is written there. Always generate a distinct path for each URL or a later file will overwrite an earlier one.

3. Choose when a page is ready

waitUntil: 'networkidle2' is the wait condition used in Puppeteer’s PDF example, but it is not a guarantee that every site’s meaningful content has loaded. Pages with analytics, chat, live updates, or long-running requests may never become idle; other pages may reach network idle before client-rendered content appears.

Approach Use it when Trade-off
load You need standard page load completion. Late API data or lazy content may still be missing.
domcontentloaded The document structure is enough to start a site-specific wait. Images, fonts, and later scripts may not be ready.
networkidle2 A generally quiet network is a useful readiness signal. Does not prove that app content is finished; persistent requests can delay it.
Specific selector The site has a stable element that appears when its data is ready. The selector is site-specific and can change.

For a page with a reliable content marker, wait for that marker after navigation rather than adding an arbitrary sleep:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector('main article', { timeout: 15_000 });
await page.pdf({ path: outputPath, format: 'A4' });

Lazy-loaded images may load only as they approach the viewport. If they matter, use site-specific scrolling or readiness logic and wait for image completion before printing. There is no single wait condition that can infer every site’s content requirements.

4. Configure print output

page.pdf() uses the page’s print CSS media type. If the website has print styles, start with that default. To make the PDF reflect screen media instead, call page.emulateMediaType('screen') before exporting. Puppeteer also adjusts colors for printing by default; CSS -webkit-print-color-adjust: exact can request exact color rendering when the page’s styles are under your control. See the PDF API remarks.

await page.emulateMediaType('screen');
await page.pdf({ path: outputPath, format: 'A4', printBackground: true });

Common PDFOptions include:

Option What it controls Practical note
format Paper format; default is Letter. For example, choose 'A4'. When set, it takes precedence over width and height.
width, height Custom paper dimensions. Use dimensions with units such as '210mm' if not using a standard format.
landscape Landscape orientation. Default is false.
margin Paper margins. Supply top, right, bottom, and left values with units for predictable layout.
printBackground Whether backgrounds and background graphics print. Default is false; set true when colored sections or background images matter.
preferCSSPageSize Whether CSS @page sizing takes priority. Default is false; otherwise CSS sizing is scaled to fit the selected paper.
pageRanges Pages to include, for example '1-5, 8'. Empty string means all pages.
scale Rendering scale. Allowed range is 0.1 to 2; default is 1.
displayHeaderFooter, headerTemplate, footerTemplate Print header and footer templates. Enable display; templates can use date, title, URL, page number, and total pages classes.
waitForFonts Wait for document fonts before printing. Default is true. A background page may need activation for fonts to resolve.
timeout PDF-generation timeout in milliseconds. Default is 30 seconds; zero disables the timeout.
path Where to write output. Relative paths resolve from the process working directory.
omitBackground Whether to hide the default white background. Can allow transparent PDF background.
tagged, outline Experimental tagged output and document outline options. Check behavior against the Puppeteer version in use.

These option names and defaults are version-sensitive; consult the PDFOptions reference for your installed release. For a robust general-purpose setting, choose an explicit paper size, margins, and background policy.

5. Decide how to handle failures and large lists

One failed URL

page.goto() can fail because a URL is invalid, the server is unreachable, TLS validation fails, navigation times out, or the main resource fails. A navigation response can also exist with an HTTP error status; decide whether that should be exported or counted as a failure. Catch failures inside the loop to continue, or move the catch outside the loop if the whole batch should stop at the first error. The Page.goto() API documents navigation behavior and errors.

Many URLs

Sequential processing is the simplest approach and limits simultaneous browser work. For larger lists, keep a bounded number of pages or browser workers in flight and record a result for every input. Unbounded parallelism can consume memory and CPU, overload the browser, or put excessive load on target websites. The right concurrency depends on the pages and runtime; measure your own job instead of assuming a universal safe number. Always close the browser in finally, as in the example.

Repeatable file names

Index-based names are simple, but they change if list order changes. For stable output names, derive a safe slug or hash from the URL and still make collisions impossible. Avoid using a raw URL as a filesystem path: it can contain slashes, query strings, reserved characters, and sensitive tokens. Save outputs in a known directory and decide whether a rerun should overwrite, skip, or version existing files.

6. Separate PDFs or one combined document?

page.pdf() produces a PDF for the current page. The direct workflow here creates one file per URL. Puppeteer does not provide a multi-URL PDF merge operation in this workflow; if you need one combined document, generate the individual PDFs and merge them in a separate step with a PDF tool or library chosen for your environment.

7. Troubleshoot common problems

Symptom Likely cause Fix
net::ERR_INVALID_URL or navigation rejects immediately Input is missing a scheme or is malformed. Validate each value with new URL(value) and require http: or https:.
Navigation timeout Site is slow or keeps network connections open. Set a suitable navigation timeout and use a readiness condition appropriate to that site instead of relying on network idle.
PDF contains an error page Main navigation returned an HTTP error response, but the script printed it. Check response.ok() and choose whether non-success status pages should be skipped or retained.
Images, data, or charts are missing Lazy content or client-side rendering had not finished. Wait for a site-specific selector or data-ready condition; trigger lazy loading where needed.
Colors or backgrounds differ PDF uses print media and background printing is disabled by default. Set printBackground: true; use screen media only if that is the intended layout.
PDF has unexpected paper size CSS @page dimensions or format precedence changed sizing. Set preferCSSPageSize deliberately; remember format wins over width and height unless CSS page size is preferred.
Fonts look wrong Web fonts have not loaded or are inaccessible. Keep waitForFonts: true, wait for the site’s font readiness when needed, and verify the font resource can load.
Files overwrite each other Every iteration uses the same output path. Include an index or unique stable identifier in the path.
Browser process remains after an error Close logic was skipped after an exception. Place browser.close() in a finally block.
Browser cannot launch in a minimal environment Required browser runtime dependencies or executable setup are missing. Use Puppeteer’s installation guidance for the deployment environment and ensure its browser dependencies are installed.

8. Performance, reliability, and cost

For a small or moderate list, one browser and one reused page keeps the workflow straightforward. Total time depends on target response time, page scripts, assets, readiness waits, and PDF rendering. More concurrency may reduce elapsed time, but uses more resources and sends more simultaneous traffic. A bounded queue, per-navigation timeout, clear per-URL result log, and retry policy for transient failures make batch jobs easier to operate.

Puppeteer itself is a browser automation library, so this workflow’s direct costs are the compute and storage used by the machine or service running it, plus any network costs in that environment. The destination pages may have their own access rules and rate limits. No output size, runtime, or universal cost can be promised without knowing the pages and deployment setup.

Or skip the browser setup

If you need a screenshot of each URL rather than a PDF document, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation. For this guide’s PDF use case, request PDF output using the documented API parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

These examples use the supplied image-output form; use the API’s documented PDF option when the result must be a PDF. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before a shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for free and get 1,000 screenshots a month with no card.

FAQ

Does Puppeteer combine all the URLs into one PDF automatically?

No. This workflow creates one PDF for each page. Combining those files is a separate step.

Can I print only selected pages from a website?

Yes. Use the pageRanges PDF option, such as '1-3, 7', to select output pages.

Does PDF generation wait for web fonts?

Yes, waitForFonts defaults to true in the documented options. If fonts still appear incorrectly, check whether the font resources load and whether the page is active when font readiness is awaited.

Can I convert a URL list from a text file?

Yes. Read the file in Node.js, split it into lines, trim whitespace, discard blank lines, then use the resulting array in the same loop. Validate each URL before navigating.

References