Best Open-Source Tools for Saving Webpages as PDF in Bulk
Compare open-source options for turning URL lists into PDFs, from ArchiveBox to browser automation and self-hosted converters.
For a ready-made workflow that imports a URL list and preserves more than PDFs, start with ArchiveBox. It can snapshot URLs into PDFs alongside HTML, screenshots, WARC, metadata, and extracted article text. If you need a custom PDF-only pipeline, use Puppeteer or Playwright to control a browser and write the batch logic yourself. For an internal HTTP conversion service, consider Gotenberg. The older wkhtmltopdf tool has batch-oriented command-line behavior, but validate it against your target sites before adopting it.
The right choice depends on whether you want URL-list intake, custom browser control, a service endpoint, or multi-format archiving. No tool documentation reviewed here establishes a universal speed or success-rate winner; test representative pages from your own collection.
What “bulk webpage to PDF” involves
A batch converter has two distinct jobs:
- Orchestration: read URLs, choose filenames, limit concurrency, retry failures, and record results.
- Rendering: load each page in a browser or conversion engine and produce a PDF.
ArchiveBox documents importing a text file of URLs, so it covers the first job as part of an archiving workflow. Puppeteer and Playwright provide page-level PDF generation; your script supplies the list processing and operational behavior. Gotenberg exposes URL-to-PDF conversion through an HTTP route, while wkhtmltopdf supports multiple page objects and repeated command input from standard input.
Tool comparison
| Tool | Best fit | Documented capability | What you operate |
|---|---|---|---|
| ArchiveBox | Self-hosted archiving of URL collections | Imports a URL list; snapshots can include PDF, HTML, screenshots, WARC, metadata, and extracted text. | An archiving system and its storage, rather than just a converter. |
| Puppeteer | Custom Node.js automation | Navigate to a page and generate a PDF with Page.pdf(); print CSS is used by default. |
URL intake, filenames, retries, concurrency, authentication, and logs. |
| Playwright | Custom browser automation | Generate a PDF with page.pdf(); print CSS is used by default, and screen media can be emulated. |
The same batch and operational logic as a custom Puppeteer script. |
| Gotenberg | A self-hosted conversion endpoint | HTTP multipart URL-to-PDF conversion using Headless Chromium; its documentation discusses JavaScript, SPAs, and dynamic content. | Deployment and calling the service; URL-list management remains a separate concern. |
| wkhtmltopdf | Existing or legacy CLI workflows | Multiple page objects and repeated command input from standard input. | Compatibility validation: it uses Qt WebKit, and the documentation reviewed does not establish modern-site compatibility or current maintenance status. |
Choose based on the job
- You have a list and want an archive: choose ArchiveBox when PDF is one useful output among several preservation formats.
- You need a tailored converter: choose Puppeteer or Playwright when code should decide readiness, authentication, naming, and failure policy.
- Your team needs an internal HTTP interface: evaluate Gotenberg as the rendering service, then build or connect the URL-list manager separately.
- You already depend on wkhtmltopdf: test current target pages and required CSS behavior before expanding its role.
For any option, make a pilot set containing a static page, a JavaScript-heavy page, a long page, and a page that requires authentication. The reviewed documentation does not promise that every protected or unusual page will render successfully.
Archive a URL list with ArchiveBox
ArchiveBox is the clearest fit in this comparison when you want to submit a collection and retain multiple snapshot formats. Its documentation describes importing URLs from a text file. A typical workflow is:
- Put one URL per line in a text file.
- Import that list using ArchiveBox’s documented list-import workflow.
- Review the snapshot outputs and storage layout, including whether PDF output is enabled for your setup.
- Check a sample of generated PDFs and preserve the accompanying formats if your retention policy needs them.
ArchiveBox is more system to run than a one-purpose converter. That tradeoff makes sense when HTML, screenshots, WARC, metadata, or extracted text are valuable alongside PDF. For a strict PDF-only pipeline, compare the storage and operational overhead with a small browser script or a conversion service.
Build a URL-to-PDF batch with Puppeteer
Puppeteer provides browser control and page-level PDF generation. The script below reads a newline-delimited URL file, generates one PDF per URL, limits concurrent browser pages, and writes a simple success/failure log. It deliberately leaves site-specific authentication and readiness rules visible for you to configure.
npm install puppeteer
// save as save-pdfs.mjs
import { readFile, appendFile } from 'node:fs/promises';
import path from 'node:path';
import puppeteer from 'puppeteer';
const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'pdfs';
const concurrency = Math.max(1, Number(process.env.CONCURRENCY ?? 2));
const timeoutMs = Math.max(1000, Number(process.env.TIMEOUT_MS ?? 45000));
const urls = (await readFile(inputFile, 'utf8'))
.split(/\r?\n/).map(s => s.trim()).filter(Boolean);
function filenameFor(url, index) {
const u = new URL(url);
const stem = `${String(index + 1).padStart(4, '0')}-${u.hostname}${u.pathname}`
.replace(/[^a-zA-Z0-9._-]+/g, '-')
.replace(/-+/g, '-')
.slice(0, 150);
return `${stem || `page-${index + 1}`}.pdf`;
}
await import('node:fs/promises').then(fs => fs.mkdir(outputDir, { recursive: true }));
const browser = await puppeteer.launch({ headless: true });
let next = 0;
async function worker() {
while (true) {
const index = next++;
if (index >= urls.length) return;
const url = urls[index];
const page = await browser.newPage();
try {
page.setDefaultNavigationTimeout(timeoutMs);
const response = await page.goto(url, { waitUntil: 'networkidle2', timeout: timeoutMs });
if (response && response.status() >= 400) {
throw new Error(`HTTP ${response.status()}`);
}
await page.pdf({
path: path.join(outputDir, filenameFor(url, index)),
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
await appendFile('results.log', `OK\t${url}\n`);
} catch (error) {
await appendFile('results.log', `FAIL\t${url}\t${String(error.message).replaceAll('\n', ' ')}\n`);
} finally {
await page.close();
}
}
}
try {
await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
} finally {
await browser.close();
}
Run it with node save-pdfs.mjs urls.txt pdfs. Each line in urls.txt should be a complete URL. The script treats navigation responses with HTTP status 400 or higher as failures and records errors for later review; it does not retry automatically.
Readiness, print styling, and authentication
- Readiness:
networkidle2is a practical starting condition, not a guarantee that a page’s content is complete. Some sites keep requests open; others render content after network activity stops. For those pages, wait for a known selector or a site-specific condition instead. - Print versus screen: Puppeteer PDF generation uses print CSS by default. A page may intentionally hide or rearrange content for print. Inspect output and adjust the page’s print styles or the capture workflow when needed.
- Authentication: use the browser context to establish the required login state or configure request credentials and cookies according to the site. Do not put secrets in a shared URL list or logs.
- Output names: the example includes an index and host to reduce collisions. If paths or query parameters distinguish pages, add a stable identifier; avoid using raw URLs as filesystem paths.
- Retries: rerun only failed rows from the log. Add bounded retries with backoff for transient navigation failures, and keep permanent HTTP errors visible instead of retrying indefinitely.
Build a URL-to-PDF batch with Playwright
Playwright also exposes page-level PDF generation. Its PDF output uses print CSS by default; the documented API can emulate screen media before generating the PDF when that is the desired rendering mode. This runnable example uses the same URL-list approach and limits simultaneous pages.
npm install playwright
// save as save-pdfs-playwright.mjs
import { readFile, appendFile, mkdir } from 'node:fs/promises';
import path from 'node:path';
import { chromium } from 'playwright';
const urls = (await readFile(process.argv[2] ?? 'urls.txt', 'utf8'))
.split(/\r?\n/).map(s => s.trim()).filter(Boolean);
const outputDir = process.argv[3] ?? 'pdfs';
const concurrency = Math.max(1, Number(process.env.CONCURRENCY ?? 2));
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
let next = 0;
function outputName(url, index) {
return `${String(index + 1).padStart(4, '0')}-${new URL(url).hostname}`
.replace(/[^a-zA-Z0-9._-]+/g, '-') + '.pdf';
}
async function worker() {
while (true) {
const i = next++;
if (i >= urls.length) return;
const url = urls[i];
const page = await browser.newPage();
try {
page.setDefaultNavigationTimeout(45000);
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (response && response.status() >= 400) throw new Error(`HTTP ${response.status()}`);
// For print output, keep the default print media behavior.
// To request screen styling instead, call: await page.emulateMedia({ media: 'screen' });
await page.pdf({
path: path.join(outputDir, outputName(url, i)),
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
await appendFile('results.log', `OK\t${url}\n`);
} catch (error) {
await appendFile('results.log', `FAIL\t${url}\t${String(error.message).replaceAll('\n', ' ')}\n`);
} finally {
await page.close();
}
}
}
try {
await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
} finally {
await browser.close();
}
Run node save-pdfs-playwright.mjs urls.txt pdfs. The example waits for the DOM to load, which can be preferable for pages whose network activity never settles. If content is inserted later, add a selector wait or another explicit readiness rule before calling page.pdf().
Use Gotenberg as an HTTP conversion service
Gotenberg packages URL-to-PDF conversion behind an HTTP multipart route and uses Headless Chromium. Its documentation calls out JavaScript, single-page applications, and dynamic content. This is useful when several internal clients need a conversion service rather than each embedding browser automation.
Use the URL conversion route documented by Gotenberg and submit the target URL as a multipart request field. The exact route and supported form fields depend on the deployed Gotenberg version, so use its official route documentation when wiring the client. Put URL iteration, output naming, request concurrency, retries, authentication handling, and result logging in your caller or job system.
For reliability, treat each conversion as an individual job: retain the source URL, service response, output path, and failure reason. Limit simultaneous requests to what your deployment can handle, and test long pages and dynamic sites. The documentation establishes the rendering approach, not a throughput guarantee.
Use wkhtmltopdf for an existing CLI workflow
The wkhtmltopdf manual documents multiple page objects and a standard-input mode for repeated invocations. Those features can fit established command-line pipelines. However, the project overview identifies its Qt WebKit basis, and the reviewed material does not establish current compatibility with modern sites or current maintenance status.
Before relying on it, test representative pages with the actual CSS, JavaScript, fonts, and authentication requirements in your workload. If important content is missing or layout differs, compare a Chromium-based option such as Puppeteer, Playwright, or Gotenberg. Avoid assuming that a successful process exit means the resulting PDF contains the expected page content.
PDF choices that affect the result
- Paper size and page breaks: use the page’s CSS page rules when appropriate, or choose a fixed paper format in the PDF API. Long pages can span many sheets; review page breaks and repeated headers or footers.
- Backgrounds: browser print output may omit background colors or images unless configured. The examples enable background printing.
- Print CSS: print styles can remove navigation or alter layout. That may improve a document or hide content you expected; inspect samples.
- Screen styling: Playwright can emulate screen media before generating a PDF if screen CSS is required. Confirm that the resulting page dimensions and pagination are acceptable.
- Dynamic content: wait for the content that matters, such as a table or article body. A generic load event cannot prove that a page finished its own asynchronous work.
- Very long pages: consider splitting the collection or handling especially large pages separately. Browser memory and conversion time depend on the page; the reviewed sources provide no universal limits.
Reliability, performance, and cost
These tools have different operating costs. ArchiveBox requires running an archiving system and storing the formats you keep. Puppeteer and Playwright require browser processes, code maintenance, and storage for generated PDFs. Gotenberg requires operating a service and its callers. wkhtmltopdf may be simple to invoke in an existing environment, but compatibility validation is essential.
For custom pipelines, start with low concurrency and increase it only after observing resource use and failure patterns in your own environment. Reuse a browser process while creating isolated pages for jobs, as the examples do. Keep timeouts bounded, log per-URL outcomes, retry transient failures selectively, and avoid redoing successful work. The official documentation reviewed does not provide comparable benchmarks, resource figures, or success rates, so no numerical performance claims are appropriate here.
For cost control, account for compute, storage, retention, and the engineering time needed to support the workflow. Archiving several output formats consumes more storage than retaining PDFs alone, but can preserve information useful later. Keep a small test corpus and review generated documents after changing browser versions, readiness logic, or PDF settings.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF is blank or missing the main content | Capture ran before client-rendered content appeared, or the page returned an error. | Check the response and logs; wait for a content-specific selector or application-ready condition before printing. |
| Navigation times out on a page that appears usable | Long-lived network requests prevent a network-idle condition, or the site is slow. | Use a less restrictive readiness event and wait for the specific content needed. Keep a bounded timeout and record failures. |
| PDF layout differs from the browser | Print CSS is active by default and may change or hide elements. | Inspect print styles. In Playwright, emulate screen media before generating the PDF if screen styling is required. |
| Colors or background images are absent | Print rendering may omit backgrounds by default. | Enable background printing in the PDF options and confirm the page’s print CSS. |
| Some URLs overwrite one another | Output names are based on a non-unique title or hostname. | Include a stable index or identifier in each filename and sanitize characters for the target filesystem. |
| Authenticated pages redirect to login | The browser or conversion service has no valid session. | Establish authentication before navigation using the site’s supported flow, then verify the final URL and page content. Protect credentials and cookies. |
| Batch stops after one bad URL | Errors are not isolated per job. | Catch and log failures around each URL, continue processing, and rerun failed entries separately. |
| High memory use or unstable jobs at larger batches | Too many heavy pages are rendering at once, or pages are not closed. | Lower concurrency, close each page in a finally block, and split large batches. Measure on your own pages. |
| Legacy converter produces broken modern pages | The rendering engine may not support the site’s current behavior. | Validate wkhtmltopdf against current pages; compare a Chromium-based renderer if required content is missing. |
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API can return a PDF as well as PNG, JPEG, or WebP. The following one-call example shows the documented request shape; see the ScreenshotNeo API documentation for PDF output configuration and the available options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
python - <<'PY'
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
PY
node --input-type=module -e "const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(\`https://api.screenshotneo.com/v1/shot?\${q}\`); if (!res.ok) throw new Error(\`HTTP \${res.status}\`); await Bun.write('shot.webp', res);"
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Frequently asked questions
Which tool accepts a URL list with the least custom batch code?
ArchiveBox is the clearest fit in this guide because its documented workflow imports a text-file list and stores multiple snapshot formats.
Do Puppeteer and Playwright automatically manage a whole URL list?
No. Their documented PDF APIs operate on pages. Your application needs to manage input, output names, concurrency, retries, authentication, and logs.
Will these tools save pages behind a login?
They can only render content available to the browser or service session. Configure authentication for the target site and verify the resulting page before treating the PDF as complete.
Which option is guaranteed to preserve every webpage accurately?
None of the reviewed documentation guarantees correct output for every site. Run a representative pilot and inspect the PDFs against your requirements.
Primary documentation
- ArchiveBox — URL-list import and snapshot formats.
- Puppeteer PDF generation.
- Playwright page PDF API.
- Gotenberg URL-to-PDF documentation.
- wkhtmltopdf usage documentation and project overview.


