ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Web Pages with Docker and Playwright

Build a repeatable Docker and Playwright workflow to capture URL lists, save viewport or full-page screenshots, and handle failures safely.

By the ScreenshotNeo team4 October 202610 min read

To bulk screenshot web pages with Docker and Playwright, run a script that reads URLs, opens each page in a browser, and saves page.screenshot() output to a mounted directory. Use a Playwright Docker image pinned to the same version as the Playwright package in your project. The image supplies browser binaries and system dependencies; install the Playwright package separately. Set fullPage: true when each image should include the page’s scrollable document.

This guide uses Node.js. It processes URLs sequentially for a straightforward baseline, gives each URL a deterministic output name, records failures, and shows how to run it in Docker. Concurrency and retries should be chosen for your workload; the official documentation does not prescribe universal values.

1. Create the project and URL list

Use a project directory with these files:

bulk-shots/
  Dockerfile
  package.json
  capture.mjs
  urls.txt
  output/

Put one URL per line in urls.txt. Blank lines and lines beginning with # are ignored.

https://example.com
https://playwright.dev/
# Add another URL here

Initialize the Node project and install Playwright. Keep the package version aligned with the Docker image tag you choose. The Docker documentation recommends pinning a specific image version where possible; verify the current supported tag in the official Playwright Docker documentation.

npm init -y
npm install playwright

2. Write a bulk capture script

This script uses one browser and one page at a time, visits each URL, and saves PNG files in output/. It writes a JSON Lines manifest containing the URL, filename, and outcome. A failed URL is recorded and does not stop later captures.

import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import path from 'node:path';

const inputFile = process.env.URLS_FILE ?? 'urls.txt';
const outputDir = process.env.OUTPUT_DIR ?? 'output';
const fullPage = process.env.FULL_PAGE !== 'false';
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);
const screenshotTimeoutMs = Number(process.env.SCREENSHOT_TIMEOUT_MS ?? 30000);

if (!Number.isFinite(navigationTimeoutMs) || navigationTimeoutMs <= 0) {
  throw new Error('NAVIGATION_TIMEOUT_MS must be a positive number');
}
if (!Number.isFinite(screenshotTimeoutMs) || screenshotTimeoutMs <= 0) {
  throw new Error('SCREENSHOT_TIMEOUT_MS must be a positive number');
}

const urls = (await readFile(inputFile, 'utf8'))
  .split(/\r?\n/)
  .map(line => line.trim())
  .filter(line => line.length > 0 && !line.startsWith('#'));

if (urls.length === 0) {
  throw new Error(`No URLs found in ${inputFile}`);
}

await mkdir(outputDir, { recursive: true });
const manifestPath = path.join(outputDir, 'manifest.jsonl');
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
const page = await context.newPage();
page.setDefaultNavigationTimeout(navigationTimeoutMs);

try {
  for (const [index, url] of urls.entries()) {
    const id = createHash('sha256').update(url).digest('hex').slice(0, 12);
    const filename = `${String(index + 1).padStart(4, '0')}-${id}.png`;
    const outputPath = path.join(outputDir, filename);
    const started = new Date().toISOString();

    try {
      const response = await page.goto(url, { waitUntil: 'load', timeout: navigationTimeoutMs });
      if (!response) {
        throw new Error('Navigation returned no main-resource response');
      }
      await page.screenshot({ path: outputPath, type: 'png', fullPage, timeout: screenshotTimeoutMs });
      await appendFile(manifestPath, JSON.stringify({
        url, file: filename, status: 'ok', httpStatus: response.status(), started
      }) + '\n');
      console.log(`OK ${url} -> ${filename} (HTTP ${response.status()})`);
    } catch (error) {
      await appendFile(manifestPath, JSON.stringify({
        url, file: filename, status: 'error', error: String(error), started
      }) + '\n');
      console.error(`FAILED ${url}: ${String(error)}`);
    }
  }
} finally {
  await context.close();
  await browser.close();
}

The SHA-256 suffix makes filenames stable for a URL and avoids unsafe path characters. The list index preserves duplicate URLs as separate entries. The manifest is append-only, so remove or rotate it before a fresh run if you want a clean record.

3. Build a version-pinned Docker image

Use the same version in the image tag and the installed package. The example below uses a placeholder so you can supply a currently documented version and keep it consistent with your package lockfile.

# Dockerfile
ARG PLAYWRIGHT_VERSION
FROM mcr.microsoft.com/playwright:v${PLAYWRIGHT_VERSION}-noble
WORKDIR /work
COPY package.json package-lock.json ./
RUN npm ci
COPY capture.mjs urls.txt ./
RUN mkdir -p /work/output
CMD ["node", "capture.mjs"]

Build with a version that matches the installed package, then mount the output directory from the host:

PLAYWRIGHT_VERSION=$(node -p "require('./node_modules/playwright/package.json').version")
docker build --build-arg PLAYWRIGHT_VERSION="$PLAYWRIGHT_VERSION" -t bulk-shots .
docker run --rm \
  -e FULL_PAGE=true \
  -e NAVIGATION_TIMEOUT_MS=30000 \
  -v "$PWD/output:/work/output" \
  bulk-shots

The official image includes browser executables and operating-system dependencies, but not the Playwright npm package. The Docker build installs that package using npm ci. Keep package-lock.json in the project so dependency installation is reproducible. If you update Playwright, update the image tag and lockfile together.

4. Choose what each screenshot captures

Choice Behavior Use it when
Viewport Captures the currently visible browser viewport; omit fullPage or set it to false. You need a consistent screen-sized preview or want to limit image dimensions.
Full page fullPage: true captures the scrollable document as a tall image. You need the complete page in one artifact and can accommodate larger images.
PNG Lossless image output; selected explicitly in the script. Text, UI detail, or downstream pixel comparison matters.
JPEG or WebP Alternative image formats supported by the screenshot API. Your downstream viewer accepts the format and image size is a consideration. Choose format-specific quality settings where supported.

Playwright also supports screenshot options such as output path and format-specific settings. See the screenshot guide and Page screenshot API for the complete options for the installed version.

For a fixed viewport, set the context’s viewport dimensions before navigation, as in the example. Full-page images can be much taller than viewport captures, so check downstream storage, image viewer, and processing limits before choosing full-page for every URL.

5. Add readiness checks and handle difficult pages

waitUntil: 'load' waits for the page’s load event. It does not guarantee that client-rendered content, animations, or data fetched after load are finished. When a page has a known content marker, wait for it before capture:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: navigationTimeoutMs });
await page.locator('main article').waitFor({ state: 'visible', timeout: 10000 });
await page.screenshot({ path: outputPath, fullPage });

Use a selector that is meaningful for the target site. A fixed delay can help with a known short animation, but arbitrary sleeps make a batch slower and still do not prove that the page is ready. Avoid assuming that network idle is suitable for every site; pages with polling or persistent connections may never become idle.

Each iteration reuses a page, which is economical for a trusted set of destinations. If pages interfere through client-side state, create a fresh browser context or page per URL and close it after capture. Contexts isolate cookies and browser storage. Set only the cookies or authentication needed for pages you are authorized to access.

6. Run and inspect the batch

Run the container, then inspect the host’s output/ directory. The files use a stable numbering and hash pattern; manifest.jsonl records successes and errors one JSON object per line.

docker run --rm -v "$PWD/output:/work/output" bulk-shots
cat output/manifest.jsonl

For another URL file, mount it at the expected path and set the environment variable:

docker run --rm \
  -e URLS_FILE=/work/input/urls.txt \
  -e OUTPUT_DIR=/work/output \
  -v "$PWD/urls.txt:/work/input/urls.txt:ro" \
  -v "$PWD/output:/work/output" \
  bulk-shots

The script treats navigation failures, missing main-resource responses, and screenshot errors as per-URL failures. It continues with the next URL and records the error text. HTTP error statuses are still captured if navigation produced a response; inspect the recorded httpStatus if your workflow should reject non-success status codes.

7. Scale the batch responsibly

The sample is sequential: it is easier to debug and puts less simultaneous load on the browser and destination sites, but takes longer as the URL list grows. A bounded worker pool can increase throughput, but it also increases memory use and simultaneous requests. There is no universal concurrency or retry number in the referenced Playwright guidance. Start sequentially, measure on representative URLs in the target container, then raise concurrency gradually while observing browser stability, memory, destination behavior, and failure rates.

  • Use a fresh page or context per concurrent task; do not drive one Page from multiple workers.
  • Set a maximum number of active captures rather than launching one browser page per URL in a large list.
  • Retry only failures that may be transient, with a small bounded attempt count and backoff. Do not repeatedly retry access denials or deterministic page errors.
  • Make output names unique across runs if you need historical versions; the sample names are deterministic and will overwrite the same screenshot on a rerun.
  • Respect each destination’s access rules and avoid sending a burst of requests that the site does not permit.

8. Reliability, rendering differences, and security

For visual baselines, capture and compare in the same environment. Rendering can differ with the operating system, browser version, settings, hardware, power source, and headless mode. Pinning the image and package helps keep those inputs stable, but does not make every external page deterministic. Content, ads, fonts, and third-party resources can change independently.

Playwright’s Docker documentation says its image is intended for testing and development and is not recommended for visiting untrusted websites. Treat arbitrary URL lists as untrusted input: only process destinations you are allowed to visit, isolate the capture environment, and avoid exposing secrets or privileged host resources to pages. See the official Docker guidance.

9. Troubleshooting

Symptom Likely cause Fix
Executable doesn’t exist or browser launch fails The Playwright package and Docker image versions do not match, or the image tag is not pinned correctly. Use the same version for the npm package and image, rebuild the image, and check the official Docker page for valid tags.
npm ci fails No lockfile is present, or it does not match package.json. Run npm install locally to create/update the lockfile, then include both package files in the Docker build.
No screenshots appear on the host The output directory was not mounted at the container’s /work/output path, or the container user cannot write there. Check the mount path and host directory permissions; verify that the script’s OUTPUT_DIR points to the mounted directory.
Navigation timeout The page is slow, blocked, or keeps loading required resources. Choose an appropriate navigation event and timeout, inspect the URL and logs, and use a known selector for page readiness where possible. Do not increase timeouts without checking why the page is waiting.
Screenshot is blank or missing content The page renders content after the chosen load event, requires a user state, or blocks automated browsing. Wait for a relevant visible selector, provide authorized state, and inspect the page in the same container. A bot check may prevent a useful capture.
Full-page capture is extremely tall or errors The document is unusually long or continuously expanding. Use viewport capture, capture a specific element, or establish an application-specific page state before taking the screenshot.
Output files overwrite previous captures The sample derives the filename from URL and list position, which stays the same across reruns. Write each run to a dated directory or include a run identifier in the output name.
Images differ between machines OS, browser version, settings, hardware, power source, or headless mode differ. Run both baseline and comparison captures in the same pinned environment.

10. Cost and performance notes

With self-hosted Playwright, the direct costs are the compute and storage resources used to run the container and retain screenshots. Full-page images, large viewports, and concurrent browser pages can increase memory, processing time, and output storage. The dossier provides no supported pages-per-minute or resource-use benchmark, so estimate capacity by measuring your own representative batch in the environment you plan to use.

Keep the output format and screenshot scope aligned with the artifact you actually need. Viewport captures produce a bounded screen view; full-page captures preserve the scrollable document but may create large images. For repeatable comparisons, reuse the same viewport, image format, browser version, and runtime environment.

Or skip the browser setup

If you need screenshots without maintaining a browser container, ScreenshotNeo is a website screenshot API and MCP server. Send one GET request with a URL to receive a PNG, JPEG, WebP, or PDF. The complete options and parameter reference are in the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Start free with 1,000 screenshots a month and no card.

FAQ

Does the Playwright Docker image include the Node package?

No. It includes browser executables and system dependencies. Install the Playwright package in your project and align its version with the image.

Should I use a new browser for every URL?

Usually a single browser with separate pages is enough for a sequential trusted batch. Use separate contexts when you need isolation of cookies and browser storage.

Can I capture authenticated pages?

Yes, when you are authorized to access them. Supply the necessary session state or authentication in a controlled way, and take care not to expose credentials in logs or images.

Will a full-page screenshot include content that loads only while scrolling?

Not necessarily. Some pages lazy-load images or sections as they approach the viewport. If the required content is absent, trigger the page’s expected loading behavior and wait for its content before taking the capture.