ScreenshotNeo

BlogHow-to

How to Batch Convert Website URLs to PNG Images with Playwright

Batch capture website URLs as PNGs with Playwright. Get runnable Node.js and Python examples, plus guidance on readiness, failures, output files, and concurrency.

By the ScreenshotNeo team4 October 202610 min read

To batch convert website URLs to PNG images with Playwright, put the URLs in a list, open each one in a browser page, wait for the page state your task needs, and save a screenshot to a unique .png path. Start sequentially for predictable results. Add bounded concurrency only after measuring the workload, since Playwright does not prescribe a universal safe worker count for arbitrary sites.

This guide uses Node.js for the main workflow and includes a Python equivalent, cURL, and ScreenshotNeo’s one-request alternative. The examples save one full-page PNG per URL and record each page’s HTTP status or failure.

1. Install Playwright and prepare a URL list

Use a current Node.js runtime and install Playwright in your project. The first command installs the package; the second downloads its browser binaries:

npm init -y
npm install playwright
npx playwright install chromium

Create a file named urls.txt with one absolute URL per line. Include the scheme, usually https://:

https://example.com
https://playwright.dev

Absolute URLs matter: Playwright navigation expects a URL such as https://example.com, not a bare hostname. Avoid putting credentials or sensitive query values in this file if you will log URLs or share the batch output.

2. Batch capture PNG files with Node.js

Save this as capture.mjs. It creates the output directory, validates the input URLs, reuses one browser context, and visits URLs sequentially. Each input gets an index-based filename, so two URLs with similar paths cannot overwrite each other.

import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';
import path from 'node:path';

const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30_000);

const raw = await readFile(inputFile, 'utf8');
const urls = raw
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line.length > 0 && !line.startsWith('#'));

if (urls.length === 0) {
  throw new Error(`No URLs found in ${inputFile}`);
}

for (const url of urls) {
  let parsed;
  try {
    parsed = new URL(url);
  } catch {
    throw new Error(`Invalid absolute URL: ${url}`);
  }
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error(`Only http and https URLs are supported: ${url}`);
  }
}

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const context = await browser.newContext({
  viewport: { width: 1440, height: 900 },
  deviceScaleFactor: 1,
});

const results = [];
try {
  for (const [index, url] of urls.entries()) {
    const filename = `${String(index + 1).padStart(4, '0')}.png`;
    const outputPath = path.join(outputDir, filename);
    const page = await context.newPage();
    let status = null;

    try {
      const response = await page.goto(url, {
        waitUntil: 'load',
        timeout: timeoutMs,
      });
      status = response?.status() ?? null;

      // Optional site-specific readiness condition goes here, for example:
      // await page.locator('main').waitFor({ state: 'visible', timeout: 10_000 });

      await page.screenshot({
        path: outputPath,
        type: 'png',
        fullPage: true,
      });
      results.push({ url, status, outputPath, ok: true });
    } catch (error) {
      results.push({
        url,
        status,
        outputPath: null,
        ok: false,
        error: error instanceof Error ? error.message : String(error),
      });
    } finally {
      await page.close();
    }
  }
} finally {
  await context.close();
  await browser.close();
}

console.log(JSON.stringify(results, null, 2));
if (results.some((result) => !result.ok)) {
  process.exitCode = 1;
}

Run it with:

node capture.mjs urls.txt screenshots

The output names are stable for a given input ordering, not for a changing URL list: inserting a URL near the start shifts later numbers. If filenames need to represent URL identity across runs, derive a readable sanitized prefix and append a short hash. Do not use an unsanitized URL as a filesystem path; URLs can contain slashes, reserved characters, and private query parameters.

3. Choose capture and readiness options

Viewport or full page

fullPage: true captures the full scrollable page; its default is false, which captures the current viewport. Use viewport capture for consistently sized previews. Use full-page capture for archives or review images that need below-the-fold content. Very long pages can produce large, tall files and may exceed limits in downstream image viewers or pipelines.

PNG output and file handling

When saving with a path, Playwright can infer the image format from the extension. This example also sets type: 'png' explicitly. PNG has no configurable screenshot quality setting; the documented quality option applies to JPEG and WebP. If storage size matters more than lossless output, consider another format only if your downstream use permits it.

You can omit path and use the returned screenshot bytes when you want to upload, transform, or inspect them in memory:

const pngBytes = await page.screenshot({ type: 'png', fullPage: true });
// Pass pngBytes to your storage or image-processing code.

Waiting for the right content

A navigation event does not guarantee that all client-side data, lazy images, fonts, animations, or embedded widgets have settled. The example waits for the page’s load event, then leaves a place for a site-specific condition. If a known element indicates readiness, wait for that element. If a page needs a short delay for a known animation, use a deliberate delay sparingly. There is no single wait rule that guarantees visual readiness on every website.

For pages with lazy-loaded content, a full-page screenshot may not cause every site’s lazy content to load. Where completeness matters, use a site-specific scroll-and-wait routine before capturing, then verify representative outputs. Also account for consent dialogs and overlays: Playwright captures what the browser renders unless your script handles them.

HTTP status and navigation response

A 404 or 500 response does not necessarily make page.goto() throw. The sample stores the navigation response status separately from whether capture succeeded. Decide whether your batch should keep error pages as screenshots or mark non-2xx responses as failed; that is an application policy. Some navigations have no main-resource response, so the code records a null status in that case.

4. Python version

Install the Python package and browser, then save the script as capture.py:

python -m pip install playwright
python -m playwright install chromium
import asyncio
import json
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright

async def main():
    input_file = Path('urls.txt')
    output_dir = Path('screenshots')
    timeout_ms = 30_000

    urls = [
        line.strip()
        for line in input_file.read_text(encoding='utf-8').splitlines()
        if line.strip() and not line.lstrip().startswith('#')
    ]
    if not urls:
        raise ValueError(f'No URLs found in {input_file}')

    for url in urls:
        parsed = urlparse(url)
        if parsed.scheme not in ('http', 'https') or not parsed.netloc:
            raise ValueError(f'Invalid absolute HTTP(S) URL: {url}')

    output_dir.mkdir(parents=True, exist_ok=True)
    results = []

    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        context = await browser.new_context(
            viewport={'width': 1440, 'height': 900},
            device_scale_factor=1,
        )
        try:
            for index, url in enumerate(urls, start=1):
                output_path = output_dir / f'{index:04d}.png'
                page = await context.new_page()
                status = None
                try:
                    response = await page.goto(
                        url, wait_until='load', timeout=timeout_ms
                    )
                    status = response.status if response else None
                    # Add a site-specific readiness wait here if needed.
                    await page.screenshot(
                        path=str(output_path), type='png', full_page=True
                    )
                    results.append({
                        'url': url, 'status': status,
                        'output_path': str(output_path), 'ok': True
                    })
                except Exception as error:
                    results.append({
                        'url': url, 'status': status,
                        'output_path': None, 'ok': False,
                        'error': str(error)
                    })
                finally:
                    await page.close()
        finally:
            await context.close()
            await browser.close()

    print(json.dumps(results, indent=2))
    if any(not result['ok'] for result in results):
        raise SystemExit(1)

asyncio.run(main())

Run it with python capture.py. The Python and Node versions use the same basic model: separate per-URL outcomes, shared context configuration, and unique output paths. Python’s official Playwright screenshot guide also documents full-page screenshots and returning screenshot bytes.

5. Add bounded concurrency when needed

Sequential execution is easiest to debug and puts one active navigation at a time on the target sites. If throughput becomes a requirement, use a bounded worker pool rather than launching one page per URL all at once. Each worker can create a page in the shared context, process one URL, save to a unique path, and close the page in a finally block.

Pick a small initial worker limit and measure browser memory, page failures, target responsiveness, and total completion time on your actual URL set. Reduce the limit if navigation failures or resource pressure rise. Avoid treating concurrency as a Playwright guarantee: the right number depends on the browser host, page weight, and the sites’ policies. Keep output paths unique even with workers, and store results by input index so logs remain readable.

6. cURL and a managed screenshot option

Playwright is useful when you need browser automation, custom interaction, or control over the capture runtime. If the task is simply turning a list of public URLs into images, a screenshot API can remove browser installation and lifecycle management from your script. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. It supports PNG, JPEG, WebP, and PDF; its other capture options are documented in the ScreenshotNeo API documentation.

For a single URL, this cURL request saves a PNG. Repeat it for each URL in your shell or application, changing the target URL and output filename:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=png \
  -o shot.png

Equivalent Python and Node.js requests:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={
        "access_key": "YOUR_API_KEY",
        "url": "https://stripe.com",
        "format": "png",
    },
    timeout=90,
)
r.raise_for_status()
open("shot.png", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'png',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.png', bytes));

7. Troubleshooting

Symptom Likely cause Fix
Executable doesn't exist or browser launch fails The Playwright package is installed but its browser binary is not. Run npx playwright install chromium for Node.js or python -m playwright install chromium for Python. Check that the runtime environment can write to its browser cache.
Navigation times out The site is slow, waiting on a resource indefinitely, or the timeout is too short for that workload. Log the URL and error, adjust the timeout for known slow pages, and choose a readiness condition that matches the site. Keep the failure isolated to that URL.
The script reports success for a 404 or 500 An HTTP error status can be returned without a navigation exception. Inspect the response status and apply your own policy for non-2xx results.
PNG is missing or overwritten The output directory was not created, paths collide, or the process lacks write permission. Create the directory first, use unique names, and check the destination’s permissions and available space.
The screenshot is blank or incomplete The page may still be rendering, depend on client-side data, or defer images until scroll. Wait for a meaningful page element, handle site-specific lazy loading, and inspect the saved navigation status. There is no universal wait condition for all sites.
Content differs between runs Live content and rendering can vary with operating system, browser version, settings, hardware, and headless mode. Keep the capture environment consistent for comparisons and account for changing page content.
A site blocks or challenges the browser The website may restrict automated traffic or require an interaction the script does not perform. Respect the site’s access rules, handle the page as an individual failure, or use an authorized capture method. Do not assume retries will resolve a site-side block.

8. Performance, reliability, and cost

  • Browser lifecycle: Launching one browser for the batch and reusing a context avoids repeating setup for every URL. Close pages after each capture and close the context and browser in cleanup paths.
  • Memory and file size: Full-page screenshots of long pages can be large. Use viewport screenshots when that meets the requirement, and monitor output storage before increasing concurrency.
  • Failure isolation: Record URL, status, output path, and error for every item. A failed URL should not silently discard results already captured.
  • Repeatability: Browser and operating system differences affect visual output. Keep the environment fixed when comparing images, and remember that live pages can change independently of your code.
  • Costs: Running Playwright has no per-capture API charge from Playwright, but your compute, storage, and bandwidth have costs. ScreenshotNeo has a free plan with 1,000 shots per month and no card; paid plans start at $5 for 3,000 shots. Only clean shots are billed, and bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing.

Or skip the browser setup

ScreenshotNeo accepts one GET request per URL, returning an image or PDF. It removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; and its MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. ScreenshotNeo also offers bulk capture for up to 100 URLs per call. See the API documentation for parameters and setup.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Sign up free for 1,000 screenshots a month, no card required.

FAQ

Can one Playwright context handle multiple URLs?

Yes. A context can contain multiple pages, and pages share context-level settings such as viewport and locale. The sequential sample reuses one context while opening and closing a page for each URL.

Should I use Playwright or a screenshot API?

Use Playwright when you need custom browser automation or control over the runtime. Use an API when you want to send URLs and receive image files without managing browser binaries and browser processes.

Can Playwright save the screenshot without writing a file?

Yes. Call page.screenshot({ type: 'png' }) without a path and use the returned bytes in your application.

Will repeated captures of a URL be pixel-identical?

Not necessarily. The page itself may change, and rendering can vary with the operating system, browser version, settings, hardware, and headless mode.

For the complete Playwright option reference, see the Page API, Pages guide, and Python screenshots guide. Learn about ScreenshotNeo if you want a screenshot API and MCP server alongside the self-managed Playwright workflow.