ScreenshotNeo

BlogHow-to

How to Take Bulk Screenshots with Playwright and Python

Capture many URLs with Playwright’s Python API using reusable browser resources, bounded concurrency, deliberate waits, and per-page error reporting.

By the ScreenshotNeo team4 October 202612 min read

Use Playwright’s async Python API to capture a list of URLs, reuse one browser for the batch, and limit the number of pages running at once. The script below saves full-page PNGs, gives each URL a safe unique filename, records success or failure for every item, and closes browser resources even if a capture fails. It uses only Python’s standard library plus Playwright.

Install Playwright and its Chromium browser, save the script as bulk_screenshots.py, then run it with a newline-separated urls.txt file:

python -m pip install playwright
python -m playwright install chromium
python bulk_screenshots.py urls.txt --output screenshots --workers 3

The worker count is a tuning setting, not a Playwright limit or universal recommendation. Start low and adjust for the machine, page weight, target-site rate limits, and the dimensions of full-page images.

1. Complete runnable async batch script

This implementation starts one Playwright manager and browser, then gives each worker its own browser context. That keeps cookies and local storage separate between captures while avoiding a browser launch for every URL. If the pages need a shared login session, use one shared context instead, as described below.

import argparse
import asyncio
import json
import re
from pathlib import Path
from urllib.parse import urlsplit

from playwright.async_api import async_playwright


def safe_filename(url: str, index: int) -> str:
    """Make a readable, filesystem-safe, unique name for one URL."""
    parsed = urlsplit(url)
    host = parsed.netloc or "page"
    path = parsed.path.strip("/") or "home"
    base = re.sub(r"[^A-Za-z0-9._-]+", "-", f"{host}-{path}").strip("-.")
    return f"{index:04d}-{base[:100] or 'page'}.png"


async def capture_one(browser, semaphore: asyncio.Semaphore, index: int,
                      url: str, output_dir: Path, timeout_ms: int) -> dict:
    result = {"index": index, "url": url, "status": "failed"}
    async with semaphore:
        context = None
        try:
            context = await browser.new_context(
                viewport={"width": 1440, "height": 1000},
                device_scale_factor=1,
            )
            page = await context.new_page()
            response = await page.goto(
                url,
                wait_until="domcontentloaded",
                timeout=timeout_ms,
            )
            # domcontentloaded means the initial document was parsed. It does
            # not promise that every app request, image, or animation is done.
            # Replace this with a site-specific locator when you know the page.
            if response is not None:
                result["http_status"] = response.status
            path = output_dir / safe_filename(url, index)
            await page.screenshot(path=str(path), full_page=True, animations="disabled")
            result.update({"status": "ok", "screenshot": str(path)})
        except Exception as exc:
            result.update({"error_type": type(exc).__name__, "error": str(exc)[:1000]})
        finally:
            if context is not None:
                await context.close()
    return result


async def run(urls: list[str], output_dir: Path, workers: int,
              timeout_ms: int) -> list[dict]:
    output_dir.mkdir(parents=True, exist_ok=True)
    semaphore = asyncio.Semaphore(workers)
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            tasks = [
                capture_one(browser, semaphore, i, url, output_dir, timeout_ms)
                for i, url in enumerate(urls, start=1)
            ]
            # gather retains input order. Each capture handles its own ordinary
            # failures, so one bad URL does not discard the other results.
            return await asyncio.gather(*tasks)
        finally:
            await browser.close()


def main() -> None:
    parser = argparse.ArgumentParser(description="Capture a URL list as full-page PNGs")
    parser.add_argument("url_file", type=Path, help="UTF-8 text file, one URL per line")
    parser.add_argument("--output", type=Path, default=Path("screenshots"))
    parser.add_argument("--workers", type=int, default=3)
    parser.add_argument("--timeout-ms", type=int, default=45000)
    args = parser.parse_args()

    if args.workers < 1:
        parser.error("--workers must be at least 1")
    if args.timeout_ms < 1:
        parser.error("--timeout-ms must be positive")

    urls = [
        line.strip() for line in args.url_file.read_text(encoding="utf-8").splitlines()
        if line.strip() and not line.lstrip().startswith("#")
    ]
    if not urls:
        parser.error("the URL file contains no URLs")

    results = asyncio.run(run(urls, args.output, args.workers, args.timeout_ms))
    manifest = args.output / "results.json"
    manifest.write_text(json.dumps(results, indent=2), encoding="utf-8")
    succeeded = sum(item["status"] == "ok" for item in results)
    failed = len(results) - succeeded
    print(f"Finished: {succeeded} succeeded, {failed} failed. Results: {manifest}")
    if failed:
        raise SystemExit(1)


if __name__ == "__main__":
    main()

Example urls.txt:

https://example.com
https://playwright.dev/python/docs/screenshots
# Lines beginning with # are ignored
https://www.python.org/

Each capture gets a numbered filename, so repeated hosts and paths do not overwrite each other. The JSON manifest preserves the input URL, status, optional HTTP status, screenshot path, and a short error description. A nonzero process exit code signals that at least one capture failed, which is useful in scheduled jobs and CI.

2. Choose what counts as ready

page.goto() supports different navigation milestones. The example uses domcontentloaded to avoid waiting for every subresource. That is a starting point, not a guarantee that a single milestone suits all sites. A client-rendered application may still be loading data; an image-heavy site may need images; a page with long-lived network connections may never become network-idle.

Readiness choice Use it when Tradeoff
commit You need the response to begin and will wait on a known element afterward. Very early; page content may not be parsed yet.
domcontentloaded The document structure is enough or you have a specific follow-up condition. Images and app data may still be loading.
load Load-event resources matter to the capture. Can take longer; some pages continue loading work afterward.
networkidle A brief network quiet period is meaningful for the target. Analytics, polling, and streaming can keep traffic active; it is not a universal “page complete” signal.

When you know the page, prefer a condition tied to its content over an arbitrary sleep. For example, after navigation, wait for the main report heading:

await page.goto(url, wait_until="domcontentloaded", timeout=timeout_ms)
await page.get_by_role("heading", name="Monthly report").wait_for(
    state="visible", timeout=15000
)
await page.screenshot(path=str(path), full_page=True)

For pages that deliberately reveal content on scroll, a full-page screenshot may not trigger every application-specific lazy-loading mechanism. If missing sections matter, scroll in controlled increments and wait for the relevant content or images before capturing; avoid assuming that one fixed delay works across sites.

3. Tune concurrency, contexts, and browser reuse

Bound the work

The semaphore in the script caps concurrent captures. More workers can overlap navigation waits, but each active page consumes memory, CPU, network capacity, and resources at the destination. Large full-page captures raise memory and image-encoding costs. There is no universal documented Playwright worker count. Increase gradually while watching memory, timeouts, and target-site responses; reduce it if failures or throttling rise.

Choose context isolation

A browser context is an isolated browser session. The script creates one context per URL so cookies and local storage do not leak between independent captures. Use a shared context when the batch is meant to use the same session, locale, or emulation settings; create pages in that context and close it after all pages finish. Use separate contexts for separate accounts or independent consent/session state. Contexts share their configured emulation and routing among pages, while separate contexts isolate browser state.

For a shared authenticated session, initialize one context before scheduling captures and pass it to workers. Protect any site-specific login setup from concurrent duplicate execution. Do not share a Playwright instance across OS threads: Playwright documents that its Python API is not thread-safe. For threaded programs, create a separate Playwright instance per thread. Also avoid cancelling a task while it is inside a Playwright call; cancellation during a running call is unsupported and can leave behavior undefined.

Keep lifecycle cleanup explicit

Use try/finally around browser and context resources. Explicitly close contexts created with browser.new_context() before closing the browser so context artifacts are flushed. A failed URL should be recorded without preventing cleanup or hiding the rest of the batch results.

4. Select screenshot scope and output

Need Playwright approach Notes
Visible browser area await page.screenshot(path="view.png") Viewport-sized image using the context viewport.
Entire scrollable document await page.screenshot(path="full.png", full_page=True) Can create a very tall image; check memory and downstream image limits.
One element await page.locator("main article").screenshot(path="article.png") Locator screenshots wait for actionability and scroll the element into view.
Image bytes for upload or processing data = await page.screenshot(full_page=True) Returns bytes, so no temporary file is needed.

A locator screenshot captures the target element; for a scrollable container it captures only the portion currently scrolled into view. Locator screenshots can disable animations, which helps reduce one source of variation. It does not make captures deterministic: page data, fonts, time, ads, and other external inputs can still change.

Set output options deliberately. type supports formats such as PNG and JPEG, and quality applies to JPEG. PNG is lossless and has no quality setting. Use omit_background=True for a transparent screenshot where the browser page and format support that use. scale="css" produces one output pixel per CSS pixel; scale="device" uses the device scale factor and can increase dimensions and file size. Context viewport and device_scale_factor control the emulated screen. Check the Playwright screenshot API documentation for the full option set and format-specific constraints.

5. Synchronous variant and other runtimes

The async API is convenient for a bounded URL batch. A synchronous script is simpler for a small, strictly sequential list. Keep one style throughout a program; do not mix sync calls into an active asyncio workflow.

from pathlib import Path
from playwright.sync_api import sync_playwright

urls = ["https://example.com", "https://playwright.dev/"]
out = Path("screenshots")
out.mkdir(parents=True, exist_ok=True)

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        context = browser.new_context(viewport={"width": 1440, "height": 1000})
        try:
            page = context.new_page()
            for index, url in enumerate(urls, start=1):
                try:
                    page.goto(url, wait_until="domcontentloaded", timeout=45000)
                    page.screenshot(path=str(out / f"{index:04d}.png"), full_page=True)
                    print("OK", url)
                except Exception as exc:
                    print("FAILED", url, type(exc).__name__, str(exc))
        finally:
            context.close()
    finally:
        browser.close()

For Python, install the package and browser binaries in the environment where the script runs. The examples use Chromium, but Playwright also provides Firefox and WebKit browser types; install the desired browser and validate rendering differences if consistency across engines matters.

6. Common failures and fixes

Symptom Likely cause Fix
Executable doesn't exist or browser launch fails The Playwright package is installed but its browser binary is absent, or the runtime environment lacks required system dependencies. Run python -m playwright install chromium in the same environment; on Linux, install the dependencies required by the chosen browser.
Navigation timeout The site is slow, unreachable, or waiting for an overly strict milestone; the timeout may be too low. Confirm the URL is reachable, select a suitable wait_until, wait for a page-specific condition, and adjust timeout deliberately. Record failures per URL.
Screenshot shows a shell or incomplete content Navigation milestone occurred before client-side data or lazy content appeared. Wait for a visible locator or a meaningful app-ready state; scroll when the site loads content on demand.
Some URLs overwrite others Output names were derived only from a hostname or reused constant. Include an input index or stable unique identifier in every output name.
Memory grows or the process is killed Too many pages run at once, or full-page screenshots are extremely tall. Lower the bounded worker count, capture only the needed scope, or process the input in smaller batches.
Images differ between runs Dynamic content, animation, time-dependent data, fonts, ads, or network responses vary. Disable animations where supported, fix viewport and locale, wait for specific content, and control external inputs when the site permits it.
One failure appears to stop all work An exception escaped the per-URL handler or the caller cancelled the batch during a Playwright call. Catch and record ordinary per-item exceptions, retain structured cleanup, and do not cancel tasks while Playwright operations are active.
Element screenshot omits part of a scrollable panel Only the panel’s currently visible portion is captured. Scroll the panel deliberately or capture the page/document instead if that matches the requirement.

7. Performance, reliability, and cost

With Playwright, the main cost is the infrastructure and engineering you operate: browser processes, CPU and memory, network transfer, storage, retries, and maintenance of browser dependencies. The dossier supplies no verified throughput benchmark or universal concurrency value, so estimate using your actual URLs and deployment environment. A representative trial should include slow pages, long pages, and failures, with metrics for elapsed time, peak memory, output size, and per-URL success.

For reliability, keep the input and result manifest, make output names unique, use explicit timeouts, and make retries selective. A retry can help with transient network errors, but repeating a deterministic 404 or blocked page wastes time and may add load. Store screenshots atomically if consumers might read the directory while a batch is running: write to a temporary path and rename after success. For very large lists, stream results or process chunks instead of creating an unbounded number of task objects; the semaphore bounds active captures, while the simple example still creates one task per input URL.

If operating browser infrastructure is not useful for the job, ScreenshotNeo is a website screenshot API and MCP server. Its per-shot API pricing is fixed by plan; the free plan includes 1,000 shots monthly with no card, and paid plans start at $5 for 3,000 shots. Yearly billing gives two months free. Every feature is on every plan. The API reports page verdict and billing status in response headers; the stated billing rule is that only clean shots are billed, while bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing.

Or skip the browser setup

Call ScreenshotNeo once per URL instead of installing and operating a browser. The API accepts a URL and returns a screenshot image or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use the screenshot tools. You get 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Create a free account at ScreenshotNeo sign-up.

8. Frequently asked questions

Can I capture URLs from a CSV?

Yes. Parse the URL column with Python’s csv module and pass its values to the same bounded capture function. Validate and normalize the column before navigation, and preserve the source row identifier in the result manifest.

Can screenshots be compared pixel by pixel?

Yes, page.screenshot() can return bytes for an image-diff pipeline. For useful comparisons, keep viewport and browser settings fixed and account for changing page content; disabling animation alone does not guarantee identical output.

Should I use one browser per URL?

Usually, reuse one browser for a batch and create contexts or pages according to the isolation you need. Launching a browser per URL adds setup and cleanup work; separate browser processes may still be appropriate when strong process isolation is a requirement.

Does full-page capture scroll the page like a person?

It captures the full scrollable document as a tall image. Sites with custom lazy-loading behavior may need deliberate scrolling and readiness checks before the screenshot call.

Primary Playwright references