ScreenshotNeo

BlogHow-to

Take screenshots of a list of URLs using Python

Use Playwright to capture a list of URLs in Python, save each image with a unique filename, and handle pages that load or fail differently.

By the ScreenshotNeo team4 October 20268 min read

Use Playwright for Python to open each URL in a browser and save a screenshot to its own file. The example below captures the visible viewport, gives every file a distinct name, and continues if one URL fails. Set full_page=True to capture the full scrollable page.

1. Install Playwright

Install the Python package and its browser binaries in the environment where the script will run:

python -m pip install playwright
python -m playwright install chromium

Playwright’s browser builds are versioned alongside the package. If you pin Playwright in a project, install the browser build for that installed version and consult the release notes when you need version-specific capabilities.

2. Capture a list of URLs

Save this as capture_urls.py. It expects one URL per line in urls.txt. Blank lines and lines beginning with # are ignored. It records successes and failures in a CSV manifest.

from pathlib import Path
from urllib.parse import urlparse
import csv
import re

from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

INPUT_FILE = Path("urls.txt")
OUTPUT_DIR = Path("screenshots")
MANIFEST_FILE = OUTPUT_DIR / "manifest.csv"
NAVIGATION_TIMEOUT_MS = 30_000
FULL_PAGE = False


def load_urls(path: Path) -> list[str]:
    return [
        line.strip()
        for line in path.read_text(encoding="utf-8").splitlines()
        if line.strip() and not line.lstrip().startswith("#")
    ]


def filename_for(index: int, url: str) -> str:
    parsed = urlparse(url)
    # Keep filenames portable and bounded; the index prevents collisions
    # between paths on the same host and repeated URLs.
    label = (parsed.netloc + parsed.path).strip("/") or "page"
    label = re.sub(r"[^A-Za-z0-9._-]+", "-", label).strip("-._")
    label = label[:100] or "page"
    return f"{index:03d}-{label}.png"


urls = load_urls(INPUT_FILE)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

with MANIFEST_FILE.open("w", newline="", encoding="utf-8") as manifest:
    writer = csv.DictWriter(manifest, fieldnames=["index", "url", "file", "status", "error"])
    writer.writeheader()

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch()
        try:
            page = browser.new_page(
                viewport={"width": 1440, "height": 900},
                device_scale_factor=1,
            )
            page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

            for index, url in enumerate(urls, start=1):
                filename = filename_for(index, url)
                output_path = OUTPUT_DIR / filename
                try:
                    response = page.goto(url, wait_until="load")
                    # HTTP errors such as 404 can still render a page. Record
                    # the status while preserving the screenshot for inspection.
                    status = response.status if response else "no response"
                    page.screenshot(path=str(output_path), full_page=FULL_PAGE)
                    writer.writerow({
                        "index": index, "url": url, "file": filename,
                        "status": f"captured (HTTP {status})", "error": "",
                    })
                except PlaywrightTimeoutError as exc:
                    writer.writerow({
                        "index": index, "url": url, "file": filename,
                        "status": "failed", "error": f"navigation timeout: {exc}",
                    })
                except Exception as exc:
                    writer.writerow({
                        "index": index, "url": url, "file": filename,
                        "status": "failed", "error": f"{type(exc).__name__}: {exc}",
                    })
                manifest.flush()
        finally:
            browser.close()

Create urls.txt beside the script:

https://example.com
https://playwright.dev/python/docs/screenshots

Run it with python capture_urls.py. The images go into screenshots/; the manifest maps each input URL to its output file and records failures. The index in each filename avoids overwriting files when URLs share a host or repeat.

3. Choose what to capture

Goal Setting or API What to expect
Visible viewport page.screenshot(path="...") Captures the current viewport. Set a fixed viewport for comparable dimensions.
Full scrollable page page.screenshot(path="...", full_page=True) Produces a taller image covering the page’s full scrollable area. Long pages can use substantial memory and create large files.
One element page.locator(".report").screenshot(path="report.png") Captures a matching element; Playwright scrolls it into view. A scrollable container shows its currently scrolled content. An overlay can obscure the element. See the Locator API.
Image bytes image_bytes = page.screenshot() Returns bytes instead of writing to a path; useful for post-processing or sending to another service.

To capture an element instead of the whole page, replace the screenshot line in the loop with a locator call:

page.locator("main article").screenshot(path=str(output_path))

Use a selector that identifies one element on every target page. If pages have different structures, handle selector failures per URL or use separate selectors by site.

4. Decide when a page is ready

The example waits for the browser’s load event. That is a useful starting point, but no single readiness condition fits every site. Some pages render important content later; others keep network requests open indefinitely.

  • wait_until="domcontentloaded" waits for the initial document to be parsed and can be quicker when later resources are irrelevant.
  • wait_until="load" waits for the page load event, as in the example.
  • wait_until="networkidle" waits for network activity to settle, but can be unsuitable for pages with polling, analytics, or other continuing requests.
  • For a page-specific signal, navigate and wait for an element that indicates the content you need:
page.goto(url, wait_until="domcontentloaded", timeout=30_000)
page.locator("main h1").wait_for(state="visible", timeout=10_000)
page.screenshot(path="ready.png")

A fixed delay is also possible when a site has a known animation or delayed widget, but it adds time to every capture and does not prove that the desired content loaded. Prefer a meaningful selector when one exists. Navigation and waiting methods are documented in the Page API.

5. Make captures more repeatable

For comparison work, keep the viewport and device scale fixed, use the same browser build, and capture at a consistent readiness point. Even then, page content can vary because of personalization, rotating ads, timestamps, consent dialogs, authentication state, and asynchronous widgets.

Playwright screenshot options can help control output. For example, animations can be disabled during capture, and screenshot styling can hide or stabilize selected elements where supported by the relevant screenshot method and installed version. Check the screenshot guide and Locator API for the options available to your method.

Use full-page capture when the entire document matters; use viewport capture when consistent dimensions or quick visual comparison matter. PNG preserves image detail but can produce larger files than compressed formats. Format support depends on the Playwright and browser versions in use, so verify current support in the release notes.

6. Handle input and output edge cases

  • Invalid or incomplete URLs: provide fully qualified URLs such as https://example.com. Validate input before launching the batch if URLs come from an untrusted source.
  • Duplicate URLs: the sequence number gives each occurrence a separate file. Remove duplicates first if one capture per unique URL is the goal.
  • Same host, different paths: the filename includes a sanitized path and index. Keep the manifest when you need the exact source-to-file mapping.
  • Non-success HTTP responses: a 404 or 500 response may still render content. The example records the HTTP status and attempts a screenshot; decide whether your workflow should instead treat particular status codes as failures.
  • Authentication: pages behind a login will show the unauthenticated state unless the browser context has the required session. Do not put credentials directly in source code or the manifest.
  • Very long pages: full-page images can be tall and expensive to process. Consider viewport screenshots, element screenshots, or splitting the task if only selected sections are needed.
  • Partial batches: per-URL exception handling lets later URLs proceed after an earlier failure. The manifest is flushed after each result so progress is retained if the process stops.
  • Concurrent execution: the sample is sequential and easy to debug. Parallel pages may improve throughput for large batches, but consume more memory, CPU, browser processes, and network capacity; choose a small bounded concurrency and measure your own workload.

7. Troubleshooting

Symptom Likely cause Fix
Executable doesn't exist or browser launch fails The Playwright package is installed, but its browser build is missing or does not match the installed version. Run python -m playwright install chromium in the same environment and check the release documentation if versions are pinned.
Navigation timeout The site is slow, unreachable, or keeps resources active; the chosen timeout or wait condition does not fit. Check the URL and network access, raise the timeout only when justified, or use domcontentloaded and then wait for a page-specific element.
Screenshot is blank or incomplete The page may render content after the chosen event, require scrolling for lazy content, or block automated browsers. Wait for the relevant content selector, inspect the page state, and handle sites that require authentication or reject browser automation. Do not assume a screenshot proves the page was fully available.
Output file is overwritten Multiple captures use the same path. Include an index or unique identifier in every filename and retain a URL-to-file manifest.
Element selector times out The selector does not exist on that URL, matches no element, or the page has not reached the expected state. Check the selector against that page, wait for the intended readiness signal, and record the failure for that URL.
Full-page capture consumes too much memory The document is very long or has large rendered content. Use viewport or element capture, reduce the batch concurrency, or process fewer pages at once.
Images differ between runs Dynamic content, browser versions, viewport differences, or timing changed. Fix the browser build and viewport, use a consistent readiness signal, and hide or stabilize known variable elements where appropriate.

8. Performance, reliability, and cost

The sample reuses one browser and one page, then visits URLs sequentially. This avoids repeatedly starting a browser, but total runtime still depends on each target site’s response, readiness condition, and screenshot size. Full-page images and many large pages increase memory use. There are no universal timing figures for this workload; measure against your actual URL list.

For more throughput, use a bounded number of pages or browser contexts rather than launching an unbounded task per URL. Keep per-URL timeouts and failure records, and close the browser in a finally block as shown. A batch script has no service charge beyond the compute, network, and storage resources of the machine running it; large or frequent jobs can still incur infrastructure costs.

Or skip the browser setup

If you need screenshots without installing and managing a browser, ScreenshotNeo is a website screenshot API and MCP server. Its API accepts one GET request with a URL and can return PNG, JPEG, WebP, or PDF. The API documentation has the request options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp

For a list, call the endpoint once per URL and save each response using a unique filename, just as in the Python batch above. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

FAQ

Can I capture a list of URLs from a CSV?

Yes. Read the URL column with Python’s CSV module and pass each value through the same per-URL capture loop. Keep the original row number in the manifest if you need to trace output back to the input.

Does a successful screenshot mean the site returned HTTP 200?

No. A browser can render an error page such as a 404. Record the navigation response status separately if that distinction matters.

Can I save screenshots as bytes instead of files?

Yes. Call page.screenshot() without a path; it returns image bytes that you can process or write elsewhere.