ScreenshotNeo

BlogHow-to

How to Capture Website Screenshots in Bulk with Python in India

Capture a list of websites with Python and Playwright. Choose viewport or full-page images, save failures, and tune the workflow for repeatable jobs.

By the ScreenshotNeo team4 October 202610 min read

To capture website screenshots in bulk with Python, prepare a list of URLs, use Playwright to open each page, and save a screenshot under a unique filename. Playwright provides the browser and screenshot operations; the URL list, output naming, failure handling, concurrency, and retries are parts of the script you build around them. This documented method is the same for developers in India; the sources do not identify an India-specific setup requirement.

The example below captures the visible viewport of each page, records failures, and closes browser resources even if something goes wrong. Use full_page=True when you need the scrollable page instead of just the initially visible area. [Playwright screenshot guide]

1. Install Playwright and prepare URLs

Install the Python package and its browser binaries:

python -m pip install playwright
python -m playwright install chromium

Save one URL per line in urls.txt, for example:

https://example.com/
https://www.python.org/
https://playwright.dev/

Use URLs you are permitted to access, and keep the list explicit so the script only visits the pages you intend to capture.

2. Capture a URL list with synchronous Python

This runnable script creates an output directory, visits each URL in order, saves a PNG, and writes any failures to a CSV file. Each URL gets an indexed filename so repeated or similar hostnames do not overwrite each other.

from pathlib import Path
from urllib.parse import urlparse
import csv

from playwright.sync_api import sync_playwright

INPUT_FILE = Path("urls.txt")
OUTPUT_DIR = Path("screenshots")
FAILURES_FILE = OUTPUT_DIR / "failures.csv"
NAVIGATION_TIMEOUT_MS = 45_000
FULL_PAGE = False


def safe_stem(index: int, url: str) -> str:
    host = urlparse(url).netloc or "page"
    host = "".join(ch if ch.isalnum() or ch in "-." else "_" for ch in host)
    return f"{index:04d}_{host}"


def main() -> None:
    urls = [line.strip() for line in INPUT_FILE.read_text(encoding="utf-8").splitlines()]
    urls = [url for url in urls if url and not url.startswith("#")]
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    failures = []

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch(headless=True)
        try:
            page = browser.new_page(viewport={"width": 1440, "height": 900})
            page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

            for index, url in enumerate(urls, start=1):
                output = OUTPUT_DIR / f"{safe_stem(index, url)}.png"
                try:
                    response = page.goto(url, wait_until="load")
                    # An HTTP error response can still have a page worth inspecting.
                    status = response.status if response else "no response"
                    page.screenshot(path=str(output), full_page=FULL_PAGE)
                    print(f"Saved {output} (HTTP {status})")
                except Exception as error:
                    failures.append((url, str(error)))
                    print(f"Failed {url}: {error}")
        finally:
            browser.close()

    if failures:
        with FAILURES_FILE.open("w", newline="", encoding="utf-8") as csvfile:
            writer = csv.writer(csvfile)
            writer.writerow(["url", "error"])
            writer.writerows(failures)
        print(f"Wrote {len(failures)} failure(s) to {FAILURES_FILE}")


if __name__ == "__main__":
    main()

Run it with python capture_bulk.py after saving the code in capture_bulk.py. page.goto() returns a response for many HTTP error statuses, so the example records the status and still captures the page. A navigation exception, by contrast, is saved as a failure.

3. Choose the capture style

Viewport or full page

The default captures the configured viewport, which is useful when you want a consistent view of what appears above the fold. Set FULL_PAGE = True, or pass full_page=True to the screenshot call, to capture the full scrollable page. Long pages can produce very tall, large images; choose full-page mode only when downstream review or processing needs it. [Playwright screenshot guide]

Whole page or one element

For a page view, use page.screenshot(). To capture a specific component, use a locator:

card = page.locator("main article").first
card.screenshot(path="article.png", animations="disabled", timeout=15_000)

Locator screenshots capture the selected element. The locator API has controls including animation handling and screenshot output options. An element obscured by another element may not appear visibly in the captured result. [Playwright locator screenshot API]

Save to a path or use image bytes

Passing path saves the screenshot directly, as in the batch example. To pass the image to another library or service without first writing a file, omit path; the call returns bytes:

image_bytes = page.screenshot(full_page=False)
# Pass image_bytes to an image-processing or storage function.

Playwright documents screenshot options for image format, clip area, quality, and other controls. Check the API reference for the exact options available on the page or locator method you use. [Guide, page screenshot API]

4. Async alternative for an async Python application

Use the async API if the surrounding program already uses asyncio. This example is sequential: it makes no throughput claim. It records failures and closes the browser in a finally block.

import asyncio
import csv
from pathlib import Path
from urllib.parse import urlparse

from playwright.async_api import async_playwright

OUTPUT_DIR = Path("screenshots_async")


def filename(index: int, url: str) -> str:
    host = urlparse(url).netloc or "page"
    host = "".join(ch if ch.isalnum() or ch in "-." else "_" for ch in host)
    return f"{index:04d}_{host}.png"


async def main() -> None:
    urls = [line.strip() for line in Path("urls.txt").read_text(encoding="utf-8").splitlines()]
    urls = [url for url in urls if url and not url.startswith("#")]
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    failures = []

    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch(headless=True)
        try:
            page = await browser.new_page(viewport={"width": 1440, "height": 900})
            page.set_default_navigation_timeout(45_000)
            for index, url in enumerate(urls, start=1):
                try:
                    response = await page.goto(url, wait_until="load")
                    await page.screenshot(path=str(OUTPUT_DIR / filename(index, url)))
                    print(f"Saved {url} (HTTP {response.status if response else 'no response'})")
                except Exception as error:
                    failures.append((url, str(error)))
                    print(f"Failed {url}: {error}")
        finally:
            await browser.close()

    if failures:
        with (OUTPUT_DIR / "failures.csv").open("w", newline="", encoding="utf-8") as file:
            writer = csv.writer(file)
            writer.writerow(["url", "error"])
            writer.writerows(failures)


if __name__ == "__main__":
    asyncio.run(main())

Playwright’s Python guide shows both synchronous and asynchronous screenshot calls. Choose the style that fits the rest of your application; treat parallelism as a separate design decision. [Playwright screenshot guide]

5. cURL and Node.js alternatives

If the batch is part of a shell pipeline, cURL can save a response directly. For browser automation in JavaScript, Playwright’s Node.js API uses the same browser workflow. These are alternatives to the Python implementation, not additional requirements.

cURL

curl -L "https://example.com/" -o page.html

cURL fetches the document; it does not render a browser screenshot. Use a browser automation tool when the output must be a rendered page image.

Node.js with Playwright

npm install playwright
npx playwright install chromium
const fs = require('node:fs/promises');
const { chromium } = require('playwright');

async function main() {
  const urls = (await fs.readFile('urls.txt', 'utf8'))
    .split(/\r?\n/)
    .map((line) => line.trim())
    .filter((line) => line && !line.startsWith('#'));
  await fs.mkdir('screenshots_node', { recursive: true });
  const browser = await chromium.launch({ headless: true });
  const failures = [];
  try {
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    page.setDefaultNavigationTimeout(45_000);
    for (const [index, url] of urls.entries()) {
      try {
        const response = await page.goto(url, { waitUntil: 'load' });
        const name = `screenshots_node/${String(index + 1).padStart(4, '0')}.png`;
        await page.screenshot({ path: name, fullPage: false });
        console.log(`Saved ${url} (HTTP ${response?.status() ?? 'no response'})`);
      } catch (error) {
        failures.push({ url, error: String(error) });
        console.error(`Failed ${url}: ${error}`);
      }
    }
  } finally {
    await browser.close();
  }
  if (failures.length) {
    await fs.writeFile('screenshots_node/failures.json', JSON.stringify(failures, null, 2));
  }
}

main().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

6. Handle scale, reliability, and repeatability

  • Use unique output names. Include a stable index or an ID derived from the input record; do not rely only on a hostname when the list can contain multiple paths on one site.
  • Keep a failure record. Log the URL and error, then rerun only failed entries instead of restarting the whole batch.
  • Set timeouts deliberately. A navigation timeout limits time spent on pages that never reach the chosen load state. Pick a value appropriate to your pages and environment; there is no universal value in the cited guide.
  • Choose a wait condition intentionally. load waits for the page load event. Some pages continue loading content after that event; if a page needs more time or a known element, add an appropriate wait for that page rather than applying an arbitrary delay to every URL.
  • Start with sequential processing. For larger lists, concurrency can reduce elapsed time but increases browser resource use and requests to destination sites. Add a bounded number of workers, throttling, and retry rules based on your workload; the documentation does not specify universal safe values or throughput guarantees.
  • Retry selectively. Retry transient navigation failures with a small bounded retry count and a delay. Avoid retrying every error blindly; invalid URLs and persistent access denials will not be fixed by repeated attempts.
  • Keep visual test environments consistent. Screenshots can vary with operating system, browser version, settings, hardware, power source, and headless mode. For baseline comparisons, pin the relevant environment and browser version. [Playwright visual comparisons]
  • Plan storage. Full-page images and high pixel dimensions consume more disk and transfer bandwidth than viewport images. Choose an image format and quality appropriate to the downstream use, and rotate or archive output when the batch is recurring.

Playwright’s screenshot documentation describes the API controls; it does not publish a universal batch throughput figure. Measure your own workload if runtime or capacity affects the design.

7. Troubleshooting common failures

Symptom Likely cause What to do
Browser executable not found The Python package is installed, but the browser binary is not. Run python -m playwright install chromium in the same environment used by the script.
Navigation timeout The page did not reach the selected load state before the timeout. Check the URL and network access, increase the timeout if justified, or use a more suitable readiness condition for that page.
Screenshot exists but looks incomplete The page renders content after the load event, or content is lazy-loaded. Wait for a page-specific selector or other readiness signal before capturing. Inspect the page in the same browser context.
Page returns an error status The server returned an HTTP error response; this is distinct from a navigation exception. Keep the status in your output log and decide whether the error page itself is useful. The sample still saves the image.
Several captures overwrite one another Names are based only on a repeated host or fixed filename. Add an input index or unique record ID to every filename.
Element screenshot is blank or obscured The locator resolved to an element covered by another element, or the target is not visible. Confirm the locator targets the intended visible element and inspect overlay behavior before capture. [Locator screenshot API]
Images differ from a previous run Rendering environment or page content changed. Keep browser and host settings consistent for visual comparison, and account for dynamic content. [Playwright visual comparisons]
Batch stops before later URLs An exception escaped the per-URL handler, or cleanup and setup failed. Keep per-URL capture inside its own exception handler and use finally to close the browser, as in the examples.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API takes one GET request with a URL and returns an image or PDF. See the ScreenshotNeo API documentation for request options. This example saves an image response; replace the target URL with the page you want to capture.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', bytes));
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Responses include X-Page-Verdict and X-Billed headers.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month with no card.

FAQ

Does “in India” require a different Playwright method?

The researched Playwright documentation describes the Python workflow without an India-specific variation. Use the same installation and capture pattern, subject to your own network and environment.

Does full_page=True make a screenshot of every route on a site?

No. It captures the full scrollable area of the currently navigated page. Your input list determines which pages are visited.

Can I process screenshot bytes without saving a local file?

Yes. Call page.screenshot() without a path and pass the returned bytes to the next step in your pipeline.

Does Playwright promise a particular bulk capture speed?

The cited documentation describes API behavior, not a throughput guarantee. Runtime depends on the pages, browser environment, network, and your orchestration choices.