ScreenshotNeo

BlogHow-to

How to Bulk Screenshot a URL List with Concurrency Controls in Selenium

Capture a URL list with Selenium using a bounded worker pool, isolated browser sessions, reliable output handling, and per-URL error reporting.

By the ScreenshotNeo team4 October 202610 min read

To bulk screenshot URLs with Selenium, submit each URL to a bounded worker pool and let each worker create, use, and close its own WebDriver session. Set the worker limit explicitly, wait for a condition that fits the pages you capture, save each screenshot to a unique path, and record failures per URL. Selenium has no built-in bulk screenshot command or universal concurrency setting.

The example below uses Python’s ThreadPoolExecutor and Chrome. It captures the current browser window after the document reaches the complete ready state. That condition does not guarantee that every application request, image, animation, or lazy-loaded section is ready.

1. Install Selenium and prepare a URL list

Install Selenium in your Python environment:

python -m pip install selenium

Save one URL per line in urls.txt:

https://example.com/
https://www.python.org/
https://www.selenium.dev/documentation/

Selenium Manager can help locate or manage browser drivers for supported setups. You still need a compatible browser installed. For reproducible batch jobs, pin and maintain your Python, Selenium, browser, and driver environment.

2. Run a bounded worker pool

Save this as bulk_screenshots.py. Each task owns one browser session, and --workers caps how many sessions run at once. The script saves a manifest after all tasks finish, so workers do not contend over a shared results file.

import argparse
import hashlib
import json
import re
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from urllib.parse import urlsplit

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait


def output_name(index: int, url: str) -> str:
    parsed = urlsplit(url)
    host = re.sub(r"[^A-Za-z0-9.-]+", "_", parsed.netloc) or "page"
    path = re.sub(r"[^A-Za-z0-9.-]+", "_", parsed.path.strip("/")) or "root"
    digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:10]
    return f"{index:05d}_{host}_{path}_{digest}.png"


def capture(index: int, url: str, output_dir: Path, page_load_timeout: int) -> dict:
    driver = None
    output_path = output_dir / output_name(index, url)
    result = {
        "index": index,
        "url": url,
        "status": "error",
        "output": str(output_path.resolve()),
        "error": None,
    }

    try:
        options = webdriver.ChromeOptions()
        options.add_argument("--headless")
        options.add_argument("--window-size=1440,1000")
        driver = webdriver.Chrome(options=options)
        driver.set_page_load_timeout(page_load_timeout)
        driver.get(url)

        # Wait for document readiness. This is a baseline, not a guarantee
        # that application data, lazy content, or animations have finished.
        WebDriverWait(driver, 20).until(
            lambda browser: browser.execute_script(
                "return document.readyState"
            ) == "complete"
        )

        saved = driver.save_screenshot(str(output_path.resolve()))
        if not saved:
            raise OSError("WebDriver reported that it could not save the screenshot")
        result["status"] = "success"
    except Exception as exc:
        result["error"] = f"{type(exc).__name__}: {exc}"
    finally:
        if driver is not None:
            try:
                driver.quit()
            except Exception as exc:
                if result["status"] == "success":
                    result["status"] = "error"
                    result["error"] = f"Driver cleanup failed: {type(exc).__name__}: {exc}"

    return result


def main() -> int:
    parser = argparse.ArgumentParser(description="Capture screenshots for a URL list")
    parser.add_argument("--input", default="urls.txt", help="One URL per line")
    parser.add_argument("--output-dir", default="screenshots")
    parser.add_argument("--workers", type=int, default=2)
    parser.add_argument("--page-load-timeout", type=int, default=45)
    args = parser.parse_args()

    if args.workers < 1:
        parser.error("--workers must be at least 1")
    if args.page_load_timeout < 1:
        parser.error("--page-load-timeout must be at least 1")

    input_path = Path(args.input)
    output_dir = Path(args.output_dir).resolve()
    output_dir.mkdir(parents=True, exist_ok=True)

    urls = [line.strip() for line in input_path.read_text(encoding="utf-8").splitlines()]
    urls = [url for url in urls if url and not url.startswith("#")]
    if not urls:
        parser.error(f"No URLs found in {input_path}")

    results = [None] * len(urls)
    with ThreadPoolExecutor(max_workers=args.workers) as executor:
        futures = {
            executor.submit(capture, index, url, output_dir, args.page_load_timeout): index
            for index, url in enumerate(urls, start=1)
        }
        for future in as_completed(futures):
            index = futures[future]
            results[index - 1] = future.result()
            item = results[index - 1]
            print(f"{item['status']}: {item['url']} -> {item['output']}")
            if item["error"]:
                print(f"  {item['error']}")

    manifest = output_dir / "manifest.json"
    manifest.write_text(json.dumps(results, indent=2), encoding="utf-8")
    failures = sum(item["status"] != "success" for item in results)
    print(f"Saved manifest: {manifest}")
    print(f"Completed {len(results)} URLs; failures: {failures}")
    return 1 if failures else 0


if __name__ == "__main__":
    raise SystemExit(main())

Run it with two simultaneous browser sessions:

python bulk_screenshots.py --input urls.txt --output-dir screenshots --workers 2

Raise or lower --workers to fit the available memory and CPU, the number of browser sessions your environment can support, and the target sites’ access policies. A worker pool bounds active tasks; it does not make a slow or resource-heavy page cheap to capture.

3. Choose readiness and screenshot extent

Wait for the right page condition

driver.get() follows the browser’s navigation behavior, but a page can continue fetching application data or rendering content afterward. The example waits for document.readyState == "complete", a useful baseline for document loading rather than a universal signal that the page is visually finished.

For a known site, wait for a meaningful element instead:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

WebDriverWait(driver, 20).until(
    EC.visibility_of_element_located((By.CSS_SELECTOR, "main article"))
)

Choose a selector that indicates the content you need, not merely that some element exists. If you need a fixed settling delay for a known animation, use one sparingly; fixed sleeps add time even when a page is already ready. Set a page-load timeout so navigation cannot wait indefinitely. A separate explicit wait also needs its own timeout, as in the example.

Current window versus full document

The ordinary Python save_screenshot() API saves the current window to a PNG file. It returns False on an I/O error, and Selenium recommends providing a full path. See the Python WebDriver API.

Do not assume that this call captures an entire long page in every browser and driver. Selenium’s Firefox WebDriver API documents separate full-document screenshot methods. If full-page output matters, choose and verify a browser-driver-specific method, or capture the page in sections. Also account for lazy content: scrolling may be necessary to trigger it before capturing.

4. Control concurrency and isolate sessions

  • Use one WebDriver per active worker. A browser session has mutable navigation and window state. Sharing one driver between concurrent tasks can mix commands, URLs, and outputs.
  • Set an explicit worker cap. Start with a small value, then adjust based on memory, CPU, browser startup overhead, and the target sites’ policies. Selenium does not prescribe a universal local worker count.
  • Keep the cap aligned with remote capacity. If using Grid, coordinate the pool size with the sessions that the Grid deployment can actually serve. Extra client workers may wait or fail to obtain a session.
  • Keep outputs unique. The example includes an input index and URL digest, avoiding collisions even if URLs share a host and path. The manifest preserves original URLs and statuses.
  • Decide how failures affect the batch. This example continues after per-URL failures and exits with status 1 if any task failed. A pipeline that must stop immediately can instead cancel pending tasks when its failure policy requires it.

For a modest batch on one machine, local sessions are simpler. For parallel execution across machines, browser versions, or platforms, Selenium Grid routes WebDriver commands to remote browser instances. Grid is optional; see the official Selenium Grid documentation and the Remote WebDriver API.

Remote WebDriver with Grid

Replace the local driver construction inside capture() with a Remote WebDriver connected to your Grid command-executor URL. The Grid URL and available browser capabilities depend on your deployment:

grid_url = "http://localhost:4444"
options = webdriver.ChromeOptions()
options.add_argument("--headless")
driver = webdriver.Remote(command_executor=grid_url, options=options)

Keep driver.quit() in the worker’s finally block. A remote session consumes Grid capacity until it is released or expires. Set --workers to a value your Grid can serve, and monitor session creation and task failures. Selenium documents the remote server model and session capabilities in its Remote WebDriver API.

5. Retries, output, and reliability

The example records one result per input and does not retry. Add retries only for errors likely to be transient, such as temporary session startup or network failures. Use a small retry limit and backoff between attempts; repeated immediate retries can increase load on a failing browser host or target site. Avoid retrying deterministic errors such as an invalid URL without changing the input.

For dependable batch runs:

  • Keep the manifest with the screenshots so each file maps back to its source URL.
  • Persist results periodically for very large batches, or write one result file per task and combine them afterward.
  • Consider writing to a temporary path and renaming only after a successful save if downstream jobs might read the output directory while captures are running.
  • Log exception type and message, and retain enough context to retry only failed URLs.
  • Close each browser with quit() in finally, including when navigation, waits, or file output fail.
  • Respect authentication requirements, robots or rate-limit policies, and the site’s terms. Configure cookies or headers only where you are authorized to do so.

6. cURL, Python, and Node.js alternatives

The Selenium implementation above is the browser-automation version of the job. For a direct screenshot API request, ScreenshotNeo accepts a URL and returns an image or PDF. See the ScreenshotNeo API documentation for request options and response details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o example.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("example.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('example.webp', bytes));

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One request captures a URL, without your script managing local browser sessions. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; responses identify page verdict and billing status. Its MCP server lets AI agents use screenshot tools.

Here is a one-call capture:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o example.webp

ScreenshotNeo includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 screenshots. See the API docs for options. Sign up for 1,000 free screenshots a month, with no card.

Troubleshooting

Symptom Likely cause Fix
Browser or driver fails to start Browser is missing, incompatible, or unavailable in the runtime. Install a supported browser, check Selenium and browser versions, and verify the runtime can launch it. For remote runs, confirm Grid is reachable and has a compatible browser slot.
Navigation times out The site is slow, the network is unavailable, or a page resource is hanging. Set a suitable page-load timeout, record the URL and exception, and decide whether a limited retry is appropriate. Do not raise concurrency to compensate for slow pages.
Screenshot is blank or missing expected content The page has not rendered application data, the wait condition is too broad, or lazy content has not loaded. Wait for a page-specific visible element, trigger lazy loading if needed, and inspect the page at the chosen viewport.
Only the visible viewport appears The regular WebDriver screenshot captures the current window, not necessarily the full document. Use a documented browser-specific full-document method, such as the Firefox API method where applicable, or capture sections and combine them.
Files overwrite one another Output names are derived from a non-unique value such as hostname alone. Include an index and stable URL digest, as in the example, or use a unique task identifier.
Screenshot save returns false WebDriver reports an I/O error, often due to an invalid path or unavailable directory. Create the output directory before capture and pass an absolute path. Check disk permissions and available space.
Sessions accumulate or the machine runs out of memory Drivers are not being quit or too many browsers are active at once. Ensure quit() runs in finally and reduce --workers. On Grid, align workers with available session capacity.
Grid tasks wait for a session The client pool is larger than the Grid’s available capacity, or browser slots are unavailable. Lower the worker limit or add capacity in the Grid deployment; inspect Grid session and node availability.

Performance and cost considerations

Each active worker launches or holds a real browser session, so memory and CPU use grow with concurrency. Browser startup, page weight, network latency, and page readiness all affect throughput. More workers can increase contention or trigger site rate limits; tune with representative pages and conservative concurrency. The dossier provides no universal worker count or benchmark, so measure your own workload.

Local Selenium shifts browser compute and maintenance to your machine or Grid. Grid distributes browser work but adds infrastructure and session management. A screenshot API removes the need to provision browser sessions in your batch runner; compare its plan and options against your capture volume. ScreenshotNeo’s listed plans are Free: 1,000 per month, Starter: $5 for 3,000, Growth: $15 for 15,000, Pro: $39 for 60,000, Scale: $99 for 250,000, and Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.

FAQ

Can multiple Selenium tasks share one browser?

Do not share a WebDriver session across concurrent URL tasks. Give each active worker its own driver so navigation and browser state remain isolated.

Does Selenium have a bulk screenshot command?

No. Read URLs from your input and schedule one normal WebDriver capture task per URL using a pool or another bounded concurrency mechanism.

What should I use as the worker count?

There is no universal value. Start small, then tune for the resources available and the capacity of your local browser host or Grid, while respecting target-site policies.

Can I use a URL list with Grid?

Yes. Keep the same task pattern and create each task’s session with webdriver.Remote pointed at your Grid endpoint. Set the client worker cap to match the Grid capacity you have configured.