ScreenshotNeo

BlogHow-to

How to Capture a Batch of Web Pages with Selenium and Save Screenshots by Domain

Capture a URL list with one Selenium browser session, save PNGs into hostname folders, and handle waits, failures, naming, and full-page limits.

By the ScreenshotNeo team4 October 202611 min read

Use one Selenium WebDriver session to visit each URL in turn, wait for the content you need, and save a PNG under a folder named for that URL’s hostname. The script below uses Python, writes a JSON manifest that maps each output back to its source URL, records failures without stopping the batch, and always closes the browser.

This example captures the browser’s current viewport. It does not promise a full-page screenshot. Selenium’s driver.get(url) waits for the page load event under the default page-load strategy, but JavaScript-rendered content and later assets may still need a page-specific wait. Selenium WebDriver API · Selenium waiting strategies

1. Install Selenium and prepare URLs

Use Python 3. Selenium 4 can manage compatible browser drivers through Selenium Manager in many standard setups; you still need a supported browser installed. If your environment manages drivers separately, configure the driver path according to that environment.

python -m pip install selenium

Save the script below as capture_batch.py. Replace the sample URLs with the pages you are authorized to capture. Each URL must include a scheme, such as https://.

2. Capture the batch into hostname folders

import hashlib
import json
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlsplit

from selenium import webdriver
from selenium.common.exceptions import TimeoutException, WebDriverException
from selenium.webdriver.support.ui import WebDriverWait

URLS = [
    "https://example.com/",
    "https://www.example.org/products",
    "https://example.com/docs?version=2",
]

OUTPUT = Path("screenshots")
VIEWPORT = (1440, 1000)
PAGE_LOAD_TIMEOUT = 30
READINESS_TIMEOUT = 10


def safe_hostname(url: str) -> str:
    """Return a portable folder name based on the exact hostname."""
    host = urlsplit(url).hostname
    if not host:
        raise ValueError("URL has no valid hostname")
    cleaned = re.sub(r"[^A-Za-z0-9.-]", "_", host.lower()).strip("._")
    return cleaned or "unknown-host"


def output_name(url: str, index: int) -> str:
    """Use an index and URL digest so same-host URLs do not overwrite."""
    digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:10]
    return f"page-{index:04d}-{digest}.png"


def main() -> None:
    OUTPUT.mkdir(parents=True, exist_ok=True)
    records = []

    options = webdriver.ChromeOptions()
    # Uncomment to run without a visible browser window in a suitable environment.
    # options.add_argument("--headless=new")
    driver = webdriver.Chrome(options=options)
    driver.set_window_size(*VIEWPORT)
    driver.set_page_load_timeout(PAGE_LOAD_TIMEOUT)

    try:
        for index, url in enumerate(URLS, start=1):
            record = {
                "index": index,
                "url": url,
                "captured_at": datetime.now(timezone.utc).isoformat(),
                "status": "failed",
            }

            try:
                parsed = urlsplit(url)
                if parsed.scheme not in ("http", "https") or not parsed.hostname:
                    raise ValueError("Use a complete http:// or https:// URL")

                domain_dir = OUTPUT / safe_hostname(url)
                domain_dir.mkdir(parents=True, exist_ok=True)
                filename = domain_dir / output_name(url, index)

                driver.get(url)

                # This confirms document readiness only. For a dynamic page,
                # replace or supplement it with a condition for meaningful content.
                WebDriverWait(driver, READINESS_TIMEOUT).until(
                    lambda d: d.execute_script("return document.readyState") == "complete"
                )

                saved = driver.get_screenshot_as_file(str(filename))
                if not saved:
                    raise OSError(f"Selenium could not write {filename}")

                record.update({"status": "captured", "file": str(filename)})
                print(f"Saved {url} -> {filename}")

            except (TimeoutException, WebDriverException, ValueError, OSError) as exc:
                record["error"] = f"{type(exc).__name__}: {exc}"
                print(f"Failed {url}: {record['error']}")

            records.append(record)
    finally:
        driver.quit()
        manifest = OUTPUT / "manifest.json"
        manifest.write_text(json.dumps(records, indent=2), encoding="utf-8")
        print(f"Wrote manifest: {manifest}")


if __name__ == "__main__":
    main()

Run it from the directory where you want the screenshots/ folder:

python capture_batch.py

The output layout will look like this:

screenshots/
  example.com/
    page-0001-....png
    page-0003-....png
  www.example.org/
    page-0002-....png
  manifest.json

The folder uses the exact hostname, so example.com and www.example.com are separate. The numbered, hashed filenames avoid collisions and preserve batch order. The manifest associates each file or error with the original URL.

3. Choose what “by domain” means

A hostname and a registrable domain are different grouping rules. A hostname includes subdomains: shop.example.com. A registrable domain is generally the site’s public-suffix-aware base, such as example.co.uk. The sample groups by exact hostname because Python’s standard URL parser exposes that value directly. Do not derive registrable domains by simply taking the last two labels: that misgroups names under suffixes such as .co.uk. If you need registrable-domain grouping, use a public-suffix-aware library and define how private suffixes and unusual hostnames should be handled.

Other grouping decisions to make explicitly:

  • Port: the hostname intentionally excludes a port, so example.com:8443 shares the example.com folder. Include the port in the folder key if separate environments should be isolated.
  • Internationalized hostnames: URL parsing may represent these in ASCII-compatible form. Decide whether folder labels should preserve that form or be normalized for display.
  • Queries and fragments: the hash is computed from the full URL string, so distinct query strings or fragments produce distinct files. Normalize URLs first if those differences should be ignored.
  • Repeated URLs: the batch index makes repeated entries distinct files; remove duplicates before the loop if that is not desired.
  • Unsafe names: the hostname sanitizer replaces characters outside letters, digits, dots, and hyphens. Keep the original host in the manifest if a normalized folder could be ambiguous.

4. Wait for the page you actually need

Navigation generally waits for the configured page-load condition, whose default is the page’s complete ready state. That does not prove a single-page application has fetched its data, rendered a chart, or loaded a lazy image. Add a condition tied to the page content that makes the screenshot useful.

For a known page with a stable content selector, use an explicit wait:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

# After driver.get(url):
WebDriverWait(driver, 15).until(
    EC.visibility_of_element_located((By.CSS_SELECTOR, "main article"))
)

Choose a selector that represents finished content, not merely a shell that appears before data arrives. A fixed sleep is sometimes useful while diagnosing a race, but as a permanent readiness rule it may be too short on a slow run and waste time on a fast one. Selenium advises using waits for the condition required by the next action; avoid mixing implicit and explicit waits because it can produce unpredictable total wait times. Waiting Strategies · Expected Conditions

If you want the browser to stop waiting for every resource’s load event and then wait on your own condition, Selenium page-load strategy can be configured as eager or none where supported by the browser. This changes navigation timing; it does not make content ready by itself. Use it only with an explicit readiness condition and verify behavior in your selected browser.

5. Configure the capture

Setting What it controls Practical choice
Window size Viewport dimensions used by the browser screenshot Set one width and height for comparable captures. Responsive layouts change with viewport width.
Page-load timeout Maximum navigation wait before Selenium raises a timeout Choose a ceiling suitable for the batch; handle timeout per URL so one page does not end the run.
Explicit wait Time allowed for a relevant selector or state to appear Set per page or page family; use the condition that signals usable content.
Headless mode Runs a browser without a visible window Useful on servers; confirm installed browser dependencies and compare output with headed mode when diagnosing rendering differences.
Session reuse Shares browser state and startup cost across URLs Efficient for a straightforward batch. Use separate sessions when cookies, logins, or isolation requirements differ.
Screenshot method Captures the current browser window as a PNG Use an explicit path with a .png extension and check the method’s Boolean result.

Selenium documents get_screenshot_as_file(filename) as saving a PNG of the current window; it returns False on an I/O error. It does not create missing parent directories, so create the destination folder first. WebDriver API

6. Viewport versus full-page screenshots

The standard WebDriver screenshot call in the script captures the current window. It is a good fit for repeatable viewport comparisons, but it should not be described as a portable full-page capture. If you need a full document image, verify support for your chosen browser and Selenium version; Firefox exposes separate full-document screenshot methods, while cross-browser behavior can differ. Another option is to capture the page in viewport-sized sections and stitch them, but that introduces risks around sticky elements, lazy loading, layout changes, and seams.

For full-page captures, decide whether the page should be scrolled to trigger lazy-loaded content, whether sticky headers should appear once or in every segment, and whether the document can change while scrolling. Keep those choices consistent across the batch.

7. cURL example for a simple sequential batch

cURL can fetch screenshot files from URLs, but it does not run a browser or save images into domain folders automatically. This shell example extracts each hostname with Python and places the returned response in that hostname’s directory. It is useful when the screenshot endpoint accepts the requested page URL.

mkdir -p screenshots
while IFS= read -r url; do
  [ -z "$url" ] && continue
  host=$(python -c 'from urllib.parse import urlsplit; import sys; print((urlsplit(sys.argv[1]).hostname or "unknown-host").lower())' "$url")
  safe_host=$(printf '%s' "$host" | tr -c 'A-Za-z0-9.-' '_')
  mkdir -p "screenshots/$safe_host"
  name=$(printf '%s' "$url" | sha256sum | cut -c1-12)
  curl --fail --show-error --location --max-time 90 \
    --get 'https://api.screenshotneo.com/v1/shot' \
    --data-urlencode 'access_key=YOUR_API_KEY' \
    --data-urlencode "url=$url" \
    -o "screenshots/$safe_host/$name.webp" || echo "Failed: $url" >&2
done < urls.txt

This cURL method uses ScreenshotNeo as the capture engine and is not Selenium. The Selenium workflow above is the do-it-yourself browser automation path. The API example uses the documented ScreenshotNeo endpoint and parameters; see ScreenshotNeo documentation for current API usage.

8. Troubleshooting

Symptom Likely cause Fix
InvalidArgumentException or navigation rejects a URL The URL is missing a scheme or is malformed. Validate inputs and provide complete https:// or http:// URLs.
Screenshot shows a loading shell or missing data The load event fired before client-side work completed. Wait for a page-specific content selector or a meaningful state; do not assume readyState guarantees app readiness.
Timeout on one URL The host is slow, unreachable, or keeps loading resources. Set a page-load timeout, record the failed URL, and continue. Review whether a different page-load strategy plus explicit wait suits the site.
Screenshot method returns False Local file I/O failed, often due to permissions, an invalid path, or a missing directory. Create parent directories, use an explicit writable path ending in .png, and check disk space and permissions.
Images are missing Lazy loading, delayed assets, blocked resources, or capture before images are ready. Scroll relevant content into view and wait for the required images or page state. There is no universal selector or readiness condition for every site.
Driver/browser startup error Browser is absent, incompatible, or unavailable in the execution environment. Install a supported browser, check Selenium/browser compatibility, and inspect driver startup logs. Selenium’s troubleshooting guide notes that some errors originate in the underlying browser driver.
Several pages share one image Filename generation used only the hostname. Include a sequence number or stable URL digest as in the script, and retain the URL in a manifest.
Unexpected grouping across subdomains The desired key is registrable domain rather than exact hostname. Choose the grouping rule and use public-suffix-aware parsing for registrable domains.
Browser remains open after an error Cleanup was not placed in a finally block. Keep driver.quit() in finally. Selenium documents that quit() closes browser windows and the driver executable.

For Selenium synchronization and driver issues, consult the official troubleshooting guide.

9. Performance, reliability, and cost

  • Reuse one session for ordinary batches. Starting a browser is overhead; navigating sequentially in one session is simple and avoids repeatedly launching a browser. It also carries cookies, cache, and other session state from one site to the next.
  • Isolate when state matters. Use a fresh browser session or a deliberate profile strategy when pages must not share authentication or cookies. This costs more startup time and system resources, but reduces state leakage between captures.
  • Bound slow work. A page-load timeout keeps one navigation from waiting without a defined ceiling. The script catches page failures per URL, then proceeds. For very large batches, checkpoint the manifest periodically so an interrupted process does not lose all progress.
  • Control concurrency. Sequential capture is easier on the host machine and target sites. Parallel browsers can increase throughput, but consume more CPU and memory and can create more traffic; use a modest, deliberate limit and respect each site’s access rules.
  • Make reruns safe. Stable URL-derived names help identify the same input across runs. Add a run identifier or separate run directory if prior captures must be preserved rather than overwritten.
  • Budget storage. PNG file sizes vary with page dimensions and content. Estimate from a representative sample in your own workload and remove or archive captures according to your retention needs.
  • Account for infrastructure. Selenium itself is open-source software, but running a browser still uses compute, storage, and network resources. Managed browser infrastructure or remote Selenium adds provider-specific costs; the dossier does not establish a universal price.

Browser screenshots may contain account details or other sensitive page content. Store them with appropriate access controls, avoid capturing pages you are not authorized to access, and set a retention policy for the output directory.

10. Or skip the browser setup

For a URL list where you need image files rather than browser control, ScreenshotNeo can return a screenshot from one GET request. Its API can also be called once per URL from a batch script. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
  • An MCP server lets AI agents use screenshot tools, including take_screenshot, get_page_info, and capture_pdf.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Sign up free for 1,000 screenshots a month with no card.

11. Frequently asked questions

Does the sample save one screenshot per URL?

Yes. Each valid URL is processed individually. Successful captures are PNG files in a hostname folder; failed entries are included in manifest.json.

Will two subdomains go in the same folder?

No. The example groups by exact hostname, so docs.example.com and www.example.com are separate folders.

Can I use another browser?

Yes, Selenium supports multiple browser implementations. Choose the corresponding WebDriver and validate screenshot and full-page behavior for that browser and version.

Why does the script use a manifest as well as descriptive filenames?

A manifest preserves the original URL and failure details even when filesystem-safe names are shortened or normalized.

Does a screenshot contain the whole page?

The standard method shown captures the current window. Use a verified browser-specific full-document method or another capture approach when the whole page is required.