How to Bulk Screenshot Websites and Skip Duplicate URLs
Build a reliable bulk screenshot workflow: normalize URLs conservatively, skip duplicates, capture with Playwright, and track failures per URL.
To bulk screenshot websites without capturing duplicate URLs, keep each original URL, derive a conservative comparison key, and capture only the first URL for each key. For a controlled job, Playwright can visit each unique URL and save a screenshot plus a manifest record; for recurring batches, a managed bulk API can accept URL lists and handle capture infrastructure.
The key is not to equate URLs just because they look similar. Lowercase only the scheme and host, apply the URI normalization rules described below, and preserve query strings unless you know the target site treats the alternatives as equivalent. For fragment URLs, choose whether fragments represent distinct browser views in your use case.
1. Decide what counts as a duplicate
Keep the original input string for logs and filenames. Create a separate key for comparison. URI syntax treats scheme and host as case-insensitive, but other components are generally case-sensitive unless a scheme says otherwise. Safe generic normalization includes lowercasing scheme and host, normalizing percent-escape hex digits, decoding percent-encoded unreserved characters, and removing dot segments where applicable. See RFC 3986.
| URL part | Default treatment | Reason |
|---|---|---|
| Scheme and host | Lowercase in the comparison key | They are case-insensitive in generic URI syntax. |
| Path | Preserve case; normalize only defined equivalents | Paths can be case-sensitive on the server. |
| Percent escapes | Uppercase hex digits and decode unreserved characters only | Reserved characters can change delimiter meaning if decoded. |
| Query | Preserve order and values by default | Parameters can identify different content; sorting or dropping values is site-specific. |
| Fragment | Drop for ordinary server-rendered documents; retain for fragment-driven app views | Fragments are not sent in the retrieval request, but browser apps may use them as client-side state. |
RFC 3986 distinguishes reserved characters because they can act as delimiters. Do not decode them or casually rewrite queries. Google’s canonicalization guidance also illustrates why content-level duplicates are not safely inferred from string comparison alone: protocol, regional, device, and filtering variants may resolve to related content, but establishing that requires site or content knowledge.
Exact duplicate strings are always safe to skip. The code below uses Python’s URI parser and a deliberately limited key: scheme and host are lowercased, fragments are dropped for the default server-rendered use case, percent-encoded unreserved characters are decoded, and reserved escapes retain their meaning. If your app uses fragments to select views, change the key function to retain the fragment.
2. Prepare a URL list and remove duplicates
Put one absolute HTTP or HTTPS URL on each line in urls.txt. This complete Python script validates input, deduplicates conservatively, captures sequentially with Playwright, and writes a JSON Lines manifest. It preserves the input URL and records the final URL, output path, status, and error for each unique URL.
from pathlib import Path
from urllib.parse import urlsplit, urlunsplit
import json
import re
import hashlib
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
INPUT = Path("urls.txt")
OUT = Path("screenshots")
MANIFEST = Path("manifest.jsonl")
VIEWPORT = {"width": 1440, "height": 900}
NAVIGATION_TIMEOUT_MS = 30_000
UNRESERVED = set("ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-._~")
PERCENT_ESCAPE = re.compile(r"%([0-9a-fA-F]{2})")
def normalize_percent_escapes(value: str) -> str:
def replace(match):
byte = int(match.group(1), 16)
char = chr(byte)
if char in UNRESERVED:
return char
return "%" + match.group(1).upper()
return PERCENT_ESCAPE.sub(replace, value)
def comparison_key(raw: str, keep_fragment: bool = False) -> str:
parts = urlsplit(raw)
scheme = parts.scheme.lower()
# Preserve user info, port, path, query, and their meanings. Host alone is case-insensitive.
netloc = parts.netloc
if parts.hostname:
host = parts.hostname.lower()
if ":" in host and not host.startswith("["):
host = f"[{host}]" # IPv6 literal
userinfo = ""
if "@" in netloc:
userinfo = netloc.rsplit("@", 1)[0] + "@"
port = f":{parts.port}" if parts.port is not None else ""
netloc = userinfo + host + port
path = normalize_percent_escapes(parts.path)
query = normalize_percent_escapes(parts.query)
fragment = normalize_percent_escapes(parts.fragment) if keep_fragment else ""
return urlunsplit((scheme, netloc, path, query, fragment))
def valid_http_url(raw: str) -> bool:
try:
parts = urlsplit(raw)
return parts.scheme.lower() in {"http", "https"} and bool(parts.hostname)
except ValueError:
return False
OUT.mkdir(parents=True, exist_ok=True)
seen = set()
unique_urls = []
invalid = []
for line_number, line in enumerate(INPUT.read_text(encoding="utf-8").splitlines(), start=1):
raw = line.strip()
if not raw or raw.startswith("#"):
continue
if not valid_http_url(raw):
invalid.append({"line": line_number, "input_url": raw, "status": "invalid_url"})
continue
key = comparison_key(raw, keep_fragment=False)
if key in seen:
continue
seen.add(key)
unique_urls.append(raw)
with MANIFEST.open("w", encoding="utf-8") as manifest:
for record in invalid:
manifest.write(json.dumps(record, ensure_ascii=False) + "\n")
with sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
context = browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
page = context.new_page()
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
for raw in unique_urls:
# Hash the original URL for a stable, filesystem-safe filename.
name = hashlib.sha256(raw.encode("utf-8")).hexdigest()[:20] + ".png"
output_path = OUT / name
record = {"input_url": raw, "output": str(output_path)}
try:
response = page.goto(raw, wait_until="load")
# A missing response can occur for nonstandard navigations; record it, but still attempt capture.
page.screenshot(path=str(output_path), full_page=True, animations="disabled")
record.update({
"status": "ok",
"final_url": page.url,
"http_status": response.status if response else None,
})
except PlaywrightTimeoutError as error:
record.update({"status": "timeout", "error": str(error), "final_url": page.url})
except Exception as error:
record.update({"status": "error", "error": str(error), "final_url": page.url})
manifest.write(json.dumps(record, ensure_ascii=False) + "\n")
manifest.flush()
context.close()
browser.close()
print(f"Captured {len(unique_urls)} unique URLs; see {MANIFEST} and {OUT}/")
Install the dependencies and run it:
python -m pip install playwright
python -m playwright install chromium
python bulk_capture.py
For reproducibility, store the browser version, operating system, viewport, device scale factor, and relevant capture settings alongside the manifest. Playwright notes that rendering can vary with OS, browser version, settings, hardware, power source, and headless mode; see its visual comparison guidance.
3. Capture URLs with Playwright
The script above is intentionally sequential and uses wait_until="load". Choose readiness based on the site: domcontentloaded is quicker but may miss later assets; load waits for load events; an explicit selector or application-ready signal can be more reliable for dynamic sites. Avoid assuming that a fixed delay guarantees the page is ready.
Full-page versus viewport capture
full_page=True captures the full document, including content below the fold, and may produce very tall images. Use full_page=False for a viewport screenshot. Lazy-loaded images may not load until scrolled into view; if they matter, scroll through the page before capturing, then wait for images to settle. Playwright documents full-page screenshots and image buffers in its screenshot guide.
Bound concurrency and isolate failures
Sequential capture is a good starting point for stability and low resource use. For more throughput, use a small fixed number of browser contexts or workers, each with its own page, and keep per-URL exception handling. Do not share one page across simultaneous navigations. Increase workers gradually while watching memory, CPU, target-site rate limits, and failure rates. The sources here establish no throughput benchmark.
For comparisons across runs, keep viewport, browser build, device scale, fonts, locale, timezone, color scheme, and animation handling consistent. A screenshot is the rendered result of the browser environment as well as the URL.
4. Choose local capture or a managed bulk API
Local Playwright gives control over browser settings, network behavior, output names, and your storage destination. It also means you own browser installation, concurrency, retries, job resumption, and durable storage. A managed API can reduce that operational work. AddScreenshots’ Swagger documentation describes asynchronous bulk capture for multiple URLs, sitemaps, URL-list files, or whole domains, with images stored in a cloud repository such as S3 or Azure Blob Storage. Verify its current endpoints, limits, integrations, security, terms, and pricing directly before relying on it; the cited documentation does not establish comparative cost or throughput. See the AddScreenshots API documentation.
Compare options against the actual job: required browser control, whether input is a list or sitemap, retry and failure visibility, naming and storage needs, maintenance ownership, and current cost. A sitemap can be a useful source of candidate URLs, but it does not guarantee that all entries are valid or distinct for your capture purpose.
Or skip the browser setup
For a managed capture call, ScreenshotNeo accepts a URL and returns an image or PDF. Its batch capture supports up to 100 URLs per call. This example shows one request; send your unique URL set through the batch workflow documented in the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots per month are free with no card, with paid plans starting at $5 for 3,000. Its responses include page-verdict and billing headers, so inspect each result when processing a batch. Sign up for 1,000 free screenshots a month, with no card.
5. Handle output naming, manifests, and reruns
Use a stable filename derived from the original URL or an explicit input ID. Avoid using raw URLs as path names: they can contain characters unsuitable for filesystems, be excessively long, or expose query data in filenames. A truncated cryptographic hash is convenient for collision-resistant naming, but keep the original URL in a protected manifest so the output remains traceable.
The JSON Lines manifest in the example allows downstream jobs to identify successes and retry only failures. For resumable jobs, read prior successful records and skip their comparison keys on rerun. Write each record as soon as it completes so an interruption does not erase the job’s progress. Store credentials and authenticated URLs carefully; manifests can contain sensitive query parameters.
6. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Two URLs that look alike produce different captures | Query values, path case, trailing slash, locale, or app state differs | Keep both unless the site owner or application rules establish they are equivalent. Do not normalize queries by assumption. |
| URLs with different capitalization are still duplicated or missed | The key does not lowercase scheme and hostname consistently | Parse the URI and normalize those components only; do not lowercase the path or query generally. |
| Fragment URLs collapse but represent different screens | The app uses a hash route or fragment state | Retain fragments in the comparison key and capture each intended view. |
| Percent-encoded URL behaves differently after normalization | A reserved character was decoded or query syntax rewritten | Decode unreserved characters only and preserve reserved escapes and query structure. |
| Navigation timeout | Slow server, long-lived requests, or a page that never reaches the selected readiness event | Set a realistic timeout, choose an appropriate readiness condition, record the failure, and retry selectively with a limit. |
| Screenshot misses lazy images | Images load only after scrolling into view | Scroll through the document and wait for relevant images before full-page capture. |
| Output differs between runs | Browser, OS, viewport, fonts, animations, or rendering environment changed | Pin and record the environment and settings; disable animations where practical. |
| Many failures after increasing worker count | Local resource exhaustion or target-side rate limiting | Reduce concurrency, add bounded retries with backoff, and respect the target site’s access policies. |
| Invalid URL or parser exception | Missing scheme, malformed authority, or invalid port | Require absolute HTTP/HTTPS URLs, validate before capture, and retain rejected lines in the manifest. |
7. Performance, reliability, and cost
- Performance: Each local capture consumes browser time and memory. Start sequentially, then add bounded workers only if needed. Full-page screenshots and high device scale factors increase image size and processing work.
- Reliability: Isolate errors per URL, flush manifest records, retry only transient failures, and cap retry attempts. A timeout or HTTP error should remain visible in the result record rather than silently disappearing.
- Cost: Local capture has no per-shot API fee established by the cited sources, but uses compute, storage, and maintenance. Managed service pricing and limits change; check the provider’s current terms. No price or throughput comparison is established here.
- Site behavior: Authentication, consent dialogs, redirects, rate limits, bot checks, and client-side rendering differ by target. Use only access and capture patterns permitted for the sites and data involved.
Frequently asked questions
Should I deduplicate by page content instead of URL?
Only if content equivalence is part of the job and you have a policy for redirects, canonical links, and rendered state. String normalization alone cannot establish that two pages show the same content.
Can I screenshot a sitemap?
Yes. A sitemap can supply candidate URLs. Validate entries, decide how fragments and query variants should be treated, then run them through the same capture and manifest workflow.
Can the script capture authenticated pages?
It can be adapted to use a browser context with the required session state or credentials. Keep secrets out of source files and manifests, and capture only pages you are authorized to access.
Why keep duplicate inputs in the logs?
They show which input lines were skipped and make the deduplication policy auditable. Preserve originals even when several inputs map to one comparison key.


