How to Bulk Save Website URLs as Full-Page PNG Evidence
Capture a URL list as full-page PNGs with Playwright, unique filenames, and a manifest that links every image to its source page.
To bulk save website URLs as full-page PNG evidence, read a stable list of URLs, open each one in a browser, wait for the content you need, and save a full-page screenshot to a unique filename. Keep a manifest that maps each original URL to its image and records capture time and browser settings. Then review the images and manifest for redirects, access prompts, errors, and missing content.
The example below uses Python and Playwright. Full-page capture includes the page’s scrollable content as rendered by the browser; it does not guarantee that delayed, interactive, authenticated, or third-party content has loaded. Playwright documents the full-page option and PNG file output in its screenshots guide and Page API.
1. Prepare a traceable URL list
Put one URL per line in urls.txt. Keep the input file unchanged during the run. The script assigns each input row a stable index, so duplicate URLs still get separate output files.
https://example.com/
https://example.org/
https://example.com/
If you use CSV, adapt the input-reading portion to select the URL column while retaining the original row number or identifier. Do not deduplicate if every submitted row must have its own evidence record.
2. Install Playwright and its browser
python -m venv .venv
# macOS or Linux:
. .venv/bin/activate
# Windows PowerShell:
# .venv\\Scripts\\Activate.ps1
python -m pip install playwright
python -m playwright install chromium
Install the browser in the same environment that runs the script. The script below uses Chromium. Playwright also supports Firefox and WebKit; choose one browser and keep that choice consistent when repeatability matters.
3. Run a sequential full-page capture
Save this as bulk_capture.py. It writes one PNG per input line, a CSV manifest, and an error entry for each failed capture. Navigation uses domcontentloaded as a starting point, followed by a short configurable settling delay. Neither that event nor a fixed delay proves that every page’s content is ready; change the readiness logic for the sites you capture and inspect the resulting images.
from datetime import datetime, timezone
from pathlib import Path
import csv
import re
from playwright.sync_api import sync_playwright
INPUT_FILE = Path("urls.txt")
OUTPUT_DIR = Path("evidence")
MANIFEST_FILE = OUTPUT_DIR / "manifest.csv"
NAVIGATION_TIMEOUT_MS = 45_000
SETTLE_MS = 1_000
VIEWPORT = {"width": 1440, "height": 900}
def safe_host(url: str) -> str:
"""Make a readable filename component from a URL's host."""
host = re.sub(r"^https?://", "", url.strip(), flags=re.I).split("/", 1)[0]
host = re.sub(r"[^A-Za-z0-9.-]+", "_", host).strip("._-")
return (host or "page")[:100]
urls = [line.strip() for line in INPUT_FILE.read_text(encoding="utf-8").splitlines()
if line.strip() and not line.lstrip().startswith("#")]
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
with MANIFEST_FILE.open("w", newline="", encoding="utf-8") as manifest_file:
fields = ["index", "url", "final_url", "filename", "captured_at_utc",
"status", "http_status", "browser", "viewport", "error"]
writer = csv.DictWriter(manifest_file, fieldnames=fields)
writer.writeheader()
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
page = context.new_page()
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
for index, url in enumerate(urls, start=1):
filename = f"{index:04d}_{safe_host(url)}.png"
output_path = OUTPUT_DIR / filename
row = {
"index": index,
"url": url,
"final_url": "",
"filename": "",
"captured_at_utc": datetime.now(timezone.utc).isoformat(),
"status": "error",
"http_status": "",
"browser": "Chromium (Playwright)",
"viewport": f"{VIEWPORT['width']}x{VIEWPORT['height']}, scale=1",
"error": "",
}
try:
response = page.goto(url, wait_until="domcontentloaded")
page.wait_for_timeout(SETTLE_MS)
page.screenshot(path=str(output_path), full_page=True, type="png")
row["final_url"] = page.url
row["filename"] = filename
row["status"] = "captured"
row["http_status"] = response.status if response else "no main response"
except Exception as exc:
row["final_url"] = page.url
row["error"] = f"{type(exc).__name__}: {exc}"
# A failed capture should not leave a stale or partial-looking artifact.
if output_path.exists():
output_path.unlink()
finally:
writer.writerow(row)
manifest_file.flush()
context.close()
browser.close()
print(f"Processed {len(urls)} URL rows. Review {MANIFEST_FILE} and {OUTPUT_DIR}.")
Run it with python bulk_capture.py. A row marked captured means the navigation and image write completed; it does not certify that the page was complete, accessible, authentic, or legally admissible. Review the HTTP status, final URL, image, and any page-specific conditions before using it as evidence.
4. Tune readiness for the pages you capture
There is no universal wait condition that guarantees every site has finished rendering. Choose a condition based on the content you need:
- Mostly static pages: use
domcontentloaded, then inspect whether the required content appears. - Pages whose main resources matter: try
wait_until="load", but account for pages with long-running or unreliable resources. - Known content marker: after navigation, wait for a site-specific locator, such as
page.locator("main article").wait_for(). A visible marker can help, but it does not prove every image or embed is ready. - Lazy-loaded images: full-page capture does not promise that every site will load all lazy content. If the page loads content only as it scrolls, use a site-specific scrolling/readiness routine, then inspect the output. Avoid assuming that one scroll strategy works for every layout.
- Interactive content: dismiss or accept a dialog, click a tab, or submit a search only when that interaction is part of the intended record. Record the action in the manifest or an accompanying notes file.
- Network-idle state: pages with analytics, polling, or open connections may never become idle. Playwright’s Page API marks
networkidleas discouraged for testing; prefer a concrete readiness condition where possible.
For a per-site readiness check, replace the settling delay with an explicit wait, for example:
page.goto(url, wait_until="domcontentloaded")
page.locator("main article").wait_for(state="visible", timeout=15_000)
page.screenshot(path=str(output_path), full_page=True, type="png")
5. Keep every image tied to its source
The CSV manifest is a practical traceability aid, not a format mandated by Playwright. For each input row, it records the original URL, final URL after redirects, output filename, capture time, browser, viewport, HTTP status when available, and error text. Add fields that matter to your review, such as a case ID, operator, authentication profile name, locale, or capture notes. Do not put passwords, session cookies, or other secrets in the manifest.
Use a new output directory for each run or include a run identifier in the directory name. Otherwise, reruns can overwrite earlier images with the same index and hostname. Preserve the input list and manifest with the image set. If integrity matters, generate file hashes and store them in a separate record using a process suitable for your organization; the Playwright sources do not prescribe chain-of-custody procedures.
6. Choose screenshot settings deliberately
| Setting | What it changes | Practical guidance |
|---|---|---|
full_page=True |
Captures the full scrollable page instead of just the visible viewport. | Use for full-page records; inspect very long pages for missing or awkwardly rendered content. |
path="...png" / type="png" |
Writes a PNG file. Screenshot format can be inferred from the extension. | Use unique paths; PNG is lossless and can be large. |
viewport |
Sets the browser’s CSS viewport dimensions. | Fix it for consistent captures. Responsive layouts can change substantially with viewport width. |
device_scale_factor |
Sets the device pixel ratio in the browser context. | Record it. A higher scale can produce larger images and affect rendering. |
browser.new_context() |
Controls isolated browser state, including viewport, locale, timezone, and other context settings. | Set required values explicitly when the page output depends on them. Use a separate context for separate identities. |
| Browser choice | Chromium, Firefox, and WebKit can render differently. | For repeat captures, keep the browser and Playwright version consistent and record them. |
Playwright documents that screenshot rendering can vary with the operating system, browser version and settings, hardware, and other environmental details. Record the relevant environment and validate comparisons instead of assuming pixel identity across machines. See the visual comparisons documentation.
7. cURL, Python, and Node.js alternatives
Browser automation is useful when each URL needs custom interaction or local browser control. For a hosted one-request-per-URL workflow, these examples use ScreenshotNeo’s screenshot API. They save the returned response bytes as a file; check the response and output before treating a file as a valid capture. See the ScreenshotNeo API documentation for parameters and response details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
To bulk capture with an API, wrap the request in a loop over the input rows, use a unique filename per row, and write a manifest just as in the browser example. Limit concurrency according to your plan and the service’s documented guidance, handle non-success responses, and retain the URL-to-file mapping. The sample API calls above use the exact target URL supplied for this article’s product examples; adapt that target to each row in your input.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns an image or PDF; for this PNG evidence workflow, set the supported output format and capture options described in the API docs, and wrap the request for each URL in your bulk loop.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Cookie and consent banners, newsletter popups, and chat widgets can be removed before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Navigation timeout | The server is slow, unreachable, or waiting on resources; the timeout may be too short. | Check the URL and network access, raise the navigation timeout for that site, and choose a meaningful readiness signal. Do not simply treat a timeout as a successful capture. |
| Invalid URL or navigation error | Malformed input, unsupported scheme, SSL issue, DNS failure, or unreachable host. | Validate the URL, include its scheme (usually https://), and record the exact error. Investigate certificate or network problems without suppressing them silently. |
| Image exists but content is missing | The screenshot call ran before delayed content appeared, lazy resources did not load, or required interaction/authentication was absent. | Add a site-specific wait or interaction, then inspect the result. Record authentication context without exposing credentials. |
| Only part of page appears | The site’s layout or browser rendering may constrain scroll height, or content may load on scroll. | Inspect page behavior, trigger the required loading behavior, and verify the full-page result visually. A full-page option captures the browser’s scrollable page; it is not a completeness guarantee. |
| Access denied, CAPTCHA, or sign-in page | The site blocks automation or requires an authorized session. | Record the page that actually appeared. Use an authorized authentication flow if appropriate; do not claim the intended page was captured when it was not. |
| Duplicate output names | Filenames are based only on domain, or a previous run used the same path. | Include the stable input index or ID in each filename and use a separate run directory. |
| Browser executable missing | Playwright’s Python package is installed but its browser was not installed in this environment. | Run python -m playwright install chromium in the same environment used to run the script. |
| Different images across runs | Page content, browser version, operating system, viewport, device scale, fonts, locale, or timing changed. | Record and standardize relevant settings, and treat genuinely dynamic content as a source of variation. |
| Very large PNG files | Full-page images can be very tall and PNG preserves image detail. | Capture only required pages, use a consistent viewport and scale, and store or transfer files with adequate capacity. If the evidence process requires PNG, do not silently switch to a lossy format. |
Performance, reliability, and cost
- Sequential versus parallel runs: the sample is sequential to make failures and outputs easy to associate with input rows. Parallel pages can reduce elapsed time, but increase CPU, memory, network load, and the chance that a site throttles or blocks requests. Add bounded concurrency only after measuring your own workload and preserving per-row error handling.
- Reuse the browser: the script launches one browser and reuses its page for the URL list. This avoids launching a new browser for every row. Separate contexts when URLs require isolated sessions or different settings.
- Long pages: full-page screenshots can consume substantial memory and disk space. Capture only what the record requires, monitor available storage, and consider splitting work into batches.
- Retries: transient network failures may justify a small, bounded retry with a recorded attempt count. Do not retry forever or overwrite a prior successful artifact without recording the event.
- Cost: local Playwright has no per-screenshot service charge in this workflow, but uses your compute, bandwidth, storage, and maintenance time. A hosted API trades local browser setup for service plan limits and charges. ScreenshotNeo’s stated plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Check the product docs for the current request options before building a production workflow.
Evidence limits and review checklist
A PNG is a record of rendered pixels at a particular time and browser configuration. It does not, by itself, establish who controlled the page, whether the content was authentic, or whether the image meets a legal evidence standard. Requirements for admissibility and chain of custody depend on context; consult the appropriate legal or organizational process when those questions matter.
- Every input row has an output filename or an explicit error in the manifest.
- Each image was opened and checked for the intended page, relevant content, and obvious access prompts.
- Redirects and HTTP statuses were reviewed, not assumed to indicate success.
- Capture time, browser, viewport, scale, and relevant interactions are recorded.
- The input URL list and manifest are retained with the images, and the output directory is protected from accidental overwrite.
- Any partial, blocked, or failed capture is labeled honestly.
FAQ
Does full-page mean every element below the fold is included?
It means Playwright captures the page’s full scrollable area as rendered. Delayed content, lazy loading, embedded frames, and interaction-dependent sections may need extra handling and visual review.
Can I capture the same URL more than once?
Yes. The sample gives each input row its own indexed filename, so repeated URLs remain separate records.
Is a screenshot alone proof that a page was authentic?
No. It records rendered pixels. Keep the surrounding context required by your review process; the Playwright documentation does not certify authenticity or legal admissibility.
Why keep a manifest if filenames already contain the domain?
Domains can repeat, URLs can redirect, and filenames omit query strings and other context. The manifest preserves the direct mapping from each input row to its captured file.


