Capture Screenshots of a List of Webpages from a CSV with Python
Use Python and Playwright to capture each URL in a CSV, save uniquely named screenshots, and track failures. Includes full-page options and an API alternative.
Read the CSV with Python’s built-in csv.DictReader, launch Playwright once, and visit each nonblank URL in turn. Save each screenshot with a row number and sanitized hostname, and record navigation errors so one unavailable page does not stop the batch. This guide uses Playwright’s synchronous Python API and supports viewport or full-page captures.
1. Prepare Python and Playwright
Use Python’s standard-library csv module for input; no CSV package is needed. Playwright provides the browser automation and screenshot functions. Install the Python package and its browser binaries:
python -m pip install playwright
python -m playwright install chromium
The examples below use Chromium. Playwright for Python also supports Firefox and WebKit; install the corresponding browser and change p.chromium to p.firefox or p.webkit if required. See the Playwright Python setup documentation and Python csv module documentation.
2. Format the input CSV
Put one URL in a column named url. A header row is required by this example:
url
https://example.com
https://www.python.org
https://playwright.dev
DictReader maps each row to a dictionary using the header names. The code opens the file with UTF-8 encoding and newline="", as recommended for CSV input. If the export uses another encoding, change encoding to match it. If it uses a different delimiter, pass that delimiter to csv.DictReader and validate the headers.
3. Capture every URL with Python
Save this as capture_csv.py next to urls.csv, then run python capture_csv.py. It writes images to screenshots/ and a per-row outcome log to capture-results.csv.
import csv
import re
from pathlib import Path
from urllib.parse import urlsplit
from playwright.sync_api import sync_playwright
INPUT_CSV = Path("urls.csv")
OUTPUT_DIR = Path("screenshots")
RESULTS_CSV = Path("capture-results.csv")
URL_COLUMN = "url"
FULL_PAGE = True
NAVIGATION_TIMEOUT_MS = 30_000
def safe_hostname(url: str) -> str:
"""Return a filesystem-safe hostname, or 'page' when parsing fails."""
try:
hostname = urlsplit(url).hostname or "page"
except ValueError:
hostname = "page"
safe = re.sub(r"[^A-Za-z0-9.-]+", "_", hostname).strip("._-")
return safe or "page"
def main() -> None:
if not INPUT_CSV.is_file():
raise FileNotFoundError(f"Input CSV not found: {INPUT_CSV}")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
outcomes = []
with INPUT_CSV.open("r", encoding="utf-8", newline="") as csv_file:
reader = csv.DictReader(csv_file)
headers = reader.fieldnames or []
if URL_COLUMN not in headers:
raise ValueError(
f"CSV must have a '{URL_COLUMN}' column; found headers: {headers}"
)
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
page = browser.new_page(viewport={"width": 1440, "height": 900})
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
for row_number, row in enumerate(reader, start=2):
url = (row.get(URL_COLUMN) or "").strip()
if not url:
outcomes.append({
"row": row_number, "url": "", "status": "skipped",
"file": "", "error": "blank URL",
})
continue
filename = f"{row_number:04d}-{safe_hostname(url)}.png"
output_path = OUTPUT_DIR / filename
try:
response = page.goto(url, wait_until="load")
# A 4xx/5xx response still rendered a page; save it and
# make the HTTP status visible in the results file.
page.screenshot(path=str(output_path), full_page=FULL_PAGE)
status = "ok" if response is None else f"http_{response.status}"
outcomes.append({
"row": row_number, "url": url, "status": status,
"file": str(output_path), "error": "",
})
except Exception as exc:
outcomes.append({
"row": row_number, "url": url, "status": "failed",
"file": "", "error": f"{type(exc).__name__}: {exc}",
})
# Discard page state after navigation/capture errors.
try:
page.goto("about:blank", wait_until="load", timeout=5_000)
except Exception:
pass
finally:
browser.close()
with RESULTS_CSV.open("w", encoding="utf-8", newline="") as results_file:
fields = ["row", "url", "status", "file", "error"]
writer = csv.DictWriter(results_file, fieldnames=fields)
writer.writeheader()
writer.writerows(outcomes)
captured = sum(item["status"] in ("ok",) or item["status"].startswith("http_")
for item in outcomes)
failed = sum(item["status"] == "failed" for item in outcomes)
skipped = sum(item["status"] == "skipped" for item in outcomes)
print(f"Captured: {captured}; failed: {failed}; skipped: {skipped}")
print(f"Images: {OUTPUT_DIR.resolve()}")
print(f"Results: {RESULTS_CSV.resolve()}")
if __name__ == "__main__":
main()
The first CSV record is row 2 in the file because row 1 is the header. Numbering plus a sanitized hostname keeps filenames unique even when a URL has a long query string or two URLs share a host. The script continues after individual navigation or screenshot exceptions. HTTP error pages are captured and labeled with their status; decide whether those images are useful for your workflow.
4. Choose what the screenshot contains
| Need | Playwright choice | Notes |
|---|---|---|
| Visible browser viewport | full_page=False |
Captures the current viewport size, set here to 1440 × 900 CSS pixels. |
| Entire scrollable document | full_page=True |
Captures beyond the viewport. Very long pages can create large images or take longer. |
| One component or region | page.locator("main").screenshot(path="main.png") |
Use a locator screenshot for one matched element; a detached or missing element can cause the capture to fail. |
For a viewport capture, change FULL_PAGE to False. For a particular element, replace the page screenshot call with a locator screenshot, for example:
page.locator("main").screenshot(path=str(output_path), animations="disabled")
Use a selector that exists on every target page, or handle missing matches as per-row failures. Playwright’s screenshot APIs also allow image format and scale choices; JPEG supports a quality setting, while PNG does not use JPEG quality. The browser viewport is expressed in CSS pixels; device scale affects the output pixel dimensions. See the Playwright screenshot documentation.
5. Handle dynamic pages and CSV variations
Wait for the right page state
The sample waits for the page’s load event. Some sites continue rendering after that event, while pages with long-lived network requests may never become network-idle. If a known element indicates readiness, wait for it after navigation:
page.goto(url, wait_until="domcontentloaded")
page.locator("main").wait_for(state="visible", timeout=10_000)
page.screenshot(path=str(output_path), full_page=True)
For a fixed delay, use page.wait_for_timeout(1_000) sparingly; a selector-based wait is usually more meaningful than guessing a delay. Lazy-loaded images may appear only after scrolling. For pages that rely on scrolling, scroll in increments and allow content to load before requesting a full-page screenshot.
Validate URLs and input shape
Blank cells are skipped and logged. If your file has no header, provide explicit field names to DictReader or use csv.reader and select the relevant column. Spreadsheet exports can use alternate delimiters or quoting. Inspect the header and a few rows if URLs seem to land in the wrong field. URLs generally need a scheme such as https://; malformed or unsupported URLs will fail navigation and appear in the results log.
Authentication and access restrictions
This script opens a fresh browser context without your normal browser profile. Pages requiring login will not automatically inherit your cookies. Sites can also return bot checks, consent screens, or access-denied pages. For authorized sites, configure a browser context with the required cookies or headers; do not assume a successful navigation means the intended content loaded. Check the returned HTTP status and inspect representative captures.
6. Run larger batches reliably
- Keep the browser open for the batch. Launching one browser and reusing a page avoids repeated startup overhead. The example processes URLs sequentially to keep memory use predictable.
- Set a timeout. The sample uses a 30-second navigation timeout. Raise it for slow destinations or lower it to move past repeatedly stalled hosts.
- Retry selectively. Add a bounded retry for transient navigation failures if your job requires it. Record attempts and error details; retries cannot fix a permanent 404, login wall, or blocked request.
- Use concurrency carefully. Parallel pages can increase throughput but also memory, CPU, and load on target websites. Start with a small number of workers, respect site policies, and keep results associated with their input rows.
- Make reruns safe. Deterministic row-and-host filenames overwrite the same output on rerun. If preserving prior runs matters, write each run to a timestamped directory.
- Plan storage. Full-page PNGs can be large. Use viewport capture, JPEG where lossy output is acceptable, or post-process images when storage is a constraint.
7. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
Executable doesn't exist or browser launch error |
The Playwright package is installed but its browser binary is not. | Run python -m playwright install chromium in the same environment. |
| CSV must have a ‘url’ column | Header is missing, spelled differently, or parsed with the wrong delimiter or encoding. | Check the first row, rename the column, or configure DictReader for the file’s actual format. |
net::ERR_NAME_NOT_RESOLVED or connection failure |
Invalid hostname, DNS issue, network restriction, or unavailable site. | Verify the URL from the same machine and network; the row’s error is retained in the results CSV. |
| Navigation timeout | The site is slow, keeps connections open, or blocks automation. | Adjust the navigation timeout, wait for a specific readiness condition, and inspect the page response. A longer timeout will not bypass access controls. |
| Screenshot is blank or shows an interstitial | The site rendered an empty page, bot check, consent screen, or error page. | Inspect the screenshot and status. Handle required consent or authorized authentication explicitly; do not treat an interstitial as the target content. |
| Full-page image misses late content | Content or images load only after scroll or interaction. | Wait for a known selector and, where appropriate, scroll through the page before capture. |
| Element screenshot times out | The selector did not match a visible element or the element was detached during rendering. | Wait for the locator to become visible, confirm the selector, and retry that row if the page is transient. |
| Duplicate or odd filenames | Raw URLs are unsuitable as filenames, or multiple rows target the same host. | Keep the row number in the filename as shown; use a different naming scheme if stable IDs are available. |
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API accepts a URL in one GET request and returns an image or PDF. See the ScreenshotNeo API documentation. For a CSV batch, loop over your rows and save each response using the same row-based filename strategy:
import csv
from pathlib import Path
import requests
OUTPUT_DIR = Path("screenshots")
OUTPUT_DIR.mkdir(exist_ok=True)
with open("urls.csv", encoding="utf-8", newline="") as source:
for row_number, row in enumerate(csv.DictReader(source), start=2):
url = (row.get("url") or "").strip()
if not url:
continue
response = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": url},
timeout=90,
)
response.raise_for_status()
(OUTPUT_DIR / f"{row_number:04d}.webp").write_bytes(response.content)
One-request examples for the same endpoint:
# cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
# Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
// Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await Bun.write('shot.webp', res);
Use your API key in place of YOUR_API_KEY. The API also supports full-page capture, element selectors, device presets, custom waits, and other capture options. Cookie banners, newsletter popups, and chat widgets are removed before capture, and each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots.
Start with 1,000 free screenshots a month; no card is required.
9. FAQ
Can I capture only selected CSV rows?
Yes. Add a condition in the loop before navigation, such as checking another CSV column for a value like enabled.
Does this script use my logged-in Chrome session?
No. It launches a separate Playwright browser context. Supply authorized cookies or authentication settings explicitly if the pages require them.
Can I save screenshots as JPEG?
Yes. Set the screenshot type to JPEG and choose a quality appropriate for your use; use PNG when lossless output or transparency is needed.
Will every URL produce the same page as a human browser?
Not necessarily. Rendering can vary with authentication, location, timing, consent, bot protection, and dynamic content. Review representative outputs for your target sites.


