How to Take Bulk Screenshots in Python with a Screenshot API
Capture many URLs with Python using Playwright or a hosted screenshot API. Get runnable code, batching patterns, error handling, and practical guidance for choosing a workflow.

To take screenshots of many URLs in Python, keep the URL list and output tracking in your code, then choose how each page gets rendered: run a browser with Playwright, or send URLs to a hosted screenshot API. Playwright documents screenshots of individual pages; its loop, retry policy, concurrency, filenames, and manifest are application code. For a managed batch workflow, choose an API whose current documentation explicitly describes multi-URL submission and job tracking. [Playwright’s screenshot guide](https://playwright.dev/python/docs/screenshots) covers path, byte-buffer, full-page, and element captures.
This guide shows a complete local Playwright batch script, a controlled asynchronous version, how to approach a documented API batch endpoint, the capture settings that matter, and ways to diagnose failures. There is no universal throughput or cost winner: measure against your pages, limits, and operating requirements.
1. Choose a capture route
| Need | Playwright running in your environment | Hosted screenshot API |
|---|---|---|
| Per-page control | Direct browser access. Set viewport and capture full pages, elements, clips, masks, image formats, and bytes or paths. | Settings depend on the provider. Check its live docs for viewport, format, waits, selectors, and other options. |
| Multiple URLs | Your program loops over URLs and tracks each result. The cited Playwright references show page capture calls, not a built-in bulk queue. | Some providers document a multi-URL batch endpoint and job-progress mechanisms. Confirm details with the provider. |
| Operations | You install and operate browser binaries, workers, retries, and storage. | The provider operates rendering infrastructure; you still need to submit work and handle returned status and outputs. |
| Costs and speed | Depends on infrastructure, page complexity, concurrency, and maintenance. | Depends on plan, quotas, rendering options, and output retention. Compare current terms; no independent benchmark is available here. |
Use Playwright when your code needs direct control over browser behavior or screenshot bytes. Use a hosted API when managed rendering and a documented batch/job interface better fit the workflow. A screenshot endpoint that only accepts one URL per request can still be called in a loop, but it does not become a batch API simply because the caller submits requests concurrently.
2. Install Playwright and prepare a URL list
Install the Python package and a browser. The browser install step matters in fresh CI workers and containers because the Python package alone does not provide a ready-to-launch browser binary.

python -m pip install playwright
python -m playwright install chromium
Put one URL on each line in urls.txt. Keep input separate from the script so a job can be rerun from the same list, and validate the scheme before navigation. Treat URLs as input: if they come from users, restrict allowed destinations to avoid making your capture worker fetch internal services or local network addresses.
https://example.com/
https://playwright.dev/
https://www.python.org/
3. Runnable Python batch capture with Playwright
This synchronous example launches Chromium once, visits each URL in order, writes one PNG per URL, and records successes and failures in a JSON manifest. It uses a conservative navigation wait and a per-navigation timeout. The timeout and wait choice are policy decisions: pages with client-side rendering may need a selector wait or a deliberate additional delay.
from pathlib import Path
from urllib.parse import urlparse
import json
import re
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
INPUT = Path("urls.txt")
OUT = Path("screenshots")
OUT.mkdir(exist_ok=True)
def safe_name(index: int, url: str) -> str:
host = urlparse(url).netloc or "page"
stem = re.sub(r"[^a-zA-Z0-9.-]+", "-", host).strip("-") or "page"
return f"{index:04d}-{stem}.png"
urls = [line.strip() for line in INPUT.read_text().splitlines()
if line.strip() and not line.lstrip().startswith("#")]
results = []
with sync_playwright() as p:
browser = p.chromium.launch()
try:
context = browser.new_context(viewport={"width": 1440, "height": 900})
try:
page = context.new_page()
for index, url in enumerate(urls, start=1):
result = {"url": url}
try:
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
raise ValueError("URL must be an absolute HTTP or HTTPS URL")
response = page.goto(url, wait_until="domcontentloaded", timeout=30_000)
result["http_status"] = response.status if response else None
path = OUT / safe_name(index, url)
page.screenshot(path=str(path), full_page=True, animations="disabled")
result.update({"ok": True, "file": str(path)})
except (PlaywrightTimeoutError, ValueError, Exception) as exc:
result.update({"ok": False, "error": f"{type(exc).__name__}: {exc}"})
results.append(result)
print(json.dumps(result))
finally:
context.close()
finally:
browser.close()
Path("manifest.json").write_text(json.dumps(results, indent=2))
The broad exception clause ensures one failed URL does not stop later URLs, but in a production library you may want to catch specific exceptions and log tracebacks to your job system. A response with an HTTP error status can still render a page; decide whether to keep its screenshot or classify it as failed according to your use case. A navigation timeout also does not prove that the page is entirely blank: capture decisions should distinguish navigation, HTTP status, and image output.
Full-page, element, and byte output
Playwright’s full_page=True captures the full scrollable document; omit it for the current viewport. Capture a single element with a locator. Calling screenshot() without a path returns bytes that can be passed to an image-processing or storage layer.
# Viewport PNG saved to disk
page.screenshot(path="viewport.png")
# Full scrollable page
page.screenshot(path="full-page.png", full_page=True)
# One element
page.locator("main article").screenshot(path="article.png")
# Bytes for upload or post-processing
image_bytes = page.screenshot(type="jpeg", quality=82)
Screenshot options include path, type (png, jpeg, or webp), quality for JPEG/WebP, full_page, clip, scale (css or device), animations, caret, mask, mask_color, omit_background, style, and timeout. JPEG does not support transparency; quality does not apply to PNG. Device scale can produce larger images than CSS-pixel scale. See the [Page screenshot API reference](https://playwright.dev/python/docs/api/class-page#page-screenshot) for exact parameter behavior.
4. Async capture and bounded concurrency
For an asyncio application, use Playwright’s async API. A semaphore limits active page tasks; there is no universal safe concurrency setting. Start low, observe memory and failure rates, then adjust for your machine and targets. This example gives each task a fresh page in one browser context and always closes it.
import asyncio
import json
from pathlib import Path
from playwright.async_api import async_playwright
URLS = ["https://example.com/", "https://playwright.dev/"]
OUT = Path("async-shots")
OUT.mkdir(exist_ok=True)
LIMIT = 3 # Tune for your environment and target sites.
async def main():
semaphore = asyncio.Semaphore(LIMIT)
async with async_playwright() as p:
browser = await p.chromium.launch()
context = await browser.new_context(viewport={"width": 1365, "height": 900})
async def capture(i, url):
async with semaphore:
page = await context.new_page()
try:
response = await page.goto(url, wait_until="domcontentloaded", timeout=30_000)
path = OUT / f"{i:04d}.png"
await page.screenshot(path=str(path), full_page=True)
return {"url": url, "ok": True, "status": response.status if response else None,
"file": str(path)}
except Exception as exc:
return {"url": url, "ok": False, "error": f"{type(exc).__name__}: {exc}"}
finally:
await page.close()
try:
results = await asyncio.gather(*(capture(i, u) for i, u in enumerate(URLS, 1)))
(OUT / "manifest.json").write_text(json.dumps(results, indent=2))
finally:
await context.close()
await browser.close()
asyncio.run(main())
For very large inputs, do not create one task per URL up front. Feed a bounded worker pool from a queue, persist each result as it completes, and resume from unfinished records after a process restart. This keeps queued work manageable and avoids losing the entire batch manifest if the process exits late.
5. Hosted API batches: submission, tracking, and output
A hosted batch API can move URL scheduling and rendering off your own browser workers. The reviewed vendor documentation describes a Python requests.post call for a single screenshot and a separate POST /api/v1/screenshot/batch route that accepts multiple URLs; it also describes a returned batch ID, polling, and server-sent event updates. Those are vendor-documented claims, not independently verified here. The research dossier does not provide the vendor’s base URL, exact request schema, auth header, or response fields, so copying a guessed host or payload would be unsafe. Check the provider’s live docs for those values before using its route.
Implement the integration in these steps:
- Read the provider’s live batch schema and authentication instructions. Keep API credentials in environment or secret configuration, not in committed source.
- Submit a bounded chunk of URLs with shared capture settings. Validate URL schemes and preserve an application-side mapping from each submitted URL to your own record ID.
- Persist the returned batch ID and submission time before polling. This lets a restarted worker continue tracking an in-flight job.
- Poll the documented status route with a delay that grows between attempts, or consume its documented event stream. Handle transient network errors and stop at a defined deadline.
- Record each URL’s terminal status and output reference. Confirm how long output URLs remain available before relying on them for archival or downstream processing.
Do not assume every request in a batch succeeds just because batch submission succeeded. A useful manifest includes the input URL, request or batch ID, final status, output location, elapsed time, and concise error detail. Preserve enough information to retry individual failures without recapturing successful URLs.
6. Capture settings that change the result
| Setting | When to use it | Things to check |
|---|---|---|
| Viewport | Make comparisons consistent or capture a desktop/mobile layout. | Responsive breakpoints, browser context dimensions, and device scale. |
| Viewport vs full page | Viewport for above-the-fold checks; full page for complete page review. | Very long documents can create large images or hit browser/service limits. |
| Navigation wait | domcontentloaded for an early page state; stronger waits if the target needs them. |
Network-idle assumptions can stall on pages with ongoing requests; test on actual targets. |
| Selector or delay | Wait for a known component or a bounded animation/data-loading period. | Selectors can change or never appear; set finite timeouts and log missing elements. |
| Format and quality | PNG for lossless UI details; JPEG/WebP for smaller lossy output where suitable. | Compare the resulting visual quality and storage size for your use case. |
| Mask/style | Hide volatile or sensitive areas for repeatable comparison images. | Do not accidentally conceal the content the screenshot is intended to verify. |
| Locale, timezone, geolocation | Reproduce region-dependent content where supported by your browser/API. | Pages can vary by locale, consent, authentication, and server-side personalization. |
For dynamic pages, a page-load event is not the same as “the screenshot is ready.” Wait for an application-specific selector when possible. If the page has no reliable readiness signal, use a deliberate, bounded delay and record that choice. For interactive states, perform the required navigation or clicks before capture; a basic screenshot call does not dismiss cookie banners or guarantee a stable page.
7. Reliability, performance, and cost
Separate capture failures from output failures. A page can navigate successfully and then fail during screenshot encoding, disk writes, upload, or manifest updates. Log these stages separately. Retry only transient failures, with a small attempt limit and backoff; repeatedly retrying a permanently invalid URL wastes capacity. Keep writes atomic where practical so a partial image is not mistaken for a complete result.
Reuse the browser process for a group of pages, but isolate page state when cookies or local storage could leak between unrelated targets. Recycle workers if memory rises across long batches. Full-page images, high device scale, and large viewports increase bytes and processing work. Smaller output formats may help storage and transfer, but compare quality before choosing them.
Cost depends on local compute and maintenance for Playwright, or provider pricing, request limits, retention, and options for a hosted service. Do not infer throughput from a vendor’s request rate or from an example script. For planning, capture a representative sample, measure your own end-to-end time and output size, and verify current API quotas and billing rules before a production rollout.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable missing | Playwright package installed without its browser binary. | Run python -m playwright install chromium in the runtime image or worker setup. |
| Navigation timeout | Slow target, stalled request, or overly strict readiness condition. | Use a finite timeout appropriate to the job; try domcontentloaded or wait for a known selector rather than an unbounded wait. |
| Screenshot is incomplete or blank | Client-rendered content has not appeared, an error page loaded, or the page is blocked. | Record response status, wait for a meaningful element, and inspect a sample capture before scaling up. |
| Element locator times out | Selector does not match this page, is in a frame, or is created only after interaction. | Check the selector against the page, account for iframe context, and wait for the correct state. |
| Images vary between runs | Animations, rotating content, personalization, fonts, or data timing vary. | Disable animations, wait for stable content, mask volatile regions, and standardize viewport and locale. |
| Duplicate or overwritten files | Names derived only from hostnames collide for multiple paths or repeat URLs. | Include a stable input index or record ID, as in the examples; sanitize filenames. |
| Batch submission rejected | Wrong endpoint schema, authentication, batch size, or quota. | Compare request details to the provider’s current docs and inspect response status/body without logging secrets. |
| Batch stuck in progress | Polling is too frequent, job is still rendering, or terminal states are not handled. | Use documented status semantics, back off polling, set a deadline, and preserve the batch ID. |
| Images exceed storage budget | Full-page captures or device-scale output create large files. | Use viewport capture or CSS scale where acceptable; consider JPEG/WebP quality and retention policy. |

Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from [ScreenshotNeo](https://screenshotneo.com). It supports bulk capture for up to 100 URLs per call, alongside single-URL captures. Its API returns PNG, JPEG, WebP, or PDF, and the service documents 63 capture options. For large Python jobs, use its bulk feature as described in the [ScreenshotNeo API docs](https://screenshotneo.com/docs/). The following is the one-call Python form for one URL; adapt the URL or use the documented bulk flow when submitting many:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Equivalent cURL and Node.js requests:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; response headers report the page verdict and billing status. An MCP server lets AI agents—including Claude, Cursor, and other MCP clients—take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. [Create a free ScreenshotNeo account](https://screenshotneo.com/account/sign-up/) to get started.
FAQ
Does Playwright have a built-in Python bulk screenshot endpoint?
The cited Python documentation describes capturing a page or locator. The loop, queue, retry rules, and batch manifest are your application’s responsibility.
Should I capture the viewport or the whole page?
Choose viewport for consistent first-screen checks and full-page when the complete document matters. Full-page output can be much taller and larger.
Can screenshots be processed without saving temporary files?
Yes. Playwright returns image bytes when no output path is supplied; pass those bytes to your storage or image-processing code.
How many pages should run at once?
There is no source-backed universal number. Begin with a small bounded worker count and tune from memory use, observed errors, and target behavior.


