How to Bulk Screenshot a List of URLs with Playwright in Python
Capture a list of URLs to PNG with Playwright in Python. Learn batching, concurrency, full-page captures, error handling, and output management.
Use Playwright’s async Python API to open each URL in a fresh page, navigate, save a screenshot, and close the page. Bound concurrent captures with an asyncio.Semaphore, handle errors per URL, and close the browser context and browser in a finally block. Use full_page=True to capture the full scrollable document; omit it for a viewport screenshot. Playwright documents both synchronous and asynchronous Python APIs, with async suited to applications already using asyncio (Playwright Python documentation).
1. Install Playwright
Install the Python package and the browser binaries. This example uses Chromium.
python -m pip install playwright
python -m playwright install chromium
Save the script below as bulk_screenshot.py, then run python bulk_screenshot.py. It creates a screenshots directory and writes one PNG per URL.
2. Capture a URL list with bounded concurrency
import asyncio
from pathlib import Path
from playwright.async_api import async_playwright
URLS = [
"https://example.com/",
"https://playwright.dev/python/",
]
OUT = Path("screenshots")
MAX_CONCURRENT_PAGES = 4 # Example limit; tune for your workload.
NAVIGATION_TIMEOUT_MS = 30_000
async def main():
OUT.mkdir(parents=True, exist_ok=True)
semaphore = asyncio.Semaphore(MAX_CONCURRENT_PAGES)
async with async_playwright() as p:
browser = await p.chromium.launch()
context = await browser.new_context(
viewport={"width": 1440, "height": 1000}
)
async def capture(index, url):
async with semaphore:
page = await context.new_page()
try:
response = await page.goto(
url,
wait_until="load",
timeout=NAVIGATION_TIMEOUT_MS,
)
status = response.status if response else None
output_path = OUT / f"{index:04d}.png"
await page.screenshot(
path=str(output_path),
full_page=True,
)
return {
"url": url,
"status": status,
"file": str(output_path),
}
except Exception as exc:
return {"url": url, "error": str(exc)}
finally:
await page.close()
try:
results = await asyncio.gather(
*(capture(index, url) for index, url in enumerate(URLS, start=1))
)
finally:
await context.close()
await browser.close()
for result in results:
print(result)
if __name__ == "__main__":
asyncio.run(main())
The semaphore limits how many capture tasks hold a page at once. Four is only an example value, not a universal recommendation. Start with a small limit and adjust for the machine, page weight, and target sites. More concurrency can use more memory and place more simultaneous load on sites. Playwright’s API supports multiple pages in one context; pages in that context share context-level settings such as viewport and device emulation (multiple pages).
3. Choose navigation and screenshot behavior
Viewport or full-page image
By default, page.screenshot() captures the current viewport. Set full_page=True when you need the complete scrollable document as one tall image. Very long pages can create large images and require more memory. Screenshot options and byte output are described in the Playwright screenshot guide.
When to wait before capture
wait_until="load" waits for the page load event; it does not guarantee that a client-rendered app has finished fetching data or painting all content. If a specific element indicates readiness, wait for it explicitly:
await page.goto(url, wait_until="domcontentloaded", timeout=30_000)
await page.locator("main article").wait_for(state="visible", timeout=15_000)
await page.screenshot(path=str(output_path), full_page=True)
Replace the selector with one that reflects the target site. For pages whose content appears after a known interaction, perform that interaction or wait for a site-specific state. Avoid relying on an arbitrary sleep unless there is no observable readiness condition.
Playwright navigation can use different lifecycle conditions such as domcontentloaded and load. Choose the earliest condition that still gives the page enough time to reach the state you need. A page can keep network connections open, so waiting for network idle is not appropriate for every site.
File paths, bytes, and stable names
The sample names files by input position, so duplicate URLs do not overwrite each other and raw URLs never become filesystem paths. For repeatable names across reordered batches, use a stable record ID or a hash of the URL, while retaining an index or manifest to map each output back to its source. Do not use an unsanitized URL as a filename.
For an upload or image-processing pipeline, request screenshot bytes instead of writing directly to disk:
image_bytes = await page.screenshot(full_page=True)
# Pass image_bytes to your storage client or image-processing code.
Writing to a path is simpler for a local batch; bytes are useful when another part of the program consumes the image. The screenshot API supports both path output and returning bytes (screenshot options).
4. Configure shared settings and isolation
A browser context holds session and emulation settings. Pages created from the same context share those settings, including the configured viewport. This is useful when every URL should use the same browser profile and dimensions. Create separate contexts when captures need isolated cookies, storage, or other session state; contexts are isolated browser sessions (browser contexts).
Set viewport when creating the context, as in the sample. You can also configure device emulation through context options. Keep the intended viewport consistent across the batch if images will be compared. If each URL needs its own cookies or authentication, use a separate context per session rather than sharing mutable session state accidentally.
Always clean up the page, context, and browser. The sample closes each page in finally and closes the context and browser even if the batch raises an error. Playwright’s browser lifecycle documentation covers launching and closing browser instances (Browser API).
5. Handle failures without losing the batch
The sample returns an error record for each failed navigation or screenshot, allowing the rest of the list to finish. In a production batch, write results to a JSONL or CSV manifest as they complete rather than keeping all results only in memory. Include at least the input identifier, URL, output path, HTTP status when available, and error message.
For transient failures, add a small, bounded retry policy around navigation and capture. Retry only errors that may be temporary, use a delay between attempts, and cap the attempts. Avoid retrying every failure indefinitely: invalid URLs, access-denied pages, and persistent timeouts will not be fixed by repeated requests. Be considerate of the target site’s rate limits and policies.
6. Synchronous alternative for a short script
If the surrounding program is not asynchronous and the list is small, the synchronous API can be simpler to read. It processes URLs sequentially:
from pathlib import Path
from playwright.sync_api import sync_playwright
urls = ["https://example.com/", "https://playwright.dev/python/"]
out = Path("screenshots")
out.mkdir(parents=True, exist_ok=True)
with sync_playwright() as p:
browser = p.chromium.launch()
context = browser.new_context(viewport={"width": 1440, "height": 1000})
try:
for index, url in enumerate(urls, start=1):
page = context.new_page()
try:
response = page.goto(url, wait_until="load", timeout=30_000)
page.screenshot(path=str(out / f"{index:04d}.png"), full_page=True)
print({"url": url, "status": response.status if response else None})
except Exception as exc:
print({"url": url, "error": str(exc)})
finally:
page.close()
finally:
context.close()
browser.close()
Use async when your application already uses asyncio or you need to coordinate many independent tasks. Async syntax alone does not guarantee higher throughput; page weight, machine resources, target response times, and your concurrency limit all matter.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
Executable doesn't exist or browser launch fails |
The Playwright package is installed, but its browser binary is not. | Run python -m playwright install chromium in the same environment where the script runs. |
| Navigation timeout | The site is slow, the connection is stalled, or the chosen lifecycle event never arrives. | Check the URL and network access. Raise the timeout only when justified, or use an earlier lifecycle event and wait for the page-specific element you actually need. |
| Screenshot is missing content | The page rendered content after the chosen navigation event. | Wait for a visible selector or another meaningful application-ready condition before capture. |
| Screenshot is only the first screen | The screenshot call used its default viewport capture. | Pass full_page=True for the full scrollable document. |
| Some outputs overwrite others | Output names are derived from a non-unique value, such as the URL basename. | Use a unique index, record ID, or collision-resistant stable key; keep a manifest mapping files to inputs. |
| Batch uses too much memory or pages fail under load | Too many heavy pages are open simultaneously, or full-page images are large. | Lower the semaphore limit, process a large input list in chunks, and consider viewport captures if a full document is unnecessary. |
| One error stops other captures | An exception escaped the per-URL task or was raised while collecting the batch. | Catch errors inside each capture task, record the failed URL, and keep cleanup in finally blocks. |
| HTTP error page was saved as an image | A navigation response can have an error status without raising a navigation exception. | Inspect response.status and decide whether to save, flag, or skip non-success statuses for your use case. |
8. Performance, reliability, and cost considerations
- Concurrency: Bound in-flight pages. Tune from a conservative starting point based on memory use, page complexity, network conditions, and how much load is appropriate for target sites. There is no universal useful concurrency setting or published benchmark in the cited Playwright documentation.
- Batch size: For a very large URL list, enqueue work in chunks or use a worker queue instead of creating an unbounded number of tasks. Persist each result as it finishes so a process interruption does not erase completed work.
- Reliability: Record failures independently, use bounded retries for transient problems, preserve a mapping from inputs to files, and close resources on all paths. Consider whether failed HTTP statuses should count as successful captures.
- Output size: Full-page captures may produce tall, large images. Use viewport screenshots where they answer the task, and plan disk or upload capacity for the expected outputs.
- Cost: Playwright is browser automation that you run in your own environment; the script itself does not charge per screenshot. Account for the machine, storage, network, and any hosted infrastructure you choose to run it on. No hosting price or performance figure is assumed here.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. ScreenshotNeo supports bulk capture of up to 100 URLs per call, along with caching, signed links, async jobs, and other capture settings. It offers 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 screenshots. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month, no card required.
FAQ
Can I capture the same URL more than once?
Yes. Each input gets its own indexed filename in the sample, so repeated URLs produce separate files.
Does a successful navigation mean the page is ready for an accurate screenshot?
Not always. Navigation lifecycle events do not know when a particular application has finished rendering its meaningful content. Wait for a page-specific signal when needed.
Should I reuse one page for every URL?
A fresh page per capture makes cleanup and per-URL state easier to reason about. Reusing pages can be appropriate in a controlled workflow, but ensure state from the previous URL cannot affect the next image.
Can I change the browser engine?
Yes. Playwright supports Chromium, Firefox, and WebKit. Install the browser you plan to launch and use its corresponding launcher from the Playwright Python API.


