How to Bulk Screenshot URLs and Add Page Titles to the Output Filenames
Use Playwright to capture a list of URLs and save each screenshot with a safe, readable page-title filename. Includes collision handling, full-page options, and a ready-to-run Python script.
Use Playwright to loop over your URLs, navigate to each page, read its document title with page.title(), sanitize that title for use as a filename, and save a screenshot with page.screenshot(path=...). Add an index or URL-derived suffix so pages with identical titles do not overwrite one another. Set full_page=True to capture the full scrollable page; omit it for the current viewport.
1. Install Playwright and a browser
This guide uses Playwright’s synchronous Python API and Chromium. Install the Python package and the browser binary from your project environment:
python -m pip install playwright
python -m playwright install chromium
The script below uses only those dependencies. It reads URLs from a text file, writes PNG files to an output directory, preserves the title in each filename, avoids collisions with a sequence number, and continues if an individual URL fails. Playwright also offers asynchronous Python APIs and other browser engines; choose the runtime and engine appropriate to your project.
2. Prepare the URL list
Create urls.txt with one absolute HTTP or HTTPS URL per line. Blank lines and lines beginning with # are ignored.
https://example.com/
https://playwright.dev/python/
# Add one URL per line
3. Run the bulk screenshot script
Save this as bulk_screenshot.py. By default it captures each viewport. Pass --full-page to include the full scrollable page, or change the format option to save JPEG or WebP files.
from __future__ import annotations
import argparse
import re
import sys
from pathlib import Path
from urllib.parse import urlsplit
from playwright.sync_api import TimeoutError as PlaywrightTimeoutError
from playwright.sync_api import sync_playwright
INVALID_FILENAME_CHARS = re.compile(r'[<>:"/\\|?*\x00-\x1f]')
WHITESPACE = re.compile(r'\s+')
def safe_stem(title: str, fallback: str, limit: int = 120) -> str:
"""Make a readable, conservative filename stem from untrusted page text."""
stem = INVALID_FILENAME_CHARS.sub('_', title)
stem = WHITESPACE.sub(' ', stem).strip(' .')
# Avoid empty names and special directory components.
if not stem or stem in {'.', '..'}:
stem = fallback
return stem[:limit].rstrip(' .') or fallback
def read_urls(path: Path) -> list[str]:
urls = []
for line_number, raw in enumerate(path.read_text(encoding='utf-8').splitlines(), 1):
value = raw.strip()
if not value or value.startswith('#'):
continue
parsed = urlsplit(value)
if parsed.scheme not in {'http', 'https'} or not parsed.netloc:
raise ValueError(f'{path}:{line_number}: expected an absolute HTTP(S) URL, got {value!r}')
urls.append(value)
return urls
def main() -> int:
parser = argparse.ArgumentParser(description='Capture URLs into title-named screenshots.')
parser.add_argument('url_file', nargs='?', default='urls.txt', help='one URL per line (default: urls.txt)')
parser.add_argument('--out', default='screenshots', help='output directory (default: screenshots)')
parser.add_argument('--full-page', action='store_true', help='capture the full scrollable page')
parser.add_argument('--format', choices=('png', 'jpeg', 'webp'), default='png')
parser.add_argument('--timeout-ms', type=int, default=30000, help='navigation timeout per URL')
parser.add_argument('--wait-until', choices=('commit', 'domcontentloaded', 'load', 'networkidle'),
default='domcontentloaded')
args = parser.parse_args()
url_file = Path(args.url_file)
out_dir = Path(args.out)
try:
urls = read_urls(url_file)
except (OSError, ValueError) as exc:
print(f'Input error: {exc}', file=sys.stderr)
return 2
if not urls:
print(f'No URLs found in {url_file}', file=sys.stderr)
return 2
out_dir.mkdir(parents=True, exist_ok=True)
failures = 0
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.set_default_navigation_timeout(args.timeout_ms)
for index, url in enumerate(urls, 1):
fallback_host = urlsplit(url).hostname or 'page'
fallback = f'{fallback_host}-page-{index:03d}'
try:
response = page.goto(url, wait_until=args.wait_until)
title = page.title()
stem = safe_stem(title, fallback)
# The index makes repeated titles unique and preserves input order.
output = out_dir / f'{index:03d}-{stem}.{args.format}'
screenshot_options = {'path': str(output), 'full_page': args.full_page}
if args.format == 'jpeg':
screenshot_options['type'] = 'jpeg'
screenshot_options['quality'] = 85
elif args.format == 'webp':
screenshot_options['type'] = 'webp'
screenshot_options['quality'] = 85
page.screenshot(**screenshot_options)
status = response.status if response else 'no response'
print(f'OK [{status}] {url} -> {output}')
except PlaywrightTimeoutError as exc:
failures += 1
print(f'TIMEOUT {url}: {exc}', file=sys.stderr)
except Exception as exc:
failures += 1
print(f'FAILED {url}: {type(exc).__name__}: {exc}', file=sys.stderr)
browser.close()
print(f'Finished: {len(urls) - failures} succeeded, {failures} failed; output: {out_dir}')
return 1 if failures else 0
if __name__ == '__main__':
raise SystemExit(main())
Run it from the directory containing urls.txt:
python bulk_screenshot.py
python bulk_screenshot.py urls.txt --out captures --full-page
python bulk_screenshot.py urls.txt --format webp --timeout-ms 60000
The code uses domcontentloaded as a practical default: it waits for the initial document to be parsed without requiring every subresource to finish. Some sites populate the title or content later, so choose another readiness condition or add a site-specific wait when needed. The returned HTTP response status is printed for diagnosis; a non-2xx page can still render and be captured.
4. Choose capture scope and page readiness
Viewport or full page
By default, Playwright saves the visible viewport. Add --full-page to capture the full scrollable page. Full-page captures can be much taller and use more memory. Long pages that lazy-load content as you scroll may need an explicit scrolling or page-specific loading step before capture; a full-page flag alone should not be treated as a guarantee that every lazy image has loaded.
Navigation wait conditions
| Condition | What it waits for | When it can help |
|---|---|---|
commit |
Navigation has committed and the response has begun. | Fast initial capture when the page’s useful content appears immediately. |
domcontentloaded |
The document’s DOM content has loaded. | A reasonable general default for this batch script. |
load |
The page load event has fired. | Pages that need more subresources before their initial rendering is useful. |
networkidle |
Network activity has been idle for the browser’s defined interval. | Some pages that make a finite set of background requests; can stall on sites with persistent traffic. |
No single wait condition suits every site. If a title appears only after a client-side application initializes, wait for a selector or a known page-specific condition before calling page.title(). A fixed delay is simple but can waste time on fast pages and still be too short on slow ones.
5. Filename safety and collision handling
Page titles come from remote content and should be treated as input data. The script replaces common reserved filename characters, collapses whitespace, trims trailing spaces and dots, limits the title portion, and uses a host-based fallback if the title is empty. The numeric prefix prevents two pages titled “Home” from overwriting each other and makes output order visible.
Filename rules differ across operating systems and filesystems. If you need reproducible names across machines, consider using a stable URL-derived hash or a stricter character allowlist instead of relying on title text alone. Avoid using a title directly as a path: slashes and other reserved characters can create invalid paths or unintended directory components.
6. Python async variant
For a large batch or an application that already uses asyncio, Playwright’s async API avoids blocking the event loop. This sequential version has the same basic behavior and filename protection. Add bounded concurrency only after measuring the target sites and available resources; opening many pages at once can increase memory, network load, and rate limiting.
import asyncio
from pathlib import Path
import re
from playwright.async_api import async_playwright
URLS = ['https://example.com/', 'https://playwright.dev/python/']
INVALID = re.compile(r'[<>:"/\\|?*\x00-\x1f]')
def stem(title: str, fallback: str) -> str:
value = INVALID.sub('_', title).strip(' .')[:120].rstrip(' .')
return value or fallback
async def main():
out = Path('screenshots')
out.mkdir(parents=True, exist_ok=True)
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
for i, url in enumerate(URLS, 1):
try:
await page.goto(url, wait_until='domcontentloaded', timeout=30000)
title = await page.title()
path = out / f'{i:03d}-{stem(title, f"page-{i:03d}")}.png'
await page.screenshot(path=str(path), full_page=True)
print(f'Saved {path}')
except Exception as exc:
print(f'Failed {url}: {type(exc).__name__}: {exc}')
await browser.close()
asyncio.run(main())
7. Use the documented title and screenshot APIs directly
The loop above combines Playwright’s individual page operations into a batch workflow; bulk iteration and the filename policy are your application logic, not a built-in bulk command. The official guides cover [Python screenshots](https://playwright.dev/python/docs/screenshots), [navigation and page titles](https://playwright.dev/python/docs/intro), and [the Page API](https://playwright.dev/python/docs/api/class-page). The Page API documents that screenshot type is inferred from the filename extension and that relative paths are resolved from the current working directory.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. The API returns a screenshot with one GET request; for bulk capture, send up to 100 URLs per call and use the usage API to inspect consumption. Its response headers identify the page verdict and whether the result was billed, and cache hits cost nothing.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/ \
-o shot.webp
See the ScreenshotNeo API documentation for authentication and request options. To name downloaded files from page titles, retrieve page information for each URL, sanitize the returned title using the filename rules above, then save each response under the resulting name. That keeps the local naming policy explicit while outsourcing browser execution. You can request PNG, JPEG, WebP, or PDF, and other screenshot APIs’ parameter names also work to simplify switching.
Cookie banners are accepted and removed along with 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, and failed loads are not billed, and cache hits cost nothing. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
9. Performance, reliability, and cost
- Throughput: This script processes URLs sequentially, which limits simultaneous browser work and makes failures easier to associate with an input. For faster batches, use a bounded number of pages or browser contexts and monitor memory, network usage, and destination-site limits. No fixed throughput is guaranteed; pages vary in size and readiness.
- Reliability: A timeout on one page is recorded and the loop continues. The process exits with status 1 if any capture failed and status 2 for invalid or empty input. Keep the URL list and output directory so a rerun can target failed inputs. Add retries only for transient errors, and use a small retry limit with a delay to avoid repeatedly hammering a site.
- Cost: Playwright itself is an open-source browser automation library, but running it consumes your machine or hosted compute, network bandwidth, and storage. Full-page images and high-resolution pages use more memory and disk. A hosted screenshot API shifts browser operation to the service and may bill according to its plan and response rules; check the provider’s current terms.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Executable doesn't exist or browser launch failure |
The Playwright package is installed, but its browser binary is missing. | Run python -m playwright install chromium in the same environment used to run the script. |
| Navigation timeout | The site is slow, keeps requests open, or the chosen wait condition is too strict. | Raise --timeout-ms, try domcontentloaded or commit, and wait for the specific content needed. |
| Blank title or generic filename | The document title is empty or is set after the initial DOM event. | Wait for a known selector or app-ready condition before reading page.title(); the host-based fallback prevents an empty filename. |
| Duplicate titles | Multiple URLs use the same document title. | Keep the sequence prefix, or append a stable URL hash if names must be stable across list reordering. |
| Invalid path or file creation error | The title contains reserved characters, trailing dots/spaces, is too long, or the output directory is not writable. | Use the sanitizer, shorten the title limit, and choose a writable --out directory. |
| Screenshot misses images or content | Content is lazy-loaded or rendered after navigation readiness. | Wait for a selector, scroll relevant sections into view, or use a page-specific readiness check before capture. |
| Access denied or CAPTCHA | The destination blocks automated browsing or requires an authenticated session. | Use an authorized session and site-approved access method. Do not attempt to bypass access controls. |
| Files are unexpectedly large or capture is slow | Full-page scope, large image dimensions, or heavy page resources increase work. | Capture the viewport when sufficient, select a smaller viewport, or use JPEG/WebP when lossy compression is acceptable. |
11. FAQ
Can the filenames contain the original URL as well as the title?
Yes. Add a short slug or stable hash derived from the URL to the filename. A hash avoids collisions without putting long query strings or sensitive URL parameters into a filename.
Will the numeric prefix stay the same if I reorder the input list?
No. It reflects the current input order. Use a URL-derived identifier if filenames must remain stable across runs with reordered inputs.
Can I capture PDFs instead of image files?
Playwright has a separate PDF workflow for supported browser configurations. If your required output is a PDF, use the browser’s PDF operation and define page size, margins, and print behavior explicitly rather than changing only the screenshot extension.
Does a successful screenshot mean the page returned HTTP 200?
No. A browser may render an error page or other non-2xx response. The script reports the navigation response status separately from whether the screenshot file was saved.


