How to Reduce Memory Use When Screenshotting Thousands of URLs
Keep bulk screenshot jobs within a predictable memory budget with bounded workers, smaller captures, prompt cleanup, and measured recovery.
To reduce memory use when screenshotting thousands of URLs, limit how many browser pages run at once, capture only the area you need, write each image to its destination promptly, and close page or context resources after every job. Read URLs from a queue instead of creating a page for every URL. Start with low concurrency, then raise it only after measuring peak memory, throughput, and failures on representative pages.
There is no universal safe worker count or memory-saving percentage: pages differ in size and behavior, and Playwright does not publish a general threshold. Its Page API documentation warns that pages can crash if they try to allocate too much memory. Treat bounded concurrency, cleanup, and measured worker recycling as operational practices, not guarantees.
1. Bound the number of active browser jobs
A queue with a fixed worker pool caps the number of simultaneous navigations, pages, and screenshot operations. The URL list can contain thousands of entries; the number of active browser jobs should stay at your configured limit.
For a first run, use one or two workers. Increase gradually while recording peak resident memory, jobs completed per minute, navigation timeouts, page crashes, and output sizes. If memory or failures rise sharply, reduce concurrency. A page-heavy site may need a lower limit than a simple static site.
Reuse a browser process where appropriate, but keep each job’s page lifecycle explicit. Close the page in a finally block. If you create a new context for each job, close that too. Recycle a worker or browser at a measured batch boundary if memory keeps climbing after cleanup; no fixed recycling interval is right for every workload.
2. Runnable Playwright example: bounded Node.js workers
This example reads URLs from a newline-delimited urls.txt, keeps at most CONCURRENCY pages active, saves viewport screenshots directly to files, and closes each page even when navigation or capture fails. Install Playwright with npm install playwright and install its browser with npx playwright install chromium. Run it with CONCURRENCY=2 node capture.mjs.
import { chromium } from 'playwright';
import { createReadStream } from 'node:fs';
import { mkdir } from 'node:fs/promises';
import { createInterface } from 'node:readline';
import { createHash } from 'node:crypto';
const concurrency = Math.max(1, Number(process.env.CONCURRENCY || 2));
const timeoutMs = Math.max(1000, Number(process.env.TIMEOUT_MS || 30000));
const outputDir = process.env.OUTPUT_DIR || 'shots';
await mkdir(outputDir, { recursive: true });
const urls = [];
const input = createInterface({
input: createReadStream(process.env.URLS_FILE || 'urls.txt'),
crlfDelay: Infinity,
});
for await (const line of input) {
const value = line.trim();
if (value && !value.startsWith('#')) urls.push(value);
}
const browser = await chromium.launch({ headless: true });
let next = 0;
let failed = 0;
async function worker() {
while (true) {
const index = next++;
if (index >= urls.length) return;
const url = urls[index];
const id = createHash('sha256').update(url).digest('hex').slice(0, 16);
let page;
try {
page = await browser.newPage({ viewport: { width: 1280, height: 800 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.screenshot({ path: `${outputDir}/${id}.png` });
console.log(`OK ${index + 1}/${urls.length} ${url}`);
} catch (error) {
failed++;
console.error(`FAIL ${url}: ${error.message}`);
} finally {
if (page) await page.close().catch(() => {});
}
}
}
try {
await Promise.all(Array.from({ length: concurrency }, () => worker()));
} finally {
await browser.close();
}
console.log(`Finished: ${urls.length - failed} succeeded, ${failed} failed`);
if (failed) process.exitCode = 1;
The example loads the URL list into memory, but not browser pages or screenshot buffers for every URL. For extremely large input lists, replace the urls array with a streaming queue so the orchestrator also has bounded input memory. The filename is a stable hash of the URL, avoiding unsafe path characters and collisions from simple URL-to-filename substitutions.
Choose the navigation wait condition deliberately
domcontentloaded is a practical starting point when the page’s initial document is enough. Use load if the screenshot depends on resources that finish loading before the load event. Use networkidle only when the page actually becomes idle; analytics, polling, and long-lived requests can make it wait too long. For a specific component, wait for that selector instead of waiting for every network request:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.locator('[data-ready="true"]').waitFor({ state: 'visible', timeout: 10000 });
await page.locator('#price-chart').screenshot({ path: `${outputDir}/${id}.png` });
Replace the selector and readiness condition with ones meaningful to the target pages. A fixed sleep can be useful for a known delayed animation, but it adds that delay to every job and does not prove the content is ready.
3. Capture less data per URL
Playwright supports viewport, full-page, and element screenshots. The screenshots guide documents these capture modes. Choose the smallest scope that answers your use case.
| Capture | Use it when | Memory consideration |
|---|---|---|
| Viewport | You need the initially visible screen | Usually avoids producing an image as tall as the entire document |
| Element | You need one chart, card, or component | Limits the screenshot output to a focused target |
| Full page | You need the entire scrollable document | Long pages can produce very large image dimensions and output |
Full-page capture is not a free substitute for viewport capture on very long pages. Use it only when the full document is required, and monitor output dimensions and memory on the longest pages in your sample. The documentation does not state a percentage reduction from choosing a smaller capture.
4. Save images promptly and release references
Playwright can save screenshots to a file path or return screenshot bytes. In a bulk pipeline, prefer writing to the final destination or processing the returned bytes immediately. Avoid collecting buffers or base64 strings in arrays while the rest of the URL queue runs. Base64 encoding also increases the in-memory representation size compared with raw bytes.
If the destination is remote object storage, upload one completed image at a time or through a separate bounded upload queue. Do not allow slow uploads to create an unbounded backlog of screenshot buffers. Apply backpressure: when the output destination is behind, pause capture workers or reduce their rate.
5. Python and cURL options
Playwright’s Python API follows the same lifecycle: launch a browser, create a page, navigate, capture to a path, and close resources. Install with pip install playwright and playwright install chromium. Save as capture.py, put one URL per line in urls.txt, and run CONCURRENCY=2 python capture.py.
import asyncio
import hashlib
import os
from pathlib import Path
from playwright.async_api import async_playwright
CONCURRENCY = max(1, int(os.getenv("CONCURRENCY", "2")))
TIMEOUT_MS = max(1000, int(os.getenv("TIMEOUT_MS", "30000")))
OUTPUT_DIR = Path(os.getenv("OUTPUT_DIR", "shots"))
URLS_FILE = Path(os.getenv("URLS_FILE", "urls.txt"))
async def main():
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
urls = [line.strip() for line in URLS_FILE.read_text().splitlines()
if line.strip() and not line.lstrip().startswith("#")]
queue = asyncio.Queue()
for index, url in enumerate(urls):
queue.put_nowait((index, url))
failures = 0
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
async def worker():
nonlocal failures
while True:
try:
index, url = queue.get_nowait()
except asyncio.QueueEmpty:
return
page = None
try:
page = await browser.new_page(viewport={"width": 1280, "height": 800})
await page.goto(url, wait_until="domcontentloaded", timeout=TIMEOUT_MS)
name = hashlib.sha256(url.encode()).hexdigest()[:16] + ".png"
await page.screenshot(path=str(OUTPUT_DIR / name))
print(f"OK {index + 1}/{len(urls)} {url}")
except Exception as exc:
failures += 1
print(f"FAIL {url}: {exc}")
finally:
if page is not None:
await page.close()
queue.task_done()
try:
await asyncio.gather(*(worker() for _ in range(min(CONCURRENCY, len(urls)))))
finally:
await browser.close()
print(f"Finished: {len(urls) - failures} succeeded, {failures} failed")
if failures:
raise SystemExit(1)
asyncio.run(main())
For cURL, a local browser endpoint or a screenshot API must provide an HTTP capture interface; cURL itself does not render web pages. If you already have a compatible endpoint that accepts a URL and returns an image, send one request per URL with a bounded request queue and stream the response to a file rather than accumulating bodies. The following is the ScreenshotNeo API example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
6. Retries, failures, and worker recovery
Retries should go back into a bounded queue, not spawn replacement pages outside the worker limit. Record the URL, attempt number, error, elapsed time, and output status. Retry transient navigation or network failures a small, configured number of times with backoff. Do not retry invalid URLs indefinitely. A timeout or browser crash should count as a failed attempt and release the page/context before another attempt starts.
If memory rises steadily across completed jobs, first verify that every page and context is closed and that image buffers are not retained. Then inspect whether a subset of pages is unusually long, interactive, or resource-heavy. If cleanup does not stabilize the worker, recycle browser workers at a batch boundary chosen from measurements. Worker recycling is a recovery policy, not a documented universal Playwright requirement.
7. Measure and tune for your workload
- Choose representative URLs, including long pages, pages with large images, and pages that often fail.
- Run at low concurrency and record peak resident memory, throughput, failure rate, timeouts, and image dimensions.
- Increase concurrency in small steps while monitoring those same measures.
- Choose the highest concurrency that stays within your memory budget and acceptable failure rate; leave headroom for unusually heavy pages.
- Repeat after changing browser version, capture scope, wait condition, or page mix.
Use the operating system’s process metrics or your container’s memory reporting to observe peak resident memory. Track per-job duration and output size so a few heavy pages do not disappear inside an overall average. There is no source-supported fixed number of tabs, megabytes per page, or safe worker count to apply to all workloads.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Memory grows throughout the batch | Pages or contexts remain open, image bytes are retained, or some pages use much more memory | Check the cleanup path and buffer references; lower concurrency; identify whether growth follows particular URLs; test measured worker recycling |
| Browser or page crashes | A page allocated too much memory or concurrent browser load is too high | Reduce concurrency, retry through the same bounded queue, log the URL, and isolate heavy pages |
| Navigation times out | The page is slow, keeps requests open, or the selected wait condition is too strict | Set an explicit timeout; use the least strict wait condition that gives correct output; wait for a relevant selector when possible |
| Screenshot misses late content | Capture starts before the required content is ready | Wait for a meaningful selector or app readiness signal; use a targeted delay only when the delay is understood |
| Full-page capture uses much more memory | The document is very tall or contains extensive content | Use viewport or element capture if sufficient; reserve full-page capture for cases that need it |
| Memory spikes during output upload | Captured data is queued faster than the destination can store it | Stream or upload promptly and bound the upload queue; apply backpressure to capture workers |
| Many duplicate output files overwrite each other | Filenames are derived from a non-unique or unsafe URL fragment | Use a stable URL hash or another collision-resistant identifier and keep the URL-to-file mapping |
9. Cost and reliability considerations
Self-hosted browser jobs trade infrastructure cost and operational work against control over capture behavior. Higher concurrency can improve throughput, but it also raises peak memory and can increase crashes or retries. Full-page images consume more output storage and transfer than smaller captures. Measure your own workload before estimating the infrastructure required; the available Playwright documentation does not provide a universal per-page cost or memory benchmark.
For reliable batches, persist job status and output mapping, make retries idempotent, and keep failed URLs available for later inspection. Limit both browser concurrency and any downstream upload concurrency. Keep enough memory headroom to handle outlier pages without losing the whole batch.
10. Or skip the browser setup
If you do not want to run and tune browser workers, ScreenshotNeo accepts a URL in one GET request and returns an image or PDF. Its API documentation covers the available parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo and the API docs, then sign up for 1,000 free screenshots a month with no card.
FAQ
Does closing a page guarantee that browser memory immediately returns to its starting level?
No. Close pages and contexts to release their resources, then observe the browser process over a representative batch. If memory continues to trend upward, investigate retained data and page behavior, and evaluate measured worker recycling.
Should I use a new browser process for every URL?
Usually, a fixed set of workers with explicit per-job cleanup is a more practical starting point. Whether to recycle browser processes depends on measured memory trends and failure behavior for your pages.
Can I process thousands of URLs without loading the entire list into memory?
Yes. Stream URLs from a file or database into a bounded queue and let a fixed number of workers consume it. The sample code loads the URL list for simplicity; the browser concurrency remains bounded either way.
Does a smaller screenshot always mean the page itself uses less memory?
No. A smaller capture reduces the screenshot scope or output, but the page still has to load and render. Measure browser memory as well as image dimensions and output size.


