How to Take Bulk Screenshots of URLs While Blocking Ads and Trackers
Use Playwright to capture a URL list with request blocking, stable filenames, and per-page error records. Learn the limits of blocking and when to use a screenshot API.
Short answer: Use Playwright to loop over a validated URL list, install request-routing rules before navigation, and save each capture under a unique filename. Block service workers in the browser context when relying on page.route(), because service-worker-intercepted requests are not handled by that route. Request filtering can reduce ads and tracking requests, but its coverage depends on your rules and browser configuration; it does not guarantee that every tracking technique is stopped.
This guide shows a runnable Node.js batch script, explains the choices that affect consistency and output, and covers failures and operational limits. Playwright’s documented primitives include navigation, routing, and screenshots; the batch loop and reporting behavior below are implementation guidance built around those primitives. The code has not been run or tested for this article.
1. Install Playwright and prepare a URL list
Use a supported Node.js installation, create a project, and install Playwright’s Chromium browser:
mkdir bulk-captures
cd bulk-captures
npm init -y
npm install playwright
npx playwright install chromium
Save one absolute HTTP or HTTPS URL per line in urls.txt:
https://example.com/
https://www.iana.org/
https://playwright.dev/
Keep input validation explicit. URLs without a scheme, malformed values, and non-web protocols should be rejected or recorded as input errors rather than passed blindly to the browser. The script below accepts only http: and https:.
2. Capture the list with Playwright and filter requests
Create capture.mjs. This example applies a small illustrative set of URL-pattern rules, takes viewport screenshots by default or full-page screenshots when requested, and writes a JSON lines report with one result for every input URL. The sample patterns are not a comprehensive or maintained ad/tracker blocklist. Replace them with a rule source that suits your use case, and review updates before applying them to production captures.
import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';
const inputPath = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'shots';
const fullPage = process.env.FULL_PAGE === '1';
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 30000);
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 900);
// Illustrative examples only. Maintain and review your own rules.
const blockedPatterns = [
/(^|\.)doubleclick\.net$/i,
/(^|\.)googlesyndication\.com$/i,
/(^|\.)google-analytics\.com$/i,
/(^|\.)googletagmanager\.com$/i,
];
function shouldBlock(urlString) {
try {
const host = new URL(urlString).hostname;
return blockedPatterns.some((pattern) => pattern.test(host));
} catch {
return false;
}
}
function fileNameFor(url) {
// A readable prefix plus a digest prevents collisions between similar URLs.
const host = new URL(url).hostname.replace(/[^a-z0-9.-]/gi, '_');
const digest = createHash('sha256').update(url).digest('hex').slice(0, 12);
return `${host}-${digest}.png`;
}
const raw = await readFile(inputPath, 'utf8');
const urls = raw.split(/\r?\n/).map((line) => line.trim()).filter(Boolean);
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
viewport: { width, height },
serviceWorkers: 'block',
});
const reportPath = `${outputDir}/results.jsonl`;
const reportLines = [];
try {
for (const original of urls) {
let url;
try {
url = new URL(original);
if (!['http:', 'https:'].includes(url.protocol)) {
throw new Error(`Unsupported protocol: ${url.protocol}`);
}
} catch (error) {
reportLines.push(JSON.stringify({ input: original, status: 'input_error', error: error.message }));
continue;
}
const page = await context.newPage();
let blockedCount = 0;
try {
// Install interception before navigation so initial document resources are covered.
await page.route('**/*', async (route) => {
if (shouldBlock(route.request().url())) {
blockedCount += 1;
await route.abort();
} else {
await route.continue();
}
});
const response = await page.goto(url.href, {
waitUntil: 'domcontentloaded',
timeout: timeoutMs,
});
// A page can continue rendering after DOMContentLoaded. This short settling
// delay is configurable; choose a page-specific readiness rule when needed.
await page.waitForTimeout(Number(process.env.SETTLE_MS ?? 1000));
const path = `${outputDir}/${fileNameFor(url.href)}`;
await page.screenshot({ path, fullPage });
reportLines.push(JSON.stringify({
input: original,
finalUrl: page.url(),
status: 'captured',
httpStatus: response?.status() ?? null,
blockedRequests: blockedCount,
file: path,
fullPage,
}));
} catch (error) {
reportLines.push(JSON.stringify({
input: original,
finalUrl: page.url(),
status: 'capture_error',
blockedRequests: blockedCount,
error: error.message,
}));
} finally {
await page.close();
}
}
} finally {
await browser.close();
await writeFile(reportPath, reportLines.join('\n') + (reportLines.length ? '\n' : ''));
}
console.log(`Processed ${urls.length} entries. Results: ${reportPath}`);
Run it:
node capture.mjs urls.txt shots
To capture full scrollable pages and adjust timing or viewport:
FULL_PAGE=1 VIEWPORT_WIDTH=1365 VIEWPORT_HEIGHT=900 TIMEOUT_MS=45000 SETTLE_MS=1500 node capture.mjs urls.txt shots
The output directory contains PNG files and results.jsonl. A URL’s filename combines its hostname and a digest of the complete URL, so two paths on the same site do not overwrite each other. The report records invalid inputs, navigation or screenshot errors, final redirected URLs, HTTP response status when available, and the number of requests aborted by the configured rules.
3. Choose rules and capture settings deliberately
Request interception and service workers
Playwright’s page.route() lets a script inspect requests and continue or abort them. Add the route before navigating so it can see page-load requests. Playwright documents a limitation: requests intercepted by a service worker are not handled by page routing. Its documentation recommends blocking service workers when request interception is needed; the sample sets serviceWorkers: 'block' on the context. This changes page behavior for sites that rely on service workers, so record that setting when reproducibility matters.
Maintain rules as data rather than scattering one-off conditions throughout a capture script. A rule set can match hosts or request URLs, and may classify resource types such as image, script, stylesheet, or font. Host-based filtering is easier to inspect but can block a site’s required first-party content. Resource-type blocking is broader and can damage layout or functionality. Start with narrow rules, log what was blocked, and inspect representative pages after rule updates.
Blocking known ad or analytics endpoints does not prove that all tracking is prevented. First-party collection, server-side measurement, fingerprinting, and endpoints absent from the rules may remain. Conversely, blocking a shared CDN or a script host can prevent the page itself from rendering correctly.
Wait condition and readiness
The sample uses domcontentloaded followed by a configurable settling delay. This avoids waiting indefinitely for pages that keep network connections open, but a fixed delay may capture before late content appears or waste time on fast pages. For a known site, wait for a meaningful selector with page.waitForSelector(), or use another navigation wait condition appropriate to that site. A network-idle condition can be unsuitable for pages with continual polling or analytics traffic. Cloudflare’s snapshot example also exposes wait and timeout options, but actual rendering time varies by page.
Viewport versus full page
A regular screenshot records the chosen viewport, which is useful for consistent visual comparisons. With Playwright’s fullPage: true, the screenshot spans the full scrollable document. Full-page images can become very tall and consume more memory and disk space; pages with sticky elements, lazy loading, or infinite scroll may not represent a simple full-document view. Playwright documents full-page capture, but lazy content behavior remains site-dependent.
File naming and reruns
Stable output names help compare reruns and prevent collisions. The example hashes the full URL, including query string. If query parameters contain secrets or volatile values, avoid exposing them in logs or input files and consider defining a sanitized identity for the capture. For versioned archives, place each run in a timestamped directory rather than overwriting prior images.
4. Python and cURL alternatives
Playwright’s official examples and APIs provide the browser primitives used in the Node.js version. If your existing pipeline is in Python, the following compact equivalent uses Playwright’s Python API. Install its package and browser first:
python -m pip install playwright
python -m playwright install chromium
from pathlib import Path
from urllib.parse import urlparse
from hashlib import sha256
import json
from playwright.sync_api import sync_playwright
urls = [line.strip() for line in Path('urls.txt').read_text().splitlines() if line.strip()]
out = Path('shots')
out.mkdir(exist_ok=True)
blocked_hosts = {'doubleclick.net', 'googlesyndication.com', 'google-analytics.com', 'googletagmanager.com'}
results = []
def should_block(request_url):
try:
host = urlparse(request_url).hostname or ''
return any(host == domain or host.endswith('.' + domain) for domain in blocked_hosts)
except Exception:
return False
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(viewport={'width': 1440, 'height': 900}, service_workers='block')
for original in urls:
try:
parsed = urlparse(original)
if parsed.scheme not in ('http', 'https') or not parsed.hostname:
raise ValueError('Expected an absolute HTTP or HTTPS URL')
page = context.new_page()
blocked = {'count': 0}
def route_request(route):
if should_block(route.request.url):
blocked['count'] += 1
route.abort()
else:
route.continue_()
page.route('**/*', route_request)
response = page.goto(original, wait_until='domcontentloaded', timeout=30000)
page.wait_for_timeout(1000)
name = parsed.hostname.replace('/', '_') + '-' + sha256(original.encode()).hexdigest()[:12] + '.png'
path = out / name
page.screenshot(path=str(path), full_page=False)
results.append({'input': original, 'status': 'captured', 'httpStatus': response.status if response else None,
'blockedRequests': blocked['count'], 'file': str(path)})
page.close()
except Exception as exc:
results.append({'input': original, 'status': 'error', 'error': str(exc)})
browser.close()
Path('shots/results.json').write_text(json.dumps(results, indent=2))
This is a sequential script intended to make each URL’s outcome easy to audit. Add bounded concurrency only after you have measured the memory and network impact in your environment.
cURL itself does not render web pages or create screenshots. It can call a browser-rendering API that offers a screenshot endpoint. A managed endpoint may accept one URL per request, so a shell loop can orchestrate a list when its documented API supports the needed capture options. Do not assume an endpoint accepts a bulk list just because it can snapshot a single page. Cloudflare documents a Browser Run /snapshot endpoint that can return rendered content and a screenshot in one request; consult its current documentation for authentication, limits, formats, and pricing before using it in a production batch. The cited endpoint documentation does not establish bulk-list input.
5. Run batches reliably
- Validate inputs. Normalize and validate schemes, remove accidental blank lines, and decide how duplicate URLs should be treated.
- Choose a stable environment. Pin the Playwright package and browser version for repeatable captures. Fix viewport, locale, timezone, and other context settings that matter to the pages you capture.
- Install filtering before navigation. Ensure routing is active before the first page request. Set the service-worker policy explicitly and understand its effect on the target sites.
- Set a timeout and readiness rule. Record the timeout and wait condition. Prefer a site-specific selector when a generic delay is unreliable.
- Keep an outcome record. Store one status per input, including errors, response status, final URL, output path, and relevant filter settings.
- Review samples after rule changes. Compare representative pages to detect missing styles, broken scripts, consent overlays, or unexpectedly blank captures.
The examples run one page at a time. This is slower than parallel capture but limits simultaneous browser work and makes failures straightforward to associate with inputs. If you add concurrency, use a small configurable worker pool, create isolated pages per task, and keep one shared browser/context only when shared cookies and storage are acceptable. For independent sessions, create separate contexts. A crashed worker should not erase completed output or the record of remaining URLs.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Many ads or analytics requests still appear | The patterns do not cover those request hosts, or traffic uses first-party endpoints or service workers. | Inspect request hosts, update and review your rules, and confirm the context’s service-worker setting. Do not treat the rules as complete tracking prevention. |
| Page layout is broken or images are missing | A rule blocked a required script, stylesheet, image, font, or shared CDN host. | Log blocked URLs and resource types, narrow the rules, and rerun a representative page. |
| Navigation times out | The page is slow, waits on ongoing traffic, or the timeout is too short for its load condition. | Choose a more suitable wait condition, increase the timeout for known slow pages, or wait for a page-specific selector. Keep timeout failures in the report. |
| Screenshot misses dynamic content | Navigation completed before client-side rendering or delayed content finished. | Wait for a meaningful selector or adjust the settling delay. Record the readiness rule so reruns are comparable. |
| Initial requests are not filtered | The route was registered after navigation, or a service worker handled the request. | Register page.route() before goto() and configure the context to block service workers where appropriate. |
| Two captures overwrite each other | The filename uses only a hostname or another non-unique field. | Include a digest of the full URL or a unique input index, as in the example. |
| Browser launch fails in a clean environment | The Playwright package is installed but its Chromium browser or system dependencies are missing. | Run Playwright’s browser installation command and follow the official installation guidance for the operating system. |
| Full-page capture is unexpectedly huge or slow | The page is very long, has infinite scrolling, or contains large images. | Prefer viewport screenshots for comparisons, or define a site-specific scrolling and stopping policy before full-page capture. |
7. Performance, reliability, and cost
Self-hosted Playwright has no per-screenshot API charge in the code shown, but it uses your compute, network, storage, and engineering time. Browser installation and upgrades, rule maintenance, retries, output retention, and monitoring are operational costs. Capture time depends on the target sites, chosen wait behavior, image sizes, and available resources; no throughput figure is implied here.
Sequential processing provides a simple baseline. Increase concurrency carefully: each active page adds browser work and network traffic, and a large pool can make captures slower or less reliable by exhausting memory, file descriptors, or target-site capacity. Use bounded retries for transient network errors, with a maximum attempt count and backoff. Avoid retrying permanent errors such as invalid input without correcting the URL. Preserve successful images and report each attempt so reruns can target failures rather than recapture the entire list.
For large or scheduled batches, consider whether you want to operate a browser fleet or call a managed rendering service. Compare authentication requirements, documented rate limits, supported formats, data handling, and current pricing directly in provider documentation. Cloudflare’s Browser Run snapshot endpoint is one documented managed option for a rendered page and screenshot; the documentation cited here does not establish a multi-URL batch call or a price comparison.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its bulk capture option accepts up to 100 URLs per call, and its API supports request blocking, selected resource types, custom headers and cookies, wait conditions, full-page capture, and other capture settings. See the API documentation for current parameter names and usage.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed, and known newsletter popups and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are never billed. Responses identify the page verdict and billing status in headers.
- An MCP server lets AI agents using Claude, Cursor, or another MCP client take screenshots, get page information, and capture PDFs.
- 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
FAQ
Can a screenshot script block every tracker?
No. It can abort requests matched by configured rules, but uncovered endpoints and tracking techniques may remain. Review requests and describe the filtering scope accurately.
Should every URL use the same wait condition?
Only if the pages have similar readiness behavior. A common baseline is convenient for batch work, while dynamic sites often need selector-based readiness or a page-specific delay.
Does cURL take a screenshot by itself?
No. cURL sends HTTP requests; screenshot capture requires a browser or a browser-rendering API endpoint.
When should I use a full-page image?
Use it when the entire scrollable document matters, such as a page review or archive. Use a fixed viewport when consistent screen-sized comparisons are the goal.
Sources
- Playwright Page API: navigation, screenshots, routing, and service-worker routing limitations.
- Playwright screenshots guide: viewport and full-page captures.
- Cloudflare Browser Run documentation: the snapshot endpoint and capture options.


