How to Bulk Capture Screenshots After Dismissing Cookie Banners
Use Playwright to handle consent banners and capture a URL list reliably, with verified dismissals, clear filenames, and useful failure logs.
To bulk capture pages without cookie banners covering the content, use a browser automation script that reads a URL list, opens each page, handles that site’s consent interface, confirms the banner is gone, waits for a meaningful page-specific signal, and saves a screenshot with a stable filename. There is no universal consent selector: the button, persistence behavior, and choice offered vary by site.
This guide uses Playwright with Node.js. It captures pages sequentially, logs failures per URL, and supports viewport, full-page, and element screenshots. Use a consent action that matches your task and the site’s choices; dismissing a visual overlay is different from recording a consent preference.
1. Set up Playwright and prepare the URL list
Install Node.js, create a project, and install Playwright. The install command downloads the browser binaries Playwright needs.
mkdir bulk-page-capture
cd bulk-page-capture
npm init -y
npm install playwright
npx playwright install chromium
Create urls.txt with one URL on each line:
https://example.com/
https://example.org/news/
https://example.net/pricing/
Use absolute URLs, including the scheme. Keep this list under version control if appropriate, but do not commit browser state files containing cookies or other reusable credentials.
2. Write a batch capture script
Save this as capture.mjs. The consent locator is deliberately site-specific: change the accessible button name for your target, or replace it with a locator for that site’s consent dialog. The script will not take a screenshot for a page if the configured consent button was found but the banner remains visible after the click.
import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const mode = process.env.CAPTURE_MODE ?? 'full'; // full, viewport, or element
const readySelector = process.env.READY_SELECTOR; // optional, e.g. main h1
const consentButtonName = process.env.CONSENT_BUTTON; // exact accessible button name
const consentDialogSelector = process.env.CONSENT_DIALOG; // optional dialog/banner selector
const elementSelector = process.env.ELEMENT_SELECTOR; // required for element mode
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 30000);
function safeName(rawUrl, index) {
const u = new URL(rawUrl);
const suffix = u.pathname.split('/').filter(Boolean).join('-') || 'home';
const stem = `${String(index).padStart(3, '0')}-${u.hostname}-${suffix}`
.toLowerCase().replace(/[^a-z0-9.-]+/g, '-').replace(/-+/g, '-').slice(0, 150);
return `${stem}.png`;
}
const urls = (await readFile(inputFile, 'utf8'))
.split(/\r?\n/).map(line => line.trim())
.filter(line => line && !line.startsWith('#'));
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width: 1440, height: 1000 }, deviceScaleFactor: 1 });
const results = [];
try {
for (let i = 0; i < urls.length; i++) {
const url = urls[i];
const page = await context.newPage();
let result = { url, status: null, consent: 'not-configured', file: null, error: null };
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
result.status = response?.status() ?? null;
if (response && response.status() >= 400) {
throw new Error(`HTTP ${response.status()}`);
}
// Prefer an application-specific readiness signal when one is known.
if (readySelector) await page.locator(readySelector).waitFor({ state: 'visible', timeout: timeoutMs });
if (consentButtonName) {
const button = page.getByRole('button', { name: consentButtonName, exact: true });
if (await button.isVisible().catch(() => false)) {
await button.click({ timeout: 5000 });
if (consentDialogSelector) {
await page.locator(consentDialogSelector).waitFor({ state: 'hidden', timeout: 5000 });
}
result.consent = 'button-clicked';
} else {
result.consent = 'button-not-visible';
}
}
if (consentDialogSelector) {
const banner = page.locator(consentDialogSelector);
if (await banner.isVisible().catch(() => false)) {
throw new Error(`Consent UI still visible: ${consentDialogSelector}`);
}
}
const file = path.join(outputDir, safeName(url, i + 1));
if (mode === 'element') {
if (!elementSelector) throw new Error('Set ELEMENT_SELECTOR when CAPTURE_MODE=element');
await page.locator(elementSelector).screenshot({ path: file, type: 'png', timeout: timeoutMs });
} else {
await page.screenshot({ path: file, type: 'png', fullPage: mode === 'full', timeout: timeoutMs });
}
result.file = file;
} catch (error) {
result.error = error instanceof Error ? error.message : String(error);
} finally {
await page.close();
results.push(result);
console.log(JSON.stringify(result));
}
}
} finally {
await browser.close();
await writeFile(path.join(outputDir, 'results.json'), JSON.stringify(results, null, 2));
}
if (results.some(item => item.error)) process.exitCode = 1;
Run it with a consent button name that actually appears on your site. For example, if the site’s button is named “Accept all cookies”:
CONSENT_BUTTON='Accept all cookies' CONSENT_DIALOG='[role="dialog"]' READY_SELECTOR='main' node capture.mjs urls.txt screenshots
If the dialog is not exposed with role dialog, set CONSENT_DIALOG to a selector based on the inspected page markup. If your task is to close the overlay without making a consent choice, configure the site’s close control instead. Do not use an “accept all” action unless that is the intended choice.
3. Choose how to wait for each page
The script navigates with domcontentloaded, then optionally waits for READY_SELECTOR. For sites with client-rendered content, use a meaningful visible selector such as a page heading, article container, or chart. Playwright documents load and domcontentloaded navigation conditions, and discourages networkidle as a generic readiness condition because pages may keep network requests active. A page-specific check is usually a better signal for the content you need. [Playwright Page API]
When a page has no stable selector, you can add a fixed delay after navigation, but delays slow every capture and cannot guarantee readiness across variable page loads. Avoid treating a successful navigation as proof that the important content loaded.
4. Handle cookie banners accurately
Inspect the actual page and target a named button, link, or control for that site’s banner. Playwright’s role-based locators make the intended control more understandable and robust than guessing a CSS class, but accessible names differ among sites. Always verify the banner has disappeared before capturing.
- Accept: use the site’s accept control only when accepting is the correct choice for the task.
- Reject: if the site offers a reject option and that is the desired preference, configure its actual accessible name instead.
- Close: use a close control only if you intend to dismiss the visual overlay without recording a broader choice.
- Settings: a settings button may open another panel rather than remove the banner. Handle and verify the resulting state explicitly.
Consent choices may persist in cookies or local storage, but behavior varies by site. A shared browser context can preserve state for pages on the same origin. To reuse state across runs, Playwright can save and load browser storage state, including cookies and local storage; this does not establish that every site’s consent system uses those mechanisms. Treat the state file as sensitive because it may contain credentials or cookies usable to impersonate a session, and keep it out of source control. [Playwright authentication and storage state]
5. Choose viewport, full-page, or element capture
| Capture | Use it for | Trade-off |
|---|---|---|
| Viewport | Consistent first-screen comparisons and visual checks | Content outside the visible area is omitted |
| Full page | Archiving or reviewing the entire scrollable document | Very long pages create large images and may expose lazy-loading or layout issues |
| Element | A chart, card, article, or other specific region | Requires a selector that resolves to the intended element |
Playwright supports screenshot paths, image formats, viewport and full-page screenshots, locator screenshots, and scale choices. With CSS scale, one output pixel corresponds to one CSS pixel; device scale uses device-pixel density and can produce larger files. [Playwright screenshots]
Set CAPTURE_MODE=viewport for the initial viewport, or CAPTURE_MODE=element ELEMENT_SELECTOR='article' for a target element. The script writes PNGs. Change type and file extensions together if you choose JPEG or another supported format. Keep viewport dimensions and device scale fixed across a batch if visual comparisons matter.
6. Scale the batch without losing track of failures
The example processes URLs sequentially. That is a conservative starting point: it limits pressure on the target sites and makes failures easier to diagnose. If throughput is important, add bounded concurrency only after observing how the sites and your machine respond. Do not infer a universal safe batch size from another implementation’s defaults.
Each URL gets its own page and result record. The JSON log records the URL, HTTP response status when available, consent handling status, output path, and error. A 404 or 500 response may be returned as a navigation response rather than thrown as a navigation exception, so check the status explicitly. The script treats HTTP status 400 or higher as a failed capture and continues with the next URL.
For repeated runs, include a run date or identifier in the output directory or filename. The generated names include a stable index, hostname, and path. Review results.json and confirm expected files exist before using the batch downstream.
7. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Banner remains in the screenshot | The configured name or selector did not match, the click missed, or settings opened instead | Inspect the accessible structure, target the actual control, and wait for the banner selector to become hidden |
| Consent click times out | The button is covered, disabled, or not yet rendered | Wait for the correct dialog/control, check overlays and frame boundaries, then use the site’s actual locator |
| Consent appears on every URL | Choices are origin-specific or not persisted in the current context | Check the site’s storage behavior; reuse a context or saved storage state only when appropriate |
| Screenshot is blank or incomplete | Navigation completed before app content rendered, or the page returned an error | Check the logged status; wait for a page-specific visible selector and inspect lazy-loaded content |
| Full-page image is unexpectedly huge | The document is very long or contains large embedded content | Use viewport or element mode if that matches the capture goal |
| Some URLs fail while others work | Timeouts, access controls, invalid URLs, or site-specific behavior differ | Use per-URL logs, increase timeout selectively, and respect site terms and access controls |
| State file exposes a session | Browser state contains reusable cookies or headers | Restrict file access, exclude it from version control, and rotate affected credentials if exposed |
8. Performance, reliability, and cost
Self-managed Playwright has no per-screenshot API fee, but you supply the browser runtime and maintain selectors, browser versions, retries, storage, and output handling. Browser memory, full-page dimensions, page scripts, and network load affect how many pages you can process at once. Begin sequentially, then increase concurrency gradually while tracking timeouts and incomplete output.
For reliability, isolate errors per URL, set a finite navigation and screenshot timeout, inspect status codes, verify consent dismissal, and use a page-specific readiness condition. Retry only failures that are plausibly transient; repeating a blocked or invalid request can waste time and may violate access controls. No single wait condition or consent selector works for every site.
If you prefer a managed endpoint, ScreenshotNeo is a website screenshot API and MCP server. Its API supports bulk capture of up to 100 URLs per call. Pricing is Free for 1,000 shots per month with no card, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. See the ScreenshotNeo documentation for request options and details.
Or skip the browser setup
Send one GET request per page to ScreenshotNeo, or use its bulk capture option for a URL list. This cURL example captures one page; replace the URL or use the documented bulk request format for batches.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify page verdict and billing status in headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Review the API documentation for options and setup, then sign up for 1,000 free screenshots a month with no card.
FAQ
Will one cookie selector work across all sites?
No. Consent interfaces and accessible names differ. Configure and verify the control for each site or site group.
Does dismissing a banner always save the choice?
No. Persistence depends on the site’s implementation and may be limited to an origin or session.
Should I use network idle before taking every screenshot?
No. Prefer a page-specific ready signal; pages with ongoing requests may never become network idle.
Can I use this for hundreds of pages?
Yes. Keep the input list manageable, start sequentially, log results, and introduce bounded concurrency only when the sites and runtime handle it reliably.


