How to Block Cookie Modals When Scraping
Learn when to dismiss, persist, or remove cookie modals in Playwright and Selenium, with reliable code, troubleshooting, and a browser-free API option.
Direct answer: Treat a cookie obstruction according to what it is. Dismiss native JavaScript dialogs with the browser dialog API, click a visible DOM banner using a site-specific selector, or load a previously recorded consent state before navigation. Do not assume that blocking cookies or deleting the overlay proves consent was rejected. For repeatable scraping, make an explicit choice such as “Reject all,” preserve the resulting cookies or storage state, and reuse that state in an isolated browser context.
This guide shows how to handle cookie modals in Playwright and Selenium, including banners inside iframes, delayed overlays, persisted consent, JavaScript hooks, screenshots, and failure diagnosis. It also covers a browser-free approach with ScreenshotNeo when your goal is a clean page image rather than DOM extraction.
1. Identify the obstruction first
Cookie interfaces are not one mechanism. Your first step determines the correct fix.
| What you see | Correct approach | Typical symptom |
|---|---|---|
Native alert, confirm, prompt, or beforeunload |
Register a dialog handler and accept or dismiss it | An action hangs until the dialog is handled |
| HTML banner or modal in the page DOM | Locate the visible reject, close, or necessary-only control and click it | Content is present but covered by an overlay |
| Consent stored in cookies or local storage | Reuse a captured storage state or set site-specific consent data before navigation | The banner returns in every fresh context |
| Banner inside an iframe | Find the frame, then locate controls within that frame | Main-page selectors find zero buttons |
| Delayed or intermittent overlay | Wait for it explicitly or use Playwright’s locator handler | Some runs fail because the modal appears after page load |
2. Handle native JavaScript dialogs in Playwright
Playwright auto-dismisses JavaScript dialogs by default. If you add a listener, the listener must resolve every dialog or the action that triggered it can stall. The following Node.js script dismisses dialogs, visits a page, and saves the resulting HTML.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
page.on('dialog', async dialog => {
console.log(`Dialog: ${dialog.type()} — ${dialog.message()}`);
await dialog.dismiss();
});
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.screenshot({ path: 'page.png', fullPage: true });
await browser.close();
Use dialog.accept() only when accepting is the documented choice for that target. A native dialog is different from a cookie banner rendered as HTML; the dialog event will not find a DOM modal.
3. Dismiss a DOM cookie banner reliably
Prefer semantic selectors such as an accessible role and name, then stable attributes supplied by the site. Avoid brittle selectors based on generated class names. Wait for the banner before clicking so a slow consent script does not race your scraper.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
const banner = page.getByRole('dialog');
if (await banner.count()) {
const reject = banner.getByRole('button', {
name: /reject|deny|necessary only|decline/i
});
if (await reject.count()) {
await reject.first().click();
}
}
await page.screenshot({ path: 'after-consent.png', fullPage: true });
await browser.close();
Some sites use a banner rather than a dialog role. In that case, scope the locator to a stable container such as [data-testid='cookie-banner'], an identified consent-management element, or text that your site documentation confirms. A selector is always site-specific; there is no universal cookie-button selector.
Handle delayed overlays with a locator handler
Playwright’s locator handler can run when an unexpected overlay blocks an action or assertion. Keep the handler self-contained because it may change focus and mouse state. After it runs, perform a fresh action rather than assuming focus is unchanged.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
await page.addLocatorHandler(page.getByRole('dialog'), async dialog => {
const reject = dialog.getByRole('button', {
name: /reject|necessary only|deny/i
});
if (await reject.count()) await reject.first().click();
});
await page.goto('https://example.com', { waitUntil: 'networkidle' });
await page.getByRole('main').screenshot({ path: 'main.png' });
await browser.close();
4. Persist an explicit consent decision
A consent-management platform commonly checks cookies or local storage. If no valid preference exists, it displays a notice, records your response, and passes consent status to tag and advertising systems. For repeatable runs, create an isolated context, make an explicit choice, save the resulting storage state, and reuse it.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
const reject = page.getByRole('button', { name: /reject|necessary only/i });
if (await reject.count()) await reject.first().click();
await context.storageState({ path: 'consent-state.json' });
await browser.close();
Use the saved state on later runs:
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext({ storageState: 'consent-state.json' });
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.screenshot({ path: 'repeatable.png', fullPage: true });
await browser.close();
Cookie domain, path, expiry, value, and consent semantics differ by site. Copying a value from one domain to another can be ineffective or misleading. Record which choice the state represents and refresh it when the site changes its consent implementation.
Set a cookie before navigation
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext();
await context.addCookies([{
name: 'site_consent',
value: 'reject',
domain: 'example.com',
path: '/',
expires: Math.floor(Date.now() / 1000) + 86400
}]);
const page = await context.newPage();
await page.goto('https://example.com');
await browser.close();
The cookie name and value above are placeholders. Inspect the site’s own response after clicking its control and use only a documented, permitted value.
Run a site-specific init script
addInitScript runs before page scripts. It is useful when a target site reads a known local-storage key or exposes a documented consent hook.
await context.addInitScript(() => {
localStorage.setItem('example_consent', JSON.stringify({ necessary: true }));
});
Do not use an invented key and assume it means consent was recorded. Confirm the site’s behavior and preserve an audit note describing the choice.
5. Cookie banners inside iframes
Query the frame before looking for buttons. A main-page locator cannot see elements inside a cross-document iframe.
const frame = page.frameLocator('iframe[title*="consent" i]');
const reject = frame.getByRole('button', { name: /reject|deny|necessary only/i });
if (await reject.count()) await reject.first().click();
If the iframe has no stable title or attribute, inspect frame URLs and choose a selector maintained for that site. Cross-origin frames can still be automated through Playwright’s frame locator, but browser security rules prevent arbitrary DOM access from page JavaScript.
6. Selenium alternative
Selenium WebDriver can drive Chrome, Firefox, and other browsers, but the browser and driver major versions must match. Selenium has no universal cookie-banner selector, so the locator remains target-specific.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
options = webdriver.ChromeOptions()
driver = webdriver.Chrome(options=options)
try:
driver.get('https://example.com')
wait = WebDriverWait(driver, 10)
reject = wait.until(EC.element_to_be_clickable((
By.XPATH,
"//button[contains(translate(., 'REJECT', 'reject'), 'reject') or contains(translate(., 'DENY', 'deny'), 'deny')]"
)))
reject.click()
driver.save_screenshot('after-consent.png')
finally:
driver.quit()
For a known cookie, set it after first opening the domain, then refresh:
driver.get('https://example.com')
driver.add_cookie({
'name': 'site_consent',
'value': 'reject',
'path': '/',
'domain': 'example.com'
})
driver.refresh()
7. Why blocking cookies alone is unreliable
Browser cookie blocking is a privacy or diagnostic setting, not a general modal-removal strategy. Firefox and Safari apply stronger third-party-cookie protections than Chrome in many default configurations, while Chrome does not block third-party cookies by default outside Incognito or explicit settings. Blocking cookies can also break embedded components and normal page functionality. A banner may still render from local storage, an iframe, a server response, or a delayed script.
Deleting the overlay with JavaScript only changes what is visible. It does not prove that consent was rejected, stop tracking code, or prevent a CMP from restoring the modal. Use DOM removal only for a controlled visual task such as taking a screenshot, and document that it is not a consent decision.
8. A practical decision checklist
- Classify the obstruction as native dialog, DOM element, iframe, or persisted state.
- Choose an explicit policy: reject all, necessary only, or another permitted option.
- Prefer a semantic, stable selector and scope it to the banner.
- Wait for delayed banners and handle them before the extraction action.
- Save storage state for repeatable sessions.
- Verify that the expected content is visible after dismissal.
- Keep consent state isolated per domain and environment.
- Respect the site’s terms, robots directives, privacy law, and permitted-access policy.
9. Troubleshooting common failures
| Error or symptom | Cause | Fix |
|---|---|---|
| Timeout while clicking a button | The banner is delayed, hidden, or covered by another layer | Wait for visibility, scope the locator to the banner, and inspect an iframe. |
| Locator finds zero elements | Wrong role or selector, changed copy, or iframe isolation | Inspect the DOM and frame list; use a stable attribute or accessible name. |
| Page hangs after a click | A native dialog remains unhandled | Register a page.on('dialog') listener that accepts or dismisses it. |
| Banner returns every run | Context is new or consent state was not saved | Save and reuse storageState, or set the documented cookie before navigation. |
| Content is still blocked after removing the modal | The CMP also gates content or requests behind consent | Make the permitted consent choice through the site control and wait for dependent requests. |
| Selenium session will not start | Browser and driver major versions differ | Install a matching driver or use Selenium Manager with a supported browser. |
| Clicks fail after a locator handler | The handler changed focus or mouse state | Use a self-contained handler and issue a fresh locator action afterward. |
10. Performance, reliability, and cost
Launching a browser is the expensive part of this workflow. Reuse a browser process while creating a fresh context per job, save consent state when policy permits, and wait for the smallest useful readiness condition rather than unconditionally using long sleeps. Network-idle waits can remain open on pages with analytics or streaming requests; a known selector or bounded delay is often more predictable.
For extraction, keep the DOM and response data instead of taking screenshots. For visual archives, restrict the viewport and disable unnecessary resources only when doing so does not change the page you need to capture. Cache consent state and avoid repeated banner interactions, but refresh it when the site changes its CMP.
11. Or skip the browser setup
When you need a clean screenshot or PDF instead of DOM-level scraping, ScreenshotNeo provides a single request. It accepts the consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
curl -G 'https://api.screenshotneo.com/v1/shot' \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo also supports full-page and element capture, dark mode, device presets, custom viewport and retina scale, waits, custom CSS and JavaScript, headers, cookies, user agents, authorization, blocking rules, caching, signed links, asynchronous jobs, bulk capture, PDFs, and HTML/CSS rendering. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
12. FAQ
Should I click Reject or delete the modal?
Click the site’s reject or necessary-only control when you are making a consent decision. Delete the element only for a controlled visual operation, and do not treat that as proof that tracking was disabled.
Can one cookie value work across every site?
No. Names, values, domains, paths, expiry rules, local-storage keys, and CMP behavior are site-specific.
Why does a banner appear after page load?
Consent scripts often load asynchronously or from an iframe. Wait for the banner, use a locator handler, or wait for a documented readiness selector.
Does cookie blocking stop third-party tracking?
It can change browser behavior but may break page features and does not reliably handle local storage, server-side state, or consent managers.
When should I use an API instead of Playwright or Selenium?
Use browser automation when you need DOM extraction and interactions. Use ScreenshotNeo when the output is a clean screenshot or PDF and you want consent overlays handled without maintaining browser infrastructure.


