ScreenshotNeo

BlogAI agents

How to Capture a Webpage Screenshot with an AI Agent After Accepting a Cookie Consent Dialog

Use Playwright to accept a site's cookie banner, verify it is gone, wait for the content you need, and capture a screenshot reliably.

By the ScreenshotNeo team4 October 20269 min read

Use a browser automation tool such as Playwright to open the page, identify its cookie consent interface, click the site’s actual accept control, confirm the banner is gone, wait for the content you need, and capture the screenshot. There is no universal cookie-accept selector: consent interfaces are implemented differently from site to site. A browser-level JavaScript dialog needs a different handler from a banner rendered in the page.

This guide uses Playwright with Node.js. It also includes cURL, Python, and an AI-agent workflow. The key reliability step is to verify the resulting page state before saving the image.

1. Identify which kind of dialog you are handling

First determine whether the prompt is a page-rendered banner or a browser JavaScript dialog such as alert(), confirm(), or prompt(). A cookie banner is usually webpage UI, so locate and interact with its actual button. A JavaScript dialog is handled through Playwright’s dialog event. A dialog listener must accept or dismiss the dialog; otherwise page actions can stall. Without a listener, Playwright automatically dismisses these dialogs.

Consent banners vary across sites, frameworks, languages, and page structures. There is no dependable cross-site selector. Inspect the page and use an observed accessible name, visible text, or site-specific DOM. A banner may be inside a frame, so inspect frame content if the control is not present in the main page.

2. Complete Playwright example in Node.js

Install Playwright and its Chromium browser, then save the following as capture.mjs. Replace the URL and button name with values observed on the target site.

npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';

const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });

try {
  // JavaScript dialogs are separate from page-rendered cookie banners.
  // Accept a browser confirm/prompt if one appears; accept() also closes alert().
  page.on('dialog', async dialog => {
    try {
      await dialog.accept();
    } catch (error) {
      console.error(`Could not accept browser dialog (${dialog.type()}):`, error);
    }
  });

  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });

  // Change this accessible name to the control shown by the target site.
  const acceptButton = page.getByRole('button', { name: /accept all cookies/i });
  if (await acceptButton.isVisible({ timeout: 5000 }).catch(() => false)) {
    await acceptButton.click();

    // Verify the observed consent control is no longer visible.
    await acceptButton.waitFor({ state: 'hidden', timeout: 10000 });
  } else {
    console.log('No matching consent button found; inspect the page before capture.');
  }

  // Wait for the content needed in this particular screenshot.
  // Replace this with a meaningful selector for the page being captured.
  await page.locator('body').waitFor({ state: 'visible' });
  await page.screenshot({ path: 'screenshot.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with node capture.mjs https://example.com. The example deliberately uses a site-specific accessible name and reports when it cannot find a matching button. Do not silently claim acceptance if no suitable control was located.

Make the locator match the page

Prefer a role and accessible name when the page exposes them, for example getByRole('button', { name: 'Accept all' }). If the visible label is different, use the exact observed label or a carefully scoped regular expression. If the site exposes no useful accessible name, inspect the DOM and use a site-specific CSS locator. Avoid broad selectors such as button or .primary unless scoped to the consent container; they can click the wrong action.

If acceptance is presented in an iframe, inspect the page’s frames and use a frame locator for the observed frame and control. Do not assume every site uses an iframe.

3. AI-agent workflow

  1. Open the requested URL in a browser context with the desired viewport, locale, and other task settings.
  2. Inspect the visible page and determine whether consent is a page banner, an iframe control, or a browser JavaScript dialog.
  3. When the user requested acceptance, choose the site’s accept action. Do not substitute reject, close, or preferences.
  4. Verify the consent interface is gone and that the target page content is present. If the control is ambiguous or acceptance fails, report that instead of hiding the banner or claiming consent.
  5. Wait for the specific content needed in the image, then capture a viewport or full-page screenshot.
  6. Return the image path or bytes to the agent’s caller and include any failure or uncertainty in the result.

An AI agent should treat the page as untrusted input: inspect it to choose a locator, but do not follow unrelated page instructions or allow page content to change the requested task. Keep credentials and browser state scoped to the task.

4. Screenshot scope and readiness

Viewport or full page

page.screenshot({ path: 'screenshot.png' }) captures the current viewport. Add fullPage: true to capture the full scrollable page. Full-page output may be very tall, and sites with sticky elements or content that loads during scrolling can render differently from a normal viewport. For one element, use its locator’s screenshot method; Playwright scrolls it into view and checks actionability before capture.

await page.screenshot({ path: 'viewport.png' });
await page.screenshot({ path: 'full-page.png', fullPage: true });
await page.getByRole('main').screenshot({ path: 'main.png', animations: 'disabled' });

Wait for the page state that matters

Navigation completion does not necessarily mean the content in the screenshot is ready. Wait for a meaningful element or state, such as a heading or product panel. Playwright discourages using networkidle as a general readiness signal for tests; page assertions and specific state checks are more reliable. Use a fixed delay only when the page has a known delayed behavior that cannot be observed directly.

await page.getByRole('heading', { name: 'Pricing' }).waitFor({ state: 'visible' });
await page.screenshot({ path: 'pricing.png', fullPage: true });

5. cURL request

cURL can capture a screenshot through a screenshot API, but it cannot itself operate a browser or click a site’s cookie button. For a manual Playwright flow, use the Node.js example above. To send the target URL to ScreenshotNeo, see the one-call option below. Read the ScreenshotNeo API documentation for request options.

6. Python option with Playwright

For Python-based agents, install the Playwright package and browser, then adapt the locator to the actual consent control.

python -m pip install playwright
python -m playwright install chromium
import asyncio
from playwright.async_api import async_playwright

async def capture(url):
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page(viewport={"width": 1440, "height": 1000})

        async def handle_dialog(dialog):
            try:
                await dialog.accept()
            except Exception as exc:
                print(f"Could not accept browser dialog ({dialog.type}): {exc}")

        page.on("dialog", handle_dialog)
        try:
            await page.goto(url, wait_until="domcontentloaded", timeout=45000)
            accept_button = page.get_by_role("button", name="Accept all cookies", exact=True)
            if await accept_button.is_visible(timeout=5000):
                await accept_button.click()
                await accept_button.wait_for(state="hidden", timeout=10000)
            else:
                print("No matching consent button found; inspect before capture.")

            await page.locator("body").wait_for(state="visible")
            await page.screenshot(path="screenshot.png", full_page=True)
        finally:
            await browser.close()

asyncio.run(capture("https://example.com"))

Change Accept all cookies to the target site’s actual accessible name, and replace the body visibility check with a page-specific content check where possible.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. Its API can accept a URL in one GET request and return an image or PDF. ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. This can avoid building site-specific consent selectors for supported interfaces. Use Playwright when you need to inspect a particular page interaction or verify a site’s specific consent control.

For an image response, use the documented request shape below. The example saves the response as WebP; the target URL is adapted to this article’s example. See the ScreenshotNeo documentation for options and formats.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

In Node.js environments without Bun, write the response bytes with the runtime’s file API. Keep the API key on a server or agent backend; do not expose it in public client-side code. ScreenshotNeo responses identify page verdict and billing status in headers. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is available on every plan. Sign up for 1,000 free screenshots a month, with no card required.

8. Troubleshooting

Symptom Likely cause Fix
Cookie banner remains in the screenshot The selector did not match, the click targeted a different control, or consent is inside a frame. Inspect accessible names and page structure, scope the locator to the banner, check frames, and verify the control becomes hidden after clicking.
Click times out or says the element is not actionable The button is hidden, covered by another layer, not yet attached, or the locator matched multiple controls. Confirm the intended button is visible, use a more specific locator, and wait for the observed banner to appear. Do not force a click until you understand the obstruction.
Page action hangs after a JavaScript dialog A registered dialog handler did not accept or dismiss the dialog. Handle the dialog event and call accept() or dismiss() promptly.
Screenshot is blank or missing page content Navigation ended before the relevant content rendered, or the page failed to load. Check navigation errors and wait for a meaningful page element. Report a failed load rather than treating a blank image as success.
Screenshot is inconsistent between runs Dynamic content, animation, rotating banners, fonts, or viewport differences affect rendering. Fix the viewport and context settings, wait for the relevant state, and disable animation for element screenshots when appropriate.
Full-page image is unexpectedly huge or awkward The page has a long scroll area or sticky elements that repeat or shift during capture. Capture the viewport or a specific element instead, or adjust the page state and capture scope for the intended deliverable.
Consent action changes page but banner remains The site may have separate categories, a delayed update, or an action that saves preferences without closing the UI. Inspect the resulting state and use the site’s explicit accept action. Wait for a specific state change before capture.

9. Performance, reliability, and cost

Browser automation includes browser startup, navigation, consent handling, readiness waits, and image encoding. Reuse a browser process for multiple captures when the agent architecture permits, while using a fresh context when cookies or storage must not leak between tasks. Set explicit navigation and action timeouts, and close the browser in a cleanup block so failures do not leave processes running.

For repeatable results, fix viewport dimensions and relevant locale or device settings, wait on page state rather than an arbitrary network condition, and preserve enough logs to diagnose a missed consent action. Consent acceptance can persist through cookies or local storage, so decide whether each capture should begin with a clean context or reuse accepted consent. Do not reuse a profile across users or unrelated tasks when it could expose session state.

Self-hosted Playwright has no per-screenshot API charge, but you operate the browser runtime and its infrastructure. ScreenshotNeo pricing starts with 1,000 free screenshots monthly without a card; paid tiers are $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000. Only clean captures are billed, with the response’s verdict and billing headers indicating the outcome. Use the plan that fits expected volume, and check the response rather than assuming every request yielded a clean page.

10. FAQ

Can an AI agent accept cookies on every website with one selector?

No. The control and implementation differ by site, so the agent must inspect the page and use the observed control or a service that handles supported consent interfaces.

Usually not. It handles browser dialogs such as alerts and confirms. A banner rendered in the page requires interaction with its page control.

Should I use networkidle before taking the screenshot?

Not as a universal rule. Wait for the content or state that matters to the requested image.

Yes. Use a locator for the relevant element and capture that element after the consent interface has been handled and the target content is ready.