ScreenshotNeo

BlogAI agents

How to Take a Website Screenshot with an AI Agent and Hide Popups

Use an AI browser agent to find and handle popups, capture the right part of a page, and verify the screenshot. Includes Playwright MCP and code examples.

By the ScreenshotNeo team4 October 202610 min read

To take a website screenshot with an AI agent and hide popups, have the agent open the page, inspect a fresh accessibility snapshot, and decide whether to dismiss the obstruction through a page control or hide it only in the captured image. Then capture the viewport, a specific element, or the full page and inspect the result. A screenshot shows appearance; use the agent’s page snapshot to find controls and understand structure.

This guide uses Playwright MCP for the browser-agent workflow, then shows equivalent Playwright code and a capture-time masking option. If you want a hosted screenshot API instead of managing a browser, see ScreenshotNeo.

1. Choose how to handle the popup

First identify what kind of popup is present and whether the page should actually change. This matters for consent banners: hiding a banner visually is not the same as accepting or rejecting consent.

What you see What it is Approach
Cookie or consent banner Ordinary page content If the task requires a real consent choice, use the appropriate visible control. For a visual-only capture, hide or mask the banner and state that this did not record a consent choice.
Signup overlay, modal, or newsletter popup Usually ordinary page content Use a clear close control if changing page state is acceptable, or hide/mask it for the screenshot only.
JavaScript alert, confirm, prompt, or beforeunload dialog Browser dialog, not a page element Handle it using Playwright’s dialog behavior; a CSS selector or screenshot mask will not work.
A new tab or window A separate browser page Listen for and handle the popup page if that is where the intended content appears.

Playwright’s guidance is to use browser_snapshot to get references for interaction and use screenshots for visual inspection. Its Page API also documents screenshot masks and styles, popup events, dialogs, and locator handlers for unexpected overlays. See the Playwright MCP screenshot guide and Playwright Page API.

2. Take a screenshot with Playwright MCP

Connect an AI agent to a Playwright MCP server and use its browser tools. Tool names can vary with the client and server version; the workflow below uses the documented browser navigation, snapshot, interaction, and screenshot concepts. Ask the agent to act on the current page, inspect the result after changes, and save the screenshot at a known path.

  1. Navigate to the URL.
  2. Wait for the content and overlays relevant to the capture to render.
  3. Read a fresh accessibility snapshot. Identify the obstruction and its close, accept, or reject controls.
  4. Choose whether to interact with the page or hide the element only during capture.
  5. After interaction, take another snapshot before using references or choosing a target. References from an old snapshot may no longer be valid.
  6. Capture the viewport, an element, or the full page, then inspect the saved image.

A useful agent instruction is: “Open the target URL, wait until the main content is visible, and inspect the accessibility snapshot. Identify any cookie banner, modal, or popup. Do not make a consent choice unless the task explicitly authorizes one. If an overlay has a clear close control and closing it is appropriate, use that control, refresh the snapshot, and capture the requested scope. Otherwise hide it only for capture. Save the screenshot and report which method you used.”

For a consent banner, choose the site’s actual consent control only when the task calls for accepting or rejecting cookies. If the goal is a clean visual, a capture-only hide or mask can avoid changing the page’s consent state.

3. Use Playwright code for repeatable captures

When you need reproducibility or more control than an agent prompt provides, use Playwright directly. Install it in a project that has Node.js available, then run the following script. It takes a screenshot of the viewport by default; set FULL_PAGE=1 for the full scrollable page. Set HIDE_SELECTOR to hide an overlay only while capturing. Replace the selector with one observed on the target page.

npm install playwright
npx playwright install chromium

# Save as screenshot.mjs
import { chromium } from 'playwright';

const url = process.env.TARGET_URL ?? 'https://example.com';
const hideSelector = process.env.HIDE_SELECTOR;
const fullPage = process.env.FULL_PAGE === '1';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });

try {
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
  await page.locator('body').waitFor({ state: 'visible', timeout: 15_000 });

  // Hiding is capture-only. It does not make a consent choice or remove page state.
  if (hideSelector) {
    await page.addStyleTag({ content: `${hideSelector} { visibility: hidden !important; }` });
  }

  await page.screenshot({ path: 'screenshot.png', fullPage });
  console.log('Saved screenshot.png');
} finally {
  await browser.close();
}

Run it with environment variables, for example TARGET_URL=https://example.com HIDE_SELECTOR='[role="dialog"]' node screenshot.mjs. A generic selector such as [role="dialog"] may match content you want to keep; inspect the page and use a specific selector. If the overlay animates or appears after navigation, wait for it or for the page’s relevant state before capture.

Dismiss an overlay with its visible control

If the task permits changing the page, prefer the actual close control over deleting arbitrary DOM. Locate it from the snapshot or inspect accessible names, click it, and then refresh the snapshot. This example uses a placeholder accessible name; replace it with the control observed on the page.

await page.getByRole('button', { name: 'Close', exact: true }).click();
// Re-check the page after it changes, then capture.
await page.screenshot({ path: 'after-close.png' });

Do not use a guessed “Accept” button just to clear a banner: accepting changes consent state. Also avoid broad scripts that remove all dialogs, fixed elements, or overlays, since they can delete intended content.

Mask an obstruction instead of dismissing it

A mask covers a locator’s bounding box in the resulting image; it does not dismiss or remove the underlying page element. A solid mask can leave a visible rectangle. A capture stylesheet can hide the matched element instead, though a broad or incorrect selector can hide wanted content. Playwright documents both mask and style screenshot options in its Page API.

await page.screenshot({
  path: 'masked.png',
  mask: [page.locator('.cookie-consent-banner')],
  maskColor: '#000000'
});

4. Select viewport, element, or full-page capture

  • Viewport: captures what is currently visible in the browser window. Use it for a quick visual check or a specific above-the-fold state.
  • Element: captures a specific chart, card, form, or other component. In Playwright, use the locator screenshot method, such as await page.locator('.chart').screenshot({ path: 'chart.png' }).
  • Full page: captures the scrollable document in one image with fullPage: true. It cannot be combined with a target element in the Playwright MCP screenshot tool; capture an element separately when that is the intended scope. Very long pages can produce large images and may trigger lazy content only as the page scrolls.

Choose the scope before troubleshooting a “missing” section. A viewport screenshot does not include below-the-fold content. For full-page captures, account for lazy-loaded images and content that appears only after scrolling or interaction.

5. Handle dialogs and popup tabs

Page overlays and browser dialogs need different treatment. Playwright automatically dismisses JavaScript dialogs when no dialog listener is registered. If you register a listener, it must accept or dismiss the dialog; otherwise actions can stall. For a popup tab or window, listen for the page’s popup event and capture that page if it contains the intended content. See the official Playwright dialogs guide and Page API.

// Handle a JavaScript dialog deliberately.
page.on('dialog', async dialog => {
  console.log('Dialog type:', dialog.type());
  await dialog.dismiss();
});

// Capture a page opened by a click in a new tab/window.
const popupPromise = page.waitForEvent('popup');
await page.getByRole('link', { name: 'Open report' }).click();
const popup = await popupPromise;
await popup.waitForLoadState('domcontentloaded');
await popup.screenshot({ path: 'report.png' });

Choose whether to accept or dismiss dialogs based on the task. For example, a beforeunload prompt can block navigation; a confirmation may represent a meaningful action and should not be accepted automatically.

6. Wait for the right page state

Navigation completing does not guarantee that the useful content, consent banner, or late-rendered popup is ready. Use a condition tied to the page you need: a visible heading, a known content selector, or a known overlay selector. For dynamic pages, a short explicit delay can help with delayed overlays, but arbitrary long sleeps make capture slow and still may miss unpredictable states.

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('main').waitFor({ state: 'visible', timeout: 15_000 });
// If the target is rendered after an app-specific transition, wait for its selector.
await page.locator('.report-chart').waitFor({ state: 'visible', timeout: 20_000 });

Network-idle waits can be unreliable on pages with analytics, polling, or persistent connections. Prefer a specific readiness signal when you know one. Refresh the accessibility snapshot after dismissing an element, navigating, or otherwise changing page state. Vercel’s agent-browser guidance likewise recommends taking a fresh snapshot after handling a covering consent banner or modal before retrying an obstructed action; see the agent-browser repository.

7. Or skip the browser setup

ScreenshotNeo’s API documentation covers its screenshot endpoint and options. One GET request can return a screenshot. Here is a cURL example; replace the target URL with the page you need and provide an API key.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, and cache hits are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

8. Troubleshooting

Symptom Likely cause Fix
The click is blocked by a banner or modal The overlay covers the target, or the locator came from an old snapshot. Inspect a fresh snapshot, use the overlay’s visible control if appropriate, refresh the snapshot, then retry. This is the workflow recommended in the agent-browser guide.
The popup is still in the screenshot after masking The selector matched the wrong element, the overlay appeared later, or the screenshot was taken from a different page. Verify the selector against the current page and wait for the relevant state. Confirm you are capturing the intended page or popup tab.
A colored block covers the popup A mask covers the element’s bounding box; masking is not removal. Use a capture stylesheet to hide the specific element if appropriate, or dismiss it through its control when page-state changes are allowed.
Wanted content disappeared too The selector was too broad, such as a generic dialog or fixed-position rule. Use a narrower selector observed in the DOM or snapshot and inspect the image again.
The screenshot is blank or the main content is missing The page had not reached the required render state, or content requires an app-specific readiness condition. Wait for a known content selector or visible heading. Check the URL and navigation outcome before capturing.
Lower content is absent The capture was viewport-only, or lazy content did not load. Use full-page capture. For lazy content, scroll the page or wait for relevant content to render before capturing.
Automation stalls on an alert A dialog listener was registered but did not accept or dismiss the dialog. Ensure every dialog event handler resolves the dialog according to the task, or remove the listener if default auto-dismissal is appropriate.
The script times out waiting for network idle Analytics, polling, or persistent network connections keep the page active. Wait for a specific page element or state instead of requiring global network silence.
The popup opened in another tab, but the screenshot shows the original page The automation kept using the original page object. Wait for the popup event and capture the returned popup page.

9. Reliability, performance, and cost

For repeatable results, keep the viewport, browser engine, target URL, wait condition, and popup-handling method consistent. Pages can vary by locale, login state, geolocation, personalization, viewport, or timing. If the goal is to compare site appearance, record these inputs alongside the image.

Viewport captures are generally quicker and smaller than full-page captures. Full-page images may consume more memory and can be very tall; use element captures when only one component matters. Waiting for a specific selector usually avoids wasting time on unrelated background requests. A screenshot API can avoid maintaining local browser installation and orchestration, while browser automation gives you direct control over page interactions and capture-time CSS.

Playwright itself is open-source browser automation; it does not require buying a hosted browser service for local use. The practical cost is the runtime and infrastructure you provide if you operate it at scale. ScreenshotNeo offers 1,000 shots per month free without a card, then plans from $5 for 3,000; see its site for the current plan details. Only use a consent interaction that matches the task, and treat a masked or hidden banner as a visual change in the image rather than a consent action.

10. Frequently asked questions

Can an AI agent remove a popup without clicking it?

For a visual-only screenshot, it can hide or mask the element during capture. That leaves the underlying page state unchanged and does not record consent. If the task requires dismissing the popup in the page itself, use its control.

Should I use a screenshot or an accessibility snapshot to find the close button?

Use the accessibility snapshot to identify controls and interact with them; use the screenshot to judge the visual result. This matches the guidance in the Playwright MCP screenshot documentation.

Can I capture a full page and a specific element at the same time?

No, the Playwright MCP screenshot tool documents that full-page capture cannot be combined with a target element. Make separate captures for those needs.

Why did the overlay return after I closed it?

The site may show it again after navigation, a reload, or a new browser context. Reinspect the current page each run rather than assuming the previous interaction persists.