ScreenshotNeo

BlogAI agents

How to Capture a Webpage Screenshot with an AI Agent in Multiple Browser Tabs

Use one Playwright page handle per browser tab, prepare each page, and capture separate viewport, full-page, or element screenshots with clear filenames.

By the ScreenshotNeo team4 October 20269 min read

To capture screenshots from multiple browser tabs with an AI agent, keep a separate page handle for each tab, bring each page to the state you need, then take a screenshot from each page and save it under a distinct filename. In Playwright, a single BrowserContext can contain multiple Page objects. Each page can capture its visible viewport, the full scrollable page, or a selected element.

Use browser snapshots or accessibility structure to find and operate controls; use screenshots to inspect visual appearance. Playwright’s documentation describes pages, contexts, and screenshots in its Pages guide, Page API, and screenshot guidance.

1. Choose how the agent gets browser access

First decide which browser state the job needs:

  • Dedicated automation context: Start a browser and create a new context for the job. This keeps the task’s pages and state separate from your personal browser session. Add authentication only when the workflow requires it.
  • Existing personal browser profile: Use this when the page is already authenticated or has necessary state. Connecting an agent to a personal profile can expose open tabs, cookies, and storage state. Chrome’s documentation describes this access for its auto-connect implementation; treat profile access as a meaningful security boundary and connect only when needed.
  • Managed browser runtime: A hosted browser can be useful when the agent needs remote execution. Cloudflare documents browser tools for agent interaction over CDP; verify current availability and terms with the provider before choosing a service.

For scripted multi-tab capture, Playwright’s Page API gives direct control of pages. For an agent-driven session, Playwright MCP offers screenshot tools and browser snapshots. Its guidance distinguishes visual evidence from interaction targets: screenshots are for inspection, while snapshots provide references for interacting with page elements.

2. Install Playwright and prepare a multi-page capture

For a new Node.js project, install Playwright and its Chromium browser:

npm install playwright
npx playwright install chromium

Save the following as capture-tabs.js, then run node capture-tabs.js. It opens two URLs in separate pages in one context and writes a separate full-page PNG for each. The example includes navigation error handling and a selector-based readiness check; replace the selectors with something meaningful for your target pages.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const context = await browser.newContext({
    viewport: { width: 1440, height: 900 },
  });

  const tasks = [
    { label: 'example', url: 'https://example.com/', readySelector: 'h1' },
    { label: 'playwright', url: 'https://playwright.dev/', readySelector: 'h1' },
  ];

  try {
    const results = await Promise.all(tasks.map(async (task) => {
      const page = await context.newPage();
      try {
        const response = await page.goto(task.url, {
          waitUntil: 'domcontentloaded',
          timeout: 45_000,
        });
        if (response && !response.ok()) {
          throw new Error(`HTTP ${response.status()} for ${task.url}`);
        }

        await page.locator(task.readySelector).waitFor({
          state: 'visible',
          timeout: 15_000,
        });

        const path = `screenshots/${task.label}.png`;
        await page.screenshot({ path, fullPage: true });
        return { label: task.label, url: task.url, path, error: null };
      } catch (error) {
        return {
          label: task.label,
          url: task.url,
          path: null,
          error: error.message,
        };
      } finally {
        await page.close();
      }
    }));

    for (const result of results) {
      console.log(result);
    }
  } finally {
    await context.close();
    await browser.close();
  }
})();

Create the output directory before running the example: mkdir -p screenshots. Each task creates and owns its own page. The labels make output names stable and prevent captures from overwriting one another. This example waits for an element to become visible, which is only a useful readiness condition if that element indicates the page is ready for your task.

3. Make the agent operate the right tab

The core rule is to pass the intended page handle to every navigation, interaction, and screenshot call. Keep task metadata with the handle so the agent can report which URL and tab produced each image.

const pageTasks = await Promise.all(tasks.map(async (task) => {
  const page = await context.newPage();
  return { ...task, page };
}));

for (const task of pageTasks) {
  await task.page.goto(task.url);
  // Perform task-specific actions through task.page.
  await task.page.screenshot({ path: `screenshots/${task.label}.png` });
}

Pages do not need to be brought to the foreground to use the Page API. If you use an agent CLI or MCP client instead of application code, list tabs, select the intended tab, inspect a browser snapshot, interact with the selected page, and request its screenshot. Keep the tab identifier or task label attached to the result; never assume that the currently focused tab is the one you meant to capture.

When pages require different viewport sizes, locale, timezone, or other emulation, create separate contexts for those settings. Pages in one context share its emulation and browser state. Avoid sharing one context when tasks need isolation from one another, such as separate logged-in accounts.

4. Choose viewport, full-page, or element screenshots

Capture type Use it for Playwright example
Viewport The part currently visible in the tab await page.screenshot({ path: 'view.png' })
Full page The whole scrollable document, including content below the fold await page.screenshot({ path: 'full.png', fullPage: true })
Element A focused component such as a login form or chart await page.locator('form').screenshot({ path: 'form.png' })

For element capture, wait for the locator to be visible and use a selector that identifies the intended element uniquely. Playwright screenshot options also support masking elements that should be obscured in the output. For example:

await page.screenshot({
  path: 'masked.png',
  fullPage: true,
  mask: [page.locator('[data-sensitive]')],
});

Playwright MCP documents PNG, JPEG, and WebP output and CSS-pixel or device-pixel scale options. Choose a format and scale appropriate to how the image will be reviewed or stored. Higher pixel density and full-page output can create larger files.

5. Handle readiness, dynamic content, and state

Navigation finishing does not guarantee that a page’s meaningful content is ready. Define readiness per site and task: wait for a content selector, a known loading indicator to disappear, or a task-specific application state. Avoid treating one fixed delay as proof that all pages are ready.

For pages with asynchronous content, a bounded selector wait is often more reliable than sleeping for the same duration on every page:

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('[data-report-ready="true"]').waitFor({
  state: 'visible',
  timeout: 20_000,
});
await page.screenshot({ path: 'report.png', fullPage: true });

If a page needs interaction first, use its snapshot or accessibility structure to locate controls, interact through that page’s handle, and then capture. A screenshot is useful for visual checking, but it does not prove that the intended state or tab was captured. Associate each output with its page label and URL, and inspect it when correctness matters.

6. Run captures concurrently or sequentially

Parallel navigation and capture can reduce elapsed time for independent pages, but it uses more browser resources at once and some sites may throttle concurrent requests. The example in section 2 runs tasks concurrently with Promise.all. To limit concurrency, process pages in batches or use a small worker pool. To run sequentially, use a regular loop and await each task before starting the next.

Keep the work bounded: set navigation and readiness timeouts, close each page after its capture, and close the context and browser in a finally block. If a page fails, record its label and error and continue with other pages when the captures are independent. If all pages must succeed as one unit, collect errors and fail the job after all tasks finish.

7. Troubleshoot common failures

Symptom Likely cause Fix
Every screenshot shows the same tab The agent or script is using a selected/focused tab instead of each page handle. Keep one Page per task and call navigation and screenshot methods on that exact page. For CLI flows, select the target tab before acting and verify its URL.
Screenshot is blank or missing content The page had not reached task-specific readiness, content is lazy-loaded, or navigation failed. Check the navigation response, wait for a meaningful selector or state, and verify the resulting page URL and screenshot.
Timeout waiting for a selector The selector is wrong, content is delayed, a consent dialog obscures the page, or the page failed to load. Inspect a browser snapshot or the page structure, confirm the selector, handle necessary dialogs, and use a bounded timeout appropriate to the page.
Images below the fold are absent Images load only as the page scrolls, or the capture is viewport-only. Use fullPage: true for a full document shot and, where needed, scroll through content before capture to trigger lazy loading.
Output files overwrite one another Every page is saved to the same path. Include a stable task label or index in each filename.
Authenticated page redirects to sign-in The new context does not have the required authentication state. Provide the required state to a dedicated context using your approved setup, or use an existing profile only when the broader access is intentional.
Browser launch fails The Playwright browser binary is not installed or the runtime lacks required browser dependencies. Run the Playwright browser install step for the browser you launch, and follow the official installation guidance for the target operating system.
Full-page capture is extremely tall or fails The document is very long, continuously expanding, or resource-heavy. Capture a viewport or a specific element, or divide the page into meaningful sections. Avoid full-page capture when the task only needs one region.

8. Performance, reliability, privacy, and cost

  • Performance: Browser startup, page loading, rendering, and image encoding all contribute to run time. Reuse a browser process for a batch when appropriate, but keep contexts separated when state isolation matters. Limit concurrent pages if memory, CPU, or site rate limits become a problem.
  • Reliability: Set explicit timeouts, use task-specific readiness conditions, retain the URL and label with every artifact, and capture page-level failures. A successful screenshot call only means an image was produced; inspect it to verify that it contains the intended content.
  • Privacy: A connected personal browser can expose cookies, open tabs, and stored session data to the agent. Use a dedicated context when possible and grant profile access only for workflows that require it. Chrome’s claims about its local DevTools-for-agents server apply to that implementation and should not be generalized to other tools.
  • Cost: Self-hosted Playwright has no per-screenshot API charge, but uses compute, browser maintenance, and engineering time. Hosted browser runtimes or screenshot APIs may charge according to their own current terms; check provider pricing directly.

9. Or skip the browser setup

If you want screenshots from URLs without managing Playwright, contexts, browser binaries, and page handles, ScreenshotNeo is a website screenshot API and MCP server for developers. A GET request captures a URL as PNG, JPEG, WebP, or PDF. For a multi-tab task, send one request per URL and store each response under a distinct name. The API parameter names used by other screenshot APIs also work, which can make switching easier. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

For multiple URLs, repeat the request with a different URL and output filename for each page. ScreenshotNeo accepts cookie banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.

10. Frequently asked questions

Can Playwright capture tabs that are not active?

Yes. A context can contain multiple pages, and you can call the screenshot API on each page handle without bringing it to the foreground.

Should an AI agent use screenshots to find buttons?

Use browser snapshots or accessibility structure to find interaction targets, then use screenshots to inspect the page’s visual appearance.

How do I capture only a login form?

Locate the form and call locator.screenshot() on it after it is visible.

Can I use a logged-in personal browser?

Some agent connections can access profile data such as cookies and storage. Use that access only when the task needs it, and account for the broader data exposure.

Does this workflow require physical screenshot hardware?

No. Browser automation captures the rendered page through browser APIs; it does not require a camera or dedicated capture device.