How to Use an AI Agent to Screenshot a Webpage with a Browser Extension Enabled
Use Playwright with persistent Chromium to capture a page with an extension enabled, or use an agent’s supported shared browser session when you need existing sign-in state.
To capture a webpage with a browser extension enabled, run the AI agent’s browser in a context where the extension is actually installed and permitted. For a reproducible workflow, use Playwright’s bundled Chromium, launch it with a persistent context and the extension loaded, then capture the page. If you need the cookies or sign-in state from a browser tab you already have open, use an agent product’s supported shared-page feature instead; an isolated browser session will not inherit that state automatically.
This guide shows both approaches, how to choose the right capture, and what to check when the extension appears to be missing. “Can an AI agent use my browser extension while it captures a webpage?” The answer depends on the browser context and the agent’s supported integration: an extension being installed in your everyday browser does not mean it is available to every agent-controlled session.
1. Choose the browser context first
| Approach | Best for | What to verify |
|---|---|---|
| Dedicated Playwright Chromium context | Repeatable scripts, automation, and CI | The extension is loaded in that context and has access to the target site. |
| Agent-controlled isolated browser | Tasks that fit the browser integration provided by the agent | Whether that integration supports extensions. Do not assume an installed local extension is present. |
| User-shared existing page | Work that depends on an existing tab’s login, cookies, or storage | The agent product supports page sharing and the extension is active in the shared page. |
Playwright’s documented extension workflow is for Chromium with a persistent browser context. Its guide recommends Playwright-bundled Chromium for sideloading and warns that Chrome and Edge removed the command-line flags needed for this method. See the Playwright extension guide. The exact behavior of an extension still depends on its permissions, settings, and the target website.
For a shared page, consult the agent product’s browser documentation. For example, VS Code documents isolated browser sessions and sharing an existing page; this describes VS Code’s integration, not a universal capability across agents. Cloudflare likewise documents its own isolated Browser Run sessions.
2. Load an extension in Playwright and capture a page
Install Playwright, put an unpacked, compatible extension in a directory, and run the following Node.js script. Replace the target URL and extension path. This uses a persistent profile directory so the browser context can load the extension.
npm install playwright
const { chromium } = require('playwright');
const path = require('path');
(async () => {
const extensionPath = path.resolve('./my-extension');
const context = await chromium.launchPersistentContext('./browser-profile', {
channel: 'chromium',
args: [
`--disable-extensions-except=${extensionPath}`,
`--load-extension=${extensionPath}`,
],
});
try {
const page = context.pages()[0] ?? await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.screenshot({ path: 'webpage.png', fullPage: true });
console.log('Saved webpage.png');
} finally {
await context.close();
}
})();
Save it as capture.js and run node capture.js. The extension directory must contain a compatible extension build, including its manifest. This example is adapted from Playwright’s documented launch configuration; it does not establish that any specific extension will work on any target page.
Confirm the extension is active
Before relying on the screenshot, check the extension’s expected visible effect or behavior on the page. If you need to inspect a Manifest V3 extension’s service worker, Playwright’s guide shows how to wait for it. The guide notes that a Manifest V3 service worker can be suspended after roughly 30 seconds without activity and restarted when needed. Do not treat a successful browser launch as proof that the extension changed the page.
Use the user’s current tab when state matters
A newly launched profile is not the same as a user’s everyday browser profile. When a task needs an existing authenticated tab, use a sharing feature supported by the agent product and ask the agent to verify it can access the intended page. VS Code’s browser tools document navigation, interaction, content reading, and screenshot capture, as well as the distinction between isolated pages and pages shared by the user. See VS Code browser tools.
3. Pick viewport, full-page, element, or clipped capture
Choose the capture scope that matches the task. A viewport screenshot captures the current visible screen; a full-page screenshot includes the scrollable page; an element screenshot or clip focuses on a component or region. Playwright documents screenshot options, including clipping, masks, PNG/JPEG output, and CSS-pixel or device-pixel scaling. See the Playwright screenshot guide.
// Current viewport
await page.screenshot({ path: 'viewport.png' });
// Entire scrollable page
await page.screenshot({ path: 'full-page.png', fullPage: true });
// A specific element
await page.locator('main article').screenshot({ path: 'article.png' });
// A viewport region in CSS pixels
await page.screenshot({
path: 'region.png',
clip: { x: 40, y: 80, width: 900, height: 600 },
});
// Choose the output format explicitly
await page.screenshot({ path: 'page.jpg', type: 'jpeg', quality: 85 });
Use CSS-pixel scaling for a one-to-one mapping between CSS pixels and image pixels. Device scaling follows the device pixel ratio and can produce a larger image. For a visual review, verify the dimensions and scope from the returned file rather than assuming the browser used the intended settings.
For AI agents, screenshots are useful for checking visual layout and canvas or chart output. If the task is to read text or understand page structure, an accessibility snapshot may be a better tool than an image; Playwright MCP documents viewport, selected-element, and full-page captures and describes this distinction in its screenshot guidance.
4. Make the capture reliable
- State the context requirement. Tell the agent whether to use a fresh automation browser or a user-shared page.
- Verify extension availability. Confirm the extension is loaded and active in the controlled context before navigating or capturing.
- Navigate to the exact URL. Wait for the content or state relevant to the task. A generic network-idle condition is not a guarantee that every application is ready.
- Capture the right scope. Use the viewport for the current screen, full-page for scrollable content, or a locator or clip for a focused region.
- Report the artifact and errors. Have the agent provide the saved filename and any extension or session errors. Do not claim a screenshot exists until the tool returns or saves one.
Some pages render content after navigation through client-side code, lazy loading, or user interaction. If the screenshot is blank or incomplete, wait for a page-specific visible element or state before capturing, then inspect the result. There is no single readiness check that fits every site and extension.
5. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The agent cannot see the extension | It opened an isolated profile or a browser integration without extension support. | Load the extension in the Playwright Chromium context, or use a product-supported shared page and verify the extension there. |
| Launch arguments have no effect in Chrome or Edge | The documented sideload approach relies on flags those browsers removed. | Use Playwright’s bundled Chromium as described in its extension guide. |
| The extension is present but the page looks unchanged | Site access permission, extension settings, or site behavior may prevent the expected effect. | Check permissions and settings, then inspect the page for the extension’s effect before taking the final screenshot. |
| The image cuts off below the fold | The capture used the viewport scope. | Set fullPage: true or capture the specific element containing the needed content. |
| The image is unexpectedly large | Device-pixel scaling follows the device pixel ratio. | Choose CSS-pixel scaling when one image pixel per CSS pixel is desired. |
| The agent cannot use an authenticated page | Its isolated browser context does not inherit the user’s cookies, storage, or sign-in state. | Use the agent product’s supported page-sharing workflow, if available, and confirm access to the intended tab. |
| The capture is blank or misses late content | The page had not reached the state needed for capture. | Wait for a page-specific element or visual state, then capture and inspect the output. |
6. Performance, reliability, and cost
A local Playwright workflow gives you control over the browser context and capture timing, but you are responsible for maintaining the extension build, browser profile, permissions, and automation environment. A persistent context is useful for loading an extension; it does not make a separate profile inherit the user’s everyday browser state. Full-page or device-scale captures can create larger image files, so use the narrowest scope and scale that answer the task.
For repeatable jobs, keep the extension version and browser setup consistent, use a dedicated profile directory, and check the produced image as part of the workflow. The cited documentation does not publish a general speed or success benchmark for extension-enabled screenshots, so performance depends on the page, extension, and environment.
7. Or skip the browser setup
If you need a clean webpage screenshot and do not need a particular installed extension to alter the page, ScreenshotNeo provides a screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. The API parameter names used by other screenshot services also work, which can make switching easier. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of these steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000. Every feature is included on every plan. An API screenshot does not run your own browser extension, so use the Playwright method when that extension’s behavior is essential.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
8. FAQ
How do I take a screenshot with a Chrome extension enabled?
Use Playwright’s bundled Chromium with a persistent context and the extension-loading arguments shown above, then verify the extension is active in that browser context.
Why does my browser automation not see the extension?
The automation may have launched an isolated profile that does not contain it, or the agent’s browser integration may not support arbitrary extensions. Check the specific integration and use a supported extension-loading or page-sharing method.
Can the agent use my browser’s signed-in session?
Only if the agent product supports sharing or controlling that existing page. An isolated session does not automatically share your cookies, storage, or sign-in state.
Should I use a screenshot to extract page text?
Use a screenshot to inspect visual appearance. For reading and understanding page structure, use the browser’s text or accessibility tools when available.


