ScreenshotNeo

BlogAI agents

How to Use an AI Agent to Capture a Webpage Screenshot and Extract Its Page Title

Use Playwright to capture a webpage and read its title directly, with runnable agent code, capture options, error handling, and a no-browser-setup alternative.

By the ScreenshotNeo team4 October 202610 min read

Use a browser automation runtime such as Playwright: navigate to the URL, save a screenshot, then retrieve the page title with page.title(). The screenshot is visual evidence; the title is text read directly from the page, so you do not need vision-based OCR to extract it.

The example below captures the full scrollable page and prints the title, final URL, and screenshot path. For a viewport-only image, omit fullPage: true. [Playwright Page API]

1. Set up Playwright

Install Playwright and its Chromium browser in a Node.js project:

npm install playwright
npx playwright install chromium

Save the following as capture-title.js. Pass the target URL as the first command-line argument.

const { chromium } = require('playwright');

async function main() {
  const inputUrl = process.argv[2];
  if (!inputUrl) {
    throw new Error('Usage: node capture-title.js https://example.com');
  }

  let parsedUrl;
  try {
    parsedUrl = new URL(inputUrl);
  } catch {
    throw new Error('The URL is invalid. Include a complete URL, such as https://example.com');
  }
  if (!['http:', 'https:'].includes(parsedUrl.protocol)) {
    throw new Error('Only HTTP and HTTPS URLs are supported by this example.');
  }

  const browser = await chromium.launch();
  try {
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    const response = await page.goto(parsedUrl.toString(), {
      waitUntil: 'domcontentloaded',
      timeout: 30000
    });

    // This example records the title after initial document navigation.
    // For an SPA that updates its title later, wait for the expected title/state.
    const title = await page.title();
    const screenshotPath = 'page.png';
    await page.screenshot({ path: screenshotPath, fullPage: true });

    console.log(JSON.stringify({
      requestedUrl: parsedUrl.toString(),
      finalUrl: page.url(),
      httpStatus: response ? response.status() : null,
      title,
      screenshotPath,
      captureScope: 'full page'
    }, null, 2));
  } finally {
    await browser.close();
  }
}

main().catch((error) => {
  console.error(JSON.stringify({
    success: false,
    error: error.message
  }, null, 2));
  process.exitCode = 1;
});

Run it with:

node capture-title.js https://example.com

The script checks that the input is a valid HTTP or HTTPS URL, reports the final URL after redirects, and includes the HTTP status when navigation produced a response. A non-2xx status does not necessarily mean navigation failed: inspect the status and page content before treating the result as successful. The finally block closes Chromium even if navigation, title retrieval, or screenshot writing throws an error.

2. Choose the right capture scope

Pick the screenshot scope to match the agent’s task. Playwright can capture the current viewport, the full scrollable page, or a particular locator. [Page API, Playwright MCP screenshot guidance]

Task Playwright code Considerations
Inspect what a visitor sees without scrolling await page.screenshot({ path: 'viewport.png' }) Uses the configured viewport dimensions.
Capture the complete scrollable document await page.screenshot({ path: 'full.png', fullPage: true }) Can produce a very tall image on long pages.
Capture one component or region await page.locator('main').screenshot({ path: 'main.png' }) Wait for the locator to exist; a selector matching multiple elements may need to be narrowed.

For a region, use a stable selector that identifies the intended element. If the selector is dynamic or ambiguous, wait for the right content and refine it rather than silently capturing the wrong match. In Playwright MCP’s screenshot tool, an element target and full-page mode are separate capture modes. [Playwright MCP screenshots]

3. Extract the title and wait for the page you need

await page.title() returns the browser page’s title text. Read it after the page has reached the state that matters to your task. domcontentloaded is a useful starting point, but it does not guarantee that every client-rendered element or late title update has finished.

If a known title is expected, wait for it explicitly instead of sleeping for an arbitrary duration:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.waitForFunction(
  (expected) => document.title === expected,
  'Product dashboard',
  { timeout: 10000 }
);
const title = await page.title();

If the expected title is not known in advance, wait for a task-specific signal: a known element, a route transition, or another condition the application defines. Then read page.title(). There is no single wait condition that fits every website.

For an SPA, navigation may complete before the client updates the document title. Wait for the relevant application transition, then read the title. If the title is empty or stale, report that uncertainty with the final URL instead of presenting it as verified metadata.

4. Return useful results from the agent

Keep the title extraction separate from visual interpretation. Return the title as text and attach or identify the screenshot as a visual artifact. Include enough provenance for a caller to understand what was captured:

  • The requested URL and final URL after redirects.
  • The returned title, including an empty string if the document has no title.
  • The screenshot path or artifact reference.
  • The capture scope: viewport, full page, or selected element.
  • The HTTP status when available and any navigation or file-writing error.

Use screenshots for layout, rendering, colors, and canvas or chart appearance. Use an accessibility snapshot when the task needs page structure, text, or interaction targets; screenshots are visual evidence and are not a substitute for a structured page representation. [Playwright MCP screenshot guidance]

If you use Playwright’s agent CLI snapshot flow, take a fresh snapshot after navigation. References belong to a particular snapshot and become invalid when the page changes. [Playwright agent CLI snapshots]

5. Run the same workflow from Python

Playwright’s Python package exposes the same browser workflow. Install the package and Chromium:

pip install playwright
playwright install chromium

Save this as capture_title.py and pass the URL on the command line:

import asyncio
import json
import sys
from urllib.parse import urlparse
from playwright.async_api import async_playwright

async def main(url):
    parsed = urlparse(url)
    if parsed.scheme not in ('http', 'https') or not parsed.netloc:
        raise ValueError('Provide a complete HTTP or HTTPS URL, such as https://example.com')

    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            page = await browser.new_page(viewport={"width": 1440, "height": 900})
            response = await page.goto(url, wait_until='domcontentloaded', timeout=30000)
            title = await page.title()
            screenshot_path = 'page.png'
            await page.screenshot(path=screenshot_path, full_page=True)
            print(json.dumps({
                'requested_url': url,
                'final_url': page.url,
                'http_status': response.status if response else None,
                'title': title,
                'screenshot_path': screenshot_path,
                'capture_scope': 'full page'
            }, indent=2))
        finally:
            await browser.close()

if __name__ == '__main__':
    if len(sys.argv) != 2:
        raise SystemExit('Usage: python capture_title.py https://example.com')
    try:
        asyncio.run(main(sys.argv[1]))
    except Exception as error:
        print(json.dumps({'success': False, 'error': str(error)}, indent=2), file=sys.stderr)
        raise SystemExit(1)

For synchronous Python, Playwright also provides a synchronous API; the navigation, screenshot, and title calls have corresponding sync forms. Keep the browser lifecycle in a try/finally or context manager so failures do not leave browser processes running.

6. Use an agent tool or hosted browser runtime

An AI agent does not need raw browser access for every operation. A useful tool boundary is a function that accepts a validated URL and returns structured fields such as title, final_url, screenshot_path, and error. Keep screenshot capture and title extraction in the same browser session so they describe the same rendered page.

A hosted browser runtime is another implementation shape: the browser runs remotely and the application controls it through the provider’s interface. Cloudflare documents CDP-controlled browser sessions for screenshots and rendered page state, including information available after JavaScript runs; its browser tools page labels those tools Beta and gives a page update date of June 24, 2026. Check the provider’s current availability and terms before choosing it. This hosted option is not required for the Playwright workflow above. [Cloudflare Agents Browser tools]

For either model, treat user-provided destinations as untrusted input and apply your application’s destination controls before navigation. Browser automation can access network resources, so the allowed destinations and credentials should be deliberate. The Playwright APIs describe browser operations, not a complete destination security policy.

7. Performance, reliability, and cost

  • Reuse browsers for batches: launching a browser for every URL adds setup work. A service that processes multiple pages can reuse a browser and create a fresh page or context for each task, while isolating state when needed.
  • Set timeouts: navigation can stall or be blocked. Bound navigation and any task-specific waits, then return a clear failure rather than waiting indefinitely.
  • Choose the smallest useful image: viewport screenshots are generally smaller than full-page captures. Full-page images can be very tall, take longer to encode, and consume more storage or transfer.
  • Keep outputs reproducible: use explicit file paths and report the requested and final URLs plus capture scope.
  • Report partial outcomes accurately: a title may be available even if screenshot writing fails, or a screenshot may exist while the title is empty. Return each result and each error separately when your agent interface allows it.
  • Account for operating costs: a local Playwright setup requires a machine or service to run the browser and store or deliver image files. Hosted browser costs and limits depend on the provider; the sources cited here do not establish a comparable price or performance benchmark.

Do not treat a successful browser navigation as proof that the page is the intended content. Check the final URL, status, title, and relevant page state, particularly when redirects, access challenges, or application errors are possible.

8. Troubleshooting

Symptom Likely cause What to do
page.goto() times out The server is slow, the requested load state never occurs, or the page is unreachable. Set a bounded timeout, choose a suitable navigation milestone such as domcontentloaded, and report the failure. Do not convert a timeout into a successful capture.
Title is blank or stale The document lacks a title or a client-rendered page updates it after initial navigation. Check the final URL and wait for a page-specific state or expected title before reading page.title().
Screenshot is blank or shows a challenge The site may have blocked automation, returned an error page, or not rendered the expected content. Inspect the final URL, status, title, and visible page. Return the observed result as a failure or blocked page rather than claiming the target was captured.
Element screenshot fails The locator does not match, is ambiguous, or the element has not appeared. Wait for the intended locator, refine the selector, and capture that element. Confirm the selector identifies the intended region.
Full-page capture is unexpectedly large The document is long or contains extensive generated content. Use viewport or element capture if that meets the task, or retain full-page mode and plan for larger files and slower transfers.
Chromium cannot launch The browser binary is not installed or the runtime lacks required system dependencies. Run the Playwright browser installation command for Chromium and follow the installation guidance for the deployment environment.
CLI element references no longer work The page changed after the snapshot that created those references. Take a new snapshot after navigation and use references from the current snapshot. [Playwright snapshots]
Screenshot file is missing The output directory may not exist or the process lacks write access. Use a known writable path, create its parent directory, and treat screenshot writing as a separate operation that can fail.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. The call below captures a page; see the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status} ${await res.text()}`);
const image = Buffer.from(await res.arrayBuffer());
await require('node:fs/promises').writeFile('shot.webp', image);

ScreenshotNeo’s MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. For this title-extraction workflow, use its page-info tool for page information and its screenshot tool for the visual artifact.

  • Cookie banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing status in headers.
  • The MCP server lets AI agents take screenshots and retrieve page information.
  • The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

10. FAQ

Can an AI agent get a page title from a screenshot alone?

It can attempt visual OCR, but browser automation can retrieve the title directly as text with page.title(), which is the more direct method for this task.

Should I return the screenshot or the title to the agent?

Return both when the task needs visual inspection and metadata. Keep the title as a text field and the screenshot as an image artifact so the caller can use each for its intended purpose.

Does full-page mode capture content that loads only after scrolling?

Full-page mode captures the scrollable page, but pages can defer loading images or other content until scrolling. If the task depends on lazy-loaded content, make the page load that content before capture and verify the result.

Can I use this approach for authenticated pages?

Yes, when your application is authorized to access them. Provide session state through a controlled browser context and avoid returning credentials or sensitive page content in logs or agent output.