ScreenshotNeo

BlogAI agents

How to Make an AI Agent Capture Only the Main Content of a Webpage as a Screenshot

Use Playwright to inspect a webpage, identify its main content, and capture just that element. Includes runnable JavaScript, Python, cURL, and Node.js examples.

By the ScreenshotNeo team4 October 20269 min read

To make an AI agent capture only a webpage’s main content, have it inspect the page structure, select the article or content element, and take an element screenshot. In Playwright, the core JavaScript call is await page.locator('main').screenshot({ path: 'main-content.png' }). Treat main as a candidate, not a universal answer: inspect the selected element and use a narrower selector if it includes navigation, sidebars, or recommendations.

This approach is for capturing the bounds of a particular DOM element. If the goal is the entire long article, use a full-page capture strategy and check how the page handles scrollable containers and lazy-loaded content.

1. How the workflow works

  1. Open the target page. Wait for the content you need to appear. A fixed delay is not a substitute for checking that the target exists.
  2. Inspect structure and text. Use the DOM or an accessibility snapshot to understand the page’s landmarks and identify likely content containers. Accessibility snapshots help with structure and text; screenshots show visual layout. Playwright agent CLI documentation.
  3. Choose a selector. Try semantic elements such as main or article, then verify the selected element excludes page furniture.
  4. Capture the element. Playwright scrolls a locator into view and clips the screenshot to the element’s position and dimensions. Playwright locator screenshot documentation.
  5. Validate the image. Confirm it contains the intended content and the expected amount of it. Refine the selector or choose full-page capture if needed.

A locator screenshot captures the matching element’s bounds. It does not automatically expand an internally scrollable element to reveal its hidden contents. A detached element can also cause capture to fail. See the locator screenshot API details.

2. Runnable JavaScript example with Playwright

This Node.js script opens a page, waits for a candidate content element, captures it, and reports useful errors. Install Playwright and its browser first with npm install playwright and npx playwright install chromium.

const { chromium } = require('playwright');

(async () => {
  const url = process.argv[2] || 'https://example.com';
  const selector = process.argv[3] || 'main';
  const browser = await chromium.launch({ headless: true });

  try {
    const page = await browser.newPage({
      viewport: { width: 1440, height: 1000 },
      deviceScaleFactor: 1,
    });
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });

    const content = page.locator(selector).first();
    await content.waitFor({ state: 'visible', timeout: 15000 });

    // Inspect the candidate before relying on it.
    const details = await content.evaluate((el) => ({
      tag: el.tagName.toLowerCase(),
      id: el.id,
      className: typeof el.className === 'string' ? el.className : '',
      textLength: (el.innerText || '').length,
      width: Math.round(el.getBoundingClientRect().width),
      height: Math.round(el.getBoundingClientRect().height),
    }));
    console.log('Selected content:', details);

    await content.screenshot({ path: 'main-content.png', animations: 'disabled' });
    console.log('Saved main-content.png');
  } catch (error) {
    console.error(`Could not capture ${selector} on ${url}: ${error.message}`);
    process.exitCode = 1;
  } finally {
    await browser.close();
  }
})();

Run it as node capture-main.js https://example.com main. Replace main with a site-specific selector when the page has no useful main landmark or the landmark contains unrelated material.

Use a more precise selector

For a page with multiple landmarks, inspect matches and choose the intended one. A site-specific article container can be more reliable than a broad landmark:

const candidates = page.locator('main, article');
console.log('Candidate count:', await candidates.count());

// Use the selector you confirmed from the page, for example:
const article = page.locator('article.post-content').first();
await article.waitFor({ state: 'visible' });
await article.screenshot({ path: 'article.png' });

Do not assume that the first main or article is correct. Check its text, dimensions, and visual result. The best selector is the narrowest stable element that contains the content requested.

3. Python equivalent

Install the Python package and browser with pip install playwright and playwright install chromium.

import sys
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

url = sys.argv[1] if len(sys.argv) > 1 else 'https://example.com'
selector = sys.argv[2] if len(sys.argv) > 2 else 'main'

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    try:
        page = browser.new_page(viewport={"width": 1440, "height": 1000})
        page.goto(url, wait_until='domcontentloaded', timeout=30000)
        content = page.locator(selector).first
        content.wait_for(state='visible', timeout=15000)

        details = content.evaluate("""el => ({
          tag: el.tagName.toLowerCase(),
          id: el.id,
          className: typeof el.className === 'string' ? el.className : '',
          textLength: (el.innerText || '').length,
          width: Math.round(el.getBoundingClientRect().width),
          height: Math.round(el.getBoundingClientRect().height)
        })""")
        print('Selected content:', details)
        content.screenshot(path='main-content.png', animations='disabled')
        print('Saved main-content.png')
    except PlaywrightTimeoutError as exc:
        print(f'Could not find a visible element matching {selector!r}: {exc}', file=sys.stderr)
        raise
    finally:
        browser.close()

Run it with python capture_main.py https://example.com main. The API call for the element capture is page.locator("main").screenshot(path="main-content.png").

4. Element screenshot or full-page screenshot?

Need Use Watch for
Only the article/content region Screenshot a verified locator The locator may include unrelated material or omit content hidden inside a scrollable child.
The full document, including content below the fold Full-page screenshot Long pages, lazy content, and fixed-position elements can affect the result; inspect the output.
A particular visible portion of an element Element screenshot at the current layout Hidden portions of an internally scrollable element are not automatically revealed.

Playwright’s agent CLI treats full-page capture and element targeting as separate modes; its documentation says the full-page option cannot be combined with an element target. Playwright agent CLI documentation. If the desired output is the whole article, decide explicitly how to capture the page rather than expecting a locator screenshot to expand every scrollable region.

5. cURL and direct API calls

cURL cannot inspect a live DOM and select an element by itself. It can call a screenshot service that performs the browser capture, but whether that service supports element selection depends on its API. For local browser automation, use the Playwright code above. For a one-request screenshot, ScreenshotNeo returns an image or PDF from a URL; its request below captures the page according to API options rather than running the DOM-inspection workflow shown above.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp

See the ScreenshotNeo API documentation for parameters, including element capture by CSS selector. If you need the agent to inspect the DOM and decide on a selector at runtime, use a browser automation tool for that decision, then pass a known selector to a screenshot API if its options support it.

6. Selector choice and capture options

Finding the main region

  • Start with semantic landmarks such as main and article.
  • Check whether there are zero, one, or multiple matches.
  • Inspect text and bounds to catch empty, tiny, or overly broad selections.
  • Prefer a stable site-specific selector if semantic elements include navigation, related stories, or other page furniture.
  • Save or inspect a sample capture before relying on the selector across different page templates.

There is no universal selector that identifies main content reliably on every site. Treat candidate selectors as hypotheses and validate them on each site.

Styling and dynamic content

The locator screenshot API supports applying a stylesheet during capture, which can help make captures repeatable or hide dynamic elements. Use styling narrowly and only when omitting those elements is appropriate. There is no single evidence-backed rule that handles every consent dialog, sticky banner, lazy-loading pattern, or site-specific delay. Inspect the page and adapt to its behavior.

For lazy-loaded content, wait for the content you intend to show to become available and verify the capture. For an internally scrollable container, scrolling the page to the locator does not reveal all content inside that container. Handle that case deliberately, for example by scrolling the container before capturing if the required result includes its hidden portions.

7. Reliability, performance, and cost considerations

  • Reliability: Wait for a visible target and handle missing or detached locators. Validate the image when page templates or content change.
  • Stability: Animations and dynamic content may change what appears in the image. Playwright’s screenshot options include animation handling and a stylesheet option; apply them only as needed.
  • Performance: Capturing one element avoids including unrelated page regions in the output. Navigation and page rendering still take time; choose a readiness condition based on the content rather than adding a large fixed sleep.
  • Cost: A self-hosted Playwright flow has browser runtime and infrastructure costs that depend on where and how it runs. Screenshot API costs depend on the chosen service and plan. ScreenshotNeo offers 1,000 screenshots per month free without a card; paid plans start at $5 for 3,000.

8. Troubleshooting

Symptom Likely cause Fix
No element found or timeout waiting for it The page has no matching landmark, content has not loaded, or the selector is wrong. Inspect the DOM or accessibility snapshot, confirm the selector, and wait for the actual content element.
Screenshot contains navigation or sidebar The selected main or article is broader than the requested content. Inspect candidate elements and use a narrower, site-specific selector.
Screenshot contains only part of the content The target or one of its children is internally scrollable, or the article extends beyond the selected bounds. Check scrollable containers and decide whether the goal is the visible element or the full article. Use full-page capture for the full document where appropriate.
Target disappears during capture The page re-rendered and detached the element. Re-query the locator after the content settles and capture promptly; handle retries only when the page behavior makes them safe.
Element is obscured or image looks wrong An overlay, sticky element, animation, or layout change affected the capture. Inspect the visible page state, wait for the intended content, and use narrowly scoped screenshot styling if appropriate.
Blank or incomplete capture Navigation or rendering did not finish as expected, or the site requires additional interaction. Check the browser page state and target visibility before capture. Do not assume one wait condition works for every site.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API supports element capture by CSS selector, alongside full-page screenshots and other capture options. A single request can return a screenshot without installing and managing a browser locally. See the API docs for the selector parameter and other options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', image);
  • Cookie and consent banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses include page-verdict and billing headers.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo and get 1,000 free screenshots a month, with no card required.

10. FAQ

Does an element screenshot include the browser’s navigation bar?

No. Playwright captures the selected page element within the page, not the browser chrome.

Can an AI agent choose the selector automatically?

It can inspect the DOM or accessibility snapshot and propose a selector, but it should verify the match and resulting image. Page semantics vary.

Should I use main or article?

Use whichever verified element best matches the requested content. Neither is guaranteed to be present or precise on every site.

Can an element screenshot capture a whole long article?

It captures the selected element’s bounds, but internally scrollable content may remain hidden. For the full document, choose a full-page approach and check the output.