How to Screenshot a Website with an AI Agent When the Page Has Infinite Scroll
Load the feed content you need before capturing it. This guide shows an AI agent workflow with Playwright, stopping rules, virtualized lists, and a one-call API option.
Short answer: An AI agent must scroll an infinite-scroll page to load the content you want before taking a full-page screenshot. In Playwright, scroll in controlled increments, wait for the site to reveal new items, stop when your target or stopping rule is reached, then capture with fullPage: true. Full-page capture includes the document’s current scrollable extent; it does not fetch content the page has not loaded.
This distinction matters: loading the feed and capturing its current state are separate tasks. An infinite feed has no natural endpoint, so decide how much content to include. For a virtualized feed that removes older items from the document, capture overlapping viewport segments as you scroll instead of relying on one tall image.
1. Choose a stopping rule
Before automating the page, define what “enough” means. A screenshot of an unbounded feed cannot contain every possible item. Choose a target that the agent can observe and report:
- Known item count: stop when at least N feed items are present.
- Target item: stop when a specific article, post, or marker appears.
- Stable end: stop after several scroll attempts produce no new items.
- Resource limit: stop at a maximum scroll count, elapsed time, or item count, and state that limit in the result.
Prefer a page-specific observable signal, such as item count increasing or a loading indicator disappearing. A fixed delay can be a fallback, but it cannot guarantee that content has loaded on a slow or variable page.
2. Inspect the page and identify its scroll container
Open the page and let its initial content render. Determine whether the document itself scrolls or whether the feed sits inside a nested scrolling element. An accessibility or DOM snapshot can help an agent locate the feed, its items, and loading indicators. Playwright MCP distinguishes screenshots for visual verification from accessibility snapshots for reading page structure and text; see the Playwright MCP screenshots documentation.
For a document-scrolling page, use window.scrollTo or mouse-wheel input. For a nested container, scroll that element’s scrollTop or move the pointer over it and use the mouse wheel. Scrolling the window will not necessarily trigger a nested feed.
3. Load content, then take a full-page screenshot with Playwright
Install Playwright and its Chromium browser:
npm init -y
npm install playwright
npx playwright install chromium
Save this as capture-infinite-scroll.js. Set TARGET_URL to the page you are authorized to access. The sample scrolls the document, checks whether the number of matching feed items increases, and stops after three unchanged rounds or 30 scrolls. Replace ITEM_SELECTOR with a selector that matches the site.
const { chromium } = require('playwright');
const TARGET_URL = 'https://example.com/feed';
const ITEM_SELECTOR = 'article';
const MAX_ROUNDS = 30;
const UNCHANGED_LIMIT = 3;
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });
try {
await page.goto(TARGET_URL, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator(ITEM_SELECTOR).first().waitFor({ state: 'visible', timeout: 15_000 });
let previousCount = await page.locator(ITEM_SELECTOR).count();
let unchangedRounds = 0;
for (let round = 0; round < MAX_ROUNDS && unchangedRounds < UNCHANGED_LIMIT; round++) {
await page.evaluate(() => window.scrollTo(0, document.documentElement.scrollHeight));
// Prefer a site-specific signal here, such as waiting for a loading
// indicator to disappear or for the item count to increase.
await page.waitForTimeout(700); // Fallback only; tune for the target site.
const currentCount = await page.locator(ITEM_SELECTOR).count();
if (currentCount > previousCount) {
unchangedRounds = 0;
} else {
unchangedRounds++;
}
previousCount = currentCount;
}
await page.screenshot({ path: 'screenshot.png', fullPage: true });
console.log(`Saved screenshot.png with ${previousCount} matching items in the DOM.`);
} finally {
await browser.close();
}
})().catch((error) => {
console.error(error);
process.exitCode = 1;
});
Run it with node capture-infinite-scroll.js. The item count is only a useful signal if the selector matches feed items and the site keeps loaded items in the DOM. Some sites append content after a longer delay, use a “Load more” button, or virtualize items. Adapt the loading condition to the site rather than treating the sample delay as universal.
Playwright defines fullPage: true as capturing the full scrollable page rather than only the viewport. Its scrolling guide describes scrolling content into view as a way to prompt an infinite list to load more. Read the primary references: Playwright Screenshots and Playwright scrolling actions.
4. Handle nested scroll containers, load buttons, and virtualized feeds
Nested scroll container
Find the element that actually scrolls, then scroll it directly. Replace the selector and item locator with site-specific values:
const feed = page.locator('[data-testid="feed"]');
await feed.evaluate((element) => {
element.scrollTop = element.scrollHeight;
});
await page.waitForTimeout(700);
Repeat the operation with the same item-count or loading-state checks as for document scrolling. If the page responds to wheel events rather than direct scrollTop changes, use page.mouse.move over the container and page.mouse.wheel(0, 800).
“Load more” button
Some feeds are not automatic infinite scroll. If a visible button loads the next batch, click it and wait for the item count or loading indicator to change. Repeat until the target or limit is met. Check that the button is enabled before clicking and stop if it disappears or becomes disabled.
Virtualized or recycled items
A virtualized list may keep only the visible items plus a small buffer in the DOM, replacing earlier nodes as you move down. In that case, the final full-page screenshot may contain only the currently rendered window or an incomplete representation. Capture viewport-sized segments while scrolling and overlap adjacent captures by part of a viewport. Keep the scroll position and segment order, then inspect the stitched result for gaps, duplicates, or repeated sticky headers. This is a practical workaround; page behavior depends on the site’s implementation.
5. Verify the screenshot
- Confirm the first expected item and the last item within your stopping rule appear.
- Check that the item count or target marker agrees with what the agent reports.
- Inspect lazy-loaded images and other assets; scrolling can trigger loading, but the capture may still happen before an asset finishes rendering.
- Look for blank areas, duplicated items, missing segments, or sticky elements repeated throughout a segmented capture.
- Record the stopping rule and extent, such as “captured 80 items” or “stopped after three scrolls with no new items.”
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot ends near the top | The page was captured before scrolling triggered more loads, or the full-page option was omitted. | Scroll first, wait for an observable load signal, then use fullPage: true. |
| Scroll loop stops too early | The item selector matches too few elements, or the site delays insertion beyond the fallback wait. | Inspect the DOM, correct the selector, and wait for the site’s actual loading signal or a count change. |
| Window scroll does nothing | The feed scrolls inside a nested container. | Identify the scrolling element and change its scrollTop or wheel over it. |
| Only the latest feed items appear | The site virtualizes or recycles older DOM nodes. | Capture overlapping viewport segments during scrolling and inspect the assembled image for continuity. |
| Images are blank or incomplete | Images are lazy-loaded or still decoding when the screenshot starts. | Scroll them into view and wait for relevant images to finish loading; use a site-specific readiness check. |
| Navigation times out | The page continues making requests, or the chosen navigation wait condition is too strict. | Use domcontentloaded when appropriate, then wait for the page’s meaningful content signal. Set a reasonable timeout and handle failures. |
| Screenshot is extremely tall or memory-heavy | The chosen extent includes too many feed items. | Set a maximum count or scroll limit, or use segmented captures. Report the extent rather than attempting an unbounded capture. |
7. Performance, reliability, and cost
Every additional scroll, wait, and rendered item adds time and browser work. Full-page screenshots of very long documents can use substantial memory and produce images that are difficult to inspect. Set explicit limits, use a specific target where possible, and avoid repeated captures if one verified capture will do.
For reliable automation, use observable conditions instead of a universal fixed sleep, keep navigation and operation timeouts bounded, and make failures visible to the calling agent. A retry can help with transient navigation or load failures, but cap retries and avoid endlessly repeating a page that never reaches the required condition. If completeness matters, verify item boundaries and the output image rather than assuming a successful screenshot call proves the feed was fully loaded.
With a self-hosted Playwright workflow, account for the compute and browser runtime you operate; actual cost depends on your infrastructure and capture volume. The page may also require authentication or site-specific state. Use only access you are permitted to use, and pass credentials through secure browser context configuration rather than embedding secrets in source code.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its one-call API is useful when the page’s current rendered extent is sufficient; a screenshot API call does not substitute for an agent-controlled scroll-and-load loop on an infinite feed. See the ScreenshotNeo API documentation for options and setup.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, and failed loads are never billed.
- An MCP server lets AI agents take screenshots with tools including
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Does fullPage: true load an infinite feed to its end?
No. It captures the current scrollable document extent. Scroll and wait for the content you want before capturing.
How do I know when to stop scrolling?
Use a target item or count when available. Otherwise, choose a bounded rule such as several rounds without new items, plus a maximum scroll count or time limit.
Can one screenshot preserve every item in a virtualized list?
Not necessarily. If earlier items are removed from the document as you scroll, capture overlapping viewport segments and check them for missing or repeated content.
Should an AI agent use a screenshot or an accessibility snapshot?
Use a screenshot to inspect visual appearance. Use an accessibility or DOM snapshot to identify structure, read text, and locate the feed or its loading signals.


