ScreenshotNeo

BlogHow-to

How to Capture Website Thumbnails for Pages With Infinite Scrolling

Capture a useful thumbnail of an infinite-scroll page by loading the content first, setting a clear boundary, then taking a full-page screenshot.

By the ScreenshotNeo team4 October 20268 min read

To capture a thumbnail that includes content loaded by infinite scrolling, first scroll through the part of the page you need and wait for new content to appear, then take a full-page screenshot. A full-page capture alone may only include content already present: it does not necessarily scroll the visible page or trigger scroll-dependent loading. Infinite feeds may have no natural end, so choose a practical boundary such as a target item count or the last section you need.

Why a full-page screenshot can miss content

Playwright describes a full-page screenshot as an image of the full scrollable page, as if it could fit on a very tall screen. But that does not mean the browser scrolls through the page in the way a visitor does. A Playwright issue describes full-page rendering that can happen beyond the visual viewport without an observable scroll event. If a site waits for scrolling or an IntersectionObserver callback before requesting more content, that content may not exist when the screenshot is taken.

So there are two separate things to consider:

  • Document height: the height of content currently laid out in the page.
  • Not-yet-loaded content: items the site will request or insert only after scrolling.

A screenshot can cover the first and still miss the second. The Playwright issue discussing this behavior proposes an opt-in lazy-content capture feature; it is a proposal, not a released Playwright option. Use the documented fullPage screenshot option, and trigger loading yourself before capture.

Playwright: scroll, wait, then capture

Install Playwright in a Node.js project and install its browser binaries:

npm install playwright
npx playwright install chromium

Save this as capture-infinite-scroll.js. It scrolls in viewport-sized steps, checks whether the document height has stopped growing for several consecutive steps, and then captures the loaded document. The stopping rule is an example, not a guarantee: adapt it to the site and your desired content boundary.

const { chromium } = require('playwright');

(async () => {
  const targetUrl = process.argv[2];
  if (!targetUrl) {
    throw new Error('Usage: node capture-infinite-scroll.js <url>');
  }

  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });

  try {
    await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 60000 });

    // Scroll through the page to trigger scroll listeners and lazy loading.
    // Stop after the document height is unchanged for several steps, or at
    // maxSteps. Tune the pause and stopping rule for the target site.
    const maxSteps = 80;
    const stableStepsRequired = 4;
    const pauseMs = 700;
    let previousHeight = 0;
    let stableSteps = 0;

    for (let step = 0; step < maxSteps; step++) {
      const height = await page.evaluate(() => document.documentElement.scrollHeight);
      if (height === previousHeight) {
        stableSteps++;
      } else {
        stableSteps = 0;
        previousHeight = height;
      }

      if (stableSteps >= stableStepsRequired) break;

      await page.evaluate(() => {
        window.scrollBy(0, Math.max(400, window.innerHeight * 0.8));
      });
      await page.waitForTimeout(pauseMs);
    }

    // Capture all content currently laid out in the document.
    await page.screenshot({ path: 'thumbnail.png', fullPage: true });
    console.log('Saved thumbnail.png');
  } finally {
    await browser.close();
  }
})();

Run it with:

node capture-infinite-scroll.js https://example.com/feed

The script’s height check is intentionally simple. Some sites append items without changing the overall height, virtualize old items out of the DOM, load content after a network request that takes longer than the pause, or keep producing items indefinitely. For those pages, stop on a known item count, a target element, or a manually chosen depth rather than relying on document height alone. Inspect the saved image for blank areas, repeated items, or a cutoff before using it.

Choose a stopping condition that fits the page

  1. Known item count: stop after the desired number of cards or rows are present.
  2. Known boundary element: stop after a footer, pagination marker, or final section appears.
  3. Maximum depth: limit the number of scroll steps for feeds with no end.
  4. Stable content: stop after loading indicators disappear and the page has stopped adding items for a site-appropriate interval.

For a feed that continuously generates more entries, “the whole page” is undefined. Decide whether the thumbnail should represent the first N items, the first few viewport heights, or a particular section. A clear boundary makes captures repeatable and controls page size.

Capture options and practical tradeoffs

Approach Use it when Tradeoff
Full-page screenshot directly All content is already in the DOM or the page is static. May omit sections gated on scrolling.
Scroll first, then full-page capture Scrolling triggers infinite loading, lazy images, or intersection observers. Needs a site-specific wait and stopping rule; scrolling can change page state.
Manual browser scrolling It is a one-off capture and only a few sections are needed. Less repeatable; browser full-page features vary and may not trigger loading automatically.

For a quick manual capture, scroll through the desired content, wait for images and new entries to appear, and then use a browser capture feature that supports full-page output. Do not assume a generic full-page command will simulate visitor scrolling. If you need repeatable results, use browser automation and record the viewport, boundary, and wait conditions alongside the image.

Make the thumbnail representative

  • Set a viewport deliberately. Responsive pages can have different layouts and loading behavior at different widths.
  • Wait for images as well as cards. A newly inserted card may appear before its image finishes loading.
  • Watch for sticky UI. Headers, back-to-top buttons, and floating controls can overlap content or change during the scroll pass.
  • Check for virtualization. Some feeds remove off-screen items to save memory. A final screenshot then may not contain everything previously visited; use a site-specific capture approach or capture sections separately.
  • Keep the loading state consistent. Authentication, consent dialogs, personalization, and network conditions can alter what appears.
  • Inspect the result. Confirm the first and last desired entries are present and that the image has no blank gaps or repeated sections.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns an image or PDF, and its API accepts common screenshot parameter names that other screenshot APIs use. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, no card required.

For infinite-scroll pages, the capture still needs a useful boundary: a screenshot service can capture the rendered page, but a feed with no end requires you to decide how much content to represent.

Troubleshooting

Symptom Likely cause What to change
Screenshot ends after the first screen The page did not grow before capture, or the script stopped early. Check that scrolling triggers loading; increase the step limit or use an item/element-based stopping condition.
Lower sections are blank Content or images were requested but had not finished loading. Wait for a site-specific loading indicator to disappear, or wait for the expected selector/image state before capture.
The capture contains only the latest few items The page virtualizes entries and removes old ones from the DOM. Capture sections as you traverse or use a page-specific method that retains the desired content.
The script runs to the maximum step count The feed keeps adding content or the height never stabilizes. Use a maximum depth or target item count; do not wait for an endless feed to become stable.
Scrolling changes what is shown Sticky elements, animations, ads, or personalization respond to scrolling. Use a consistent viewport and state; disable motion or hide obstructing elements when appropriate for the capture.
Navigation times out The page keeps network connections open or takes longer to reach the chosen load state. Use a less restrictive navigation milestone such as domcontentloaded, then wait for the specific content you need. Avoid assuming all network activity will become idle.

Performance, reliability, and cost

Each scroll-and-wait step adds time, and a long feed can consume significant browser memory when every item stays in the DOM. Set a maximum depth, use a meaningful stopping rule, and avoid fixed pauses longer than the site needs. A shorter pause improves speed but can capture before asynchronous content arrives; there is no universal delay that works for all pages.

For repeatable captures, keep the viewport, target boundary, navigation condition, and site state consistent. Full-page images of very long pages can be large and may be slow to generate or inspect. If the target is a changing feed, record the capture boundary and time so a later run can be compared fairly. Do not interpret a visually complete image as proof that every item in an unbounded feed was captured.

Playwright itself is browser automation software; this workflow has no per-screenshot service price stated in the cited Playwright documentation. Infrastructure costs depend on where and how you run the browser. If you prefer a managed API, ScreenshotNeo offers a free allowance and paid tiers described above; use a defined page boundary for dynamic feeds regardless of capture method.

FAQ

Does fullPage: true scroll through an infinite page?

It captures the full scrollable content currently laid out. It does not guarantee that scroll-triggered content will be requested first.

Can an infinite page have a complete screenshot?

Only after you define what “complete” means for the capture, such as a target number of items or a known endpoint. A feed that can continue forever has no finite full-page boundary.

Will scrolling first affect a visual regression screenshot?

It can. Scrolling may trigger animations, sticky states, lazy loading, or other page changes. Use consistent capture conditions and inspect the output when comparing renders.

Is there a universal wait duration?

No. Loading depends on the site and its network requests. Wait for the relevant page condition and use a practical maximum so the capture cannot run indefinitely.

Sources