ScreenshotNeo

BlogHow-to

Puppeteer Screenshot After Infinite Scroll Stops Adding Items

Use a bounded Puppeteer loop to detect when an infinite feed appears stable, then capture the viewport or accumulated page. Includes selectors, limits, and fixes for virtualized feeds.

By the ScreenshotNeo team4 October 20269 min read

To screenshot a Puppeteer page after an infinite feed stops adding items, scroll in a bounded loop, wait for a page-specific signal such as the item count to increase, and stop only after that signal remains unchanged for a chosen number of rounds. Then call page.screenshot(). Puppeteer has no universal infinite-scroll completion event: the selector, scroll target, and stability window must match the site.

The example below is for an append-only feed whose items match .feed-item and which responds to scrolling the window. Replace that selector and, if needed, the scroll operation with ones that match the page.

1. Install Puppeteer and prepare the page

In a new project, install Puppeteer. Its package downloads a compatible browser by default. If your environment supplies a browser separately, consult the Puppeteer configuration documentation for using that executable.

npm install puppeteer

Save the following as screenshot-feed.js. Set TARGET_URL to a page you are authorized to access.

const puppeteer = require('puppeteer');

const TARGET_URL = process.env.TARGET_URL || 'https://example.com/feed';
const ITEM_SELECTOR = '.feed-item';
const OUTPUT = 'feed.png';
const MAX_ROUNDS = 40;
const STABLE_ROUNDS_REQUIRED = 3;
const CHANGE_TIMEOUT_MS = 5000;
const PAUSE_BETWEEN_ROUNDS_MS = 250;

async function main() {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });
    page.setDefaultNavigationTimeout(30000);
    await page.goto(TARGET_URL, { waitUntil: 'domcontentloaded' });

    // Wait until the feed itself exists. Adapt this if it is rendered later.
    await page.waitForSelector(ITEM_SELECTOR, { timeout: 15000 });

    let previousCount = await page.locator(ITEM_SELECTOR).count();
    let stableRounds = 0;
    let rounds = 0;

    while (stableRounds < STABLE_ROUNDS_REQUIRED && rounds < MAX_ROUNDS) {
      rounds += 1;
      await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));

      try {
        await page.waitForFunction(
          (selector, count) => document.querySelectorAll(selector).length > count,
          { timeout: CHANGE_TIMEOUT_MS, polling: 'mutation' },
          ITEM_SELECTOR,
          previousCount,
        );
        previousCount = await page.locator(ITEM_SELECTOR).count();
        stableRounds = 0;
      } catch (error) {
        // A wait timeout means the count did not increase in this window.
        // Other errors (for example, a closed target) should not be hidden.
        if (error.name !== 'TimeoutError') throw error;
        stableRounds += 1;
      }

      if (PAUSE_BETWEEN_ROUNDS_MS > 0) {
        await new Promise(resolve => setTimeout(resolve, PAUSE_BETWEEN_ROUNDS_MS));
      }
    }

    const stopReason = stableRounds >= STABLE_ROUNDS_REQUIRED
      ? 'stable item count'
      : 'maximum scroll rounds reached';
    console.log(`Stopped after ${rounds} rounds (${stopReason}); observed ${previousCount} items.`);

    // fullPage captures the accumulated document; omit it for viewport only.
    await page.screenshot({ path: OUTPUT, fullPage: true });
    console.log(`Saved ${OUTPUT}`);
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it with TARGET_URL set to the feed URL. For example, in a POSIX shell:

TARGET_URL='https://example.com/feed' node screenshot-feed.js

This loop treats a timeout as one unchanged observation, resets the stability counter when new items appear, and stops after a finite number of rounds. The three-round threshold, five-second wait, and forty-round cap are example settings, not Puppeteer guarantees. Tune them for the site’s load behavior and your runtime budget.

2. Choose a completion signal that fits the feed

The item count works when the site appends new item elements and retains earlier ones. A feed may instead update existing cards, recycle DOM nodes, or load content in a nested scrolling panel. In those cases, a count can report stability while the visible content is still changing.

Signal Useful when Limit
Item count New items are appended and kept in the DOM. Misses replacement or virtualized items.
Stable item IDs Cards have durable identifiers, including in feeds that replace nodes. Requires reading and comparing the site’s identifiers.
Loading indicator The page exposes a reliable spinner or loading state. Some sites remove it between batches or never show one.
End-of-feed marker The page renders an explicit terminal marker. Only works if that marker accurately means there are no more items.
Document height Content growth changes the document’s scroll height. Height can stay constant during replacement, or change for unrelated layout reasons.

For a site with an explicit end marker, check it each round and stop when it appears. For a virtualized feed, compare stable IDs or another page-specific value rather than relying on the number of DOM nodes. Keep the maximum round count even when using a better signal: a broken or changed page should not create an unbounded job.

Scroll the right container

Some feeds scroll the window; others place the list in an element with its own scrollbar. Identify which element actually scrolls. Puppeteer’s interaction guide documents Locator.scroll() for scrolling an element. For a known container, you can also scroll it in page context:

const feed = await page.waitForSelector('.feed-scroll-container');
await feed.evaluate(element => {
  element.scrollTop = element.scrollHeight;
});

Use the container’s item selector and observe its content. Scrolling the window while the feed is inside a panel may never trigger another batch.

Wait for a height change instead

When document growth is the relevant signal, record the previous height and wait for it to grow. This is still only a heuristic; a stable height does not prove a feed is finished.

const oldHeight = await page.evaluate(() => document.documentElement.scrollHeight);
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
try {
  await page.waitForFunction(
    height => document.documentElement.scrollHeight > height,
    { timeout: 5000, polling: 'mutation' },
    oldHeight,
  );
  // Height grew; continue observing.
} catch (error) {
  if (error.name !== 'TimeoutError') throw error;
  // No growth in the observation window; count one stable round.
}

3. Capture the viewport, full page, or an element

page.screenshot() captures the viewport by default. Set fullPage: true to request the full document after scrolling has accumulated its content. Full-page capture cannot restore items that the site has removed from the DOM through virtualization.

// The visible viewport only
await page.screenshot({ path: 'viewport.png' });

// The full accumulated document
await page.screenshot({ path: 'full-page.png', fullPage: true });

To capture one element, use its element handle’s screenshot method. Puppeteer scrolls the element into view if necessary. It throws if the element detaches before capture, which can happen when a virtualized feed recycles nodes.

const card = await page.waitForSelector('.feed-item[data-id="item-123"]');
await card.screenshot({ path: 'item.png' });

If the element may be replaced during the wait, locate it again immediately before capture and handle a missing or detached element explicitly. The Puppeteer guides document the screenshot and element screenshot behavior; see the screenshot guide, waitForFunction API, page interactions guide, and ElementHandle screenshot API.

4. Tune the wait and stopping rules

  • Observation timeout: The sample waits five seconds for a count increase after each scroll. Increase it if the site’s normal response is slower; reduce it for a fast local page.
  • Stable rounds: Several unchanged rounds reduce the chance of stopping during a brief pause. They do not prove the feed is exhausted. Use an end marker if the site provides one.
  • Maximum rounds: This bounds time, page growth, and output size. Choose a limit appropriate to the task and log when it is reached.
  • Polling: waitForFunction accepts raf, mutation, or numeric polling. Mutation polling suits DOM-based changes; use numeric polling if the relevant signal changes without DOM mutations. Its documented default timeout is 30 seconds, but the example supplies a shorter per-wait timeout.
  • Navigation: domcontentloaded avoids waiting for every network resource before starting. If the feed requires a later app bootstrap, wait for its container or a site-specific ready condition.

Do not set a huge per-round timeout and a high round cap without considering their combined worst case. If every round times out, the loop can take roughly the timeout multiplied by the maximum rounds, plus navigation and capture time.

5. Troubleshooting

Symptom Likely cause Fix
The script stops after the first scroll. The selector is wrong, the feed uses a different scroll target, or the first batch has not rendered. Inspect the page markup, wait for the actual feed container, and scroll the element that owns the feed.
It stops while more content exists. The timeout is shorter than the site’s response delay, or the signal misses replacements. Lengthen the observation window and monitor stable IDs, a loading state, or an end marker that matches the page.
It never reaches a stable count. Ads, live updates, or unrelated matching elements keep changing the count. Narrow the selector to feed items and define stability around the relevant feed state.
The full-page image omits earlier feed items. The site virtualizes or removes off-screen nodes. Capture batches as they appear, disable virtualization if the site supports it, or use an explicit export/API when authorized. A full-page screenshot only captures content present in the rendered document.
Element screenshot fails because the node detached. The feed recycled or replaced the element during capture. Re-query by a stable item identifier immediately before capture, or capture a stable parent/container.
Navigation or selector wait times out. The page is slow, blocked, requires authentication, or the selector does not match. Check access and URL, use the correct selector, and set a bounded timeout that reflects the site. Do not treat a timeout as proof the feed is complete.
Browser launch fails in deployment. The environment lacks required browser dependencies or has a mismatched executable. Use a supported Puppeteer installation and browser configuration for that environment; inspect launch errors and configure the executable only when managing the browser separately.

6. Performance, reliability, and cost

Each round adds a scroll action and may wait up to the configured timeout. The worst case is bounded by MAX_ROUNDS, but a generous timeout and cap can still make a job slow. Keep the item selector narrow, use a signal that changes only when the feed changes, and avoid waiting for whole-page network idle if the site keeps analytics or streaming requests open.

Large feeds can consume substantial memory and produce very tall images. Set a maximum round count, consider viewport or per-item captures when a single huge image is unnecessary, and check the output dimensions and file size in your pipeline. For repeatable captures, use the same viewport and page state, and record the stop reason and observed item count so incomplete runs are visible. A stable-window heuristic is probabilistic: sites can pause longer than expected, fail to load, or stop exposing content without a terminal marker.

Running Puppeteer has browser and compute costs in your own environment. The exact cost depends on the runtime and workload; there is no universal benchmark for this page-specific loop.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. For a direct screenshot of a URL, one GET request returns an image or PDF. The API captures a page as requested; it does not run the infinite-scroll loop above to accumulate feed items before capture.

See the ScreenshotNeo API documentation for request options. For example:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers say what happened. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Visit ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.

FAQ

Does Puppeteer know when an infinite feed is finished?

No universal event is documented. Observe a page-specific signal and impose a finite stopping rule.

Should I use network idle to decide when to stop?

Not by itself. A feed can load after a pause, and background requests can prevent network idle. Prefer a signal tied to the feed’s content or state.

Will fullPage: true include every item I scrolled past?

Only if those items remain in the document. Virtualized feeds may remove or recycle off-screen content.

Can this capture an authenticated feed?

Yes, if the browser context has valid access, such as a session established for the authorized account. Protect credentials and avoid saving sensitive screenshots to public locations.