ScreenshotNeo

BlogHow-to

Puppeteer Screenshot of an Infinite Scroll Page Without Duplicate Sections

Load and verify an infinite feed before capturing it with Puppeteer. Learn how to detect duplicate items, handle nested scrollers and virtualization, and troubleshoot full-page screenshots.

By the ScreenshotNeo team4 October 202611 min read

Short answer: load the feed before capturing it. Puppeteer’s fullPage: true captures the page content available at screenshot time; it does not scroll an infinite feed to make the site load more items. Scroll the correct container in bounded steps, wait for page-specific evidence of new items or an end marker, check that rendered items are unique, then take the screenshot.

There is no universal infinite-scroll loop or duplicate-removal switch. The right selectors, scroll container, and stopping condition depend on the site. Puppeteer provides the browser-side primitives to implement that page-specific logic: evaluate(), waitForFunction(), and screenshot(). Puppeteer’s screenshot guide documents Page.screenshot(); its ScreenshotOptions defines fullPage as capturing the full page.

1. Why infinite-scroll screenshots miss or repeat content

A normal page screenshot reflects the current rendered page. On an infinite-scroll site, additional content may only load after scrolling reaches a trigger, often through an intersection observer or scroll handler. A full-page capture does not itself trigger every such load. A generic navigation condition such as networkidle2 can help the initial page settle, but it does not prove that the feed is finished: the next request may not begin until you scroll again.

Repeated sections have several possible causes. The page might already have duplicate records in its DOM; the script might capture while content is still changing; the page might use a nested scroller; or the feed might virtualize its items, removing older ones as newer ones appear. The Puppeteer APIs do not prescribe a universal fix. Inspect the rendered DOM and the target site’s behavior before choosing a correction.

2. Identify the feed, its scroll container, and a completion signal

Before automating the page, inspect it in a browser and determine:

  • Item selector: a selector that matches one feed item, such as [data-feed-item].
  • Stable item key: an attribute or child value that uniquely identifies an item, such as data-id. Prefer this over text when available.
  • Scroll container: the document, or a particular element with its own scrollbar.
  • Completion signal: an end-of-feed marker, a known target item count, or a defensible rule that no new unique items appeared for several scroll rounds.
  • Virtualization: whether older items disappear from the DOM as you move down the feed.

Use a page-specific completion signal where possible. Network inactivity is a settling condition, not a documented indication that an infinite feed is complete. Puppeteer’s Page API documents the relevant evaluation and waiting methods.

3. Complete Node.js example: bounded scroll, unique-item check, screenshot

This example uses the document as the scroll container and [data-feed-item] / data-id as example selectors. Replace those with selectors and semantics from the page you are capturing. The loop has a maximum number of scrolls, requires multiple stable rounds before stopping, and reports when it reaches its limit. It checks for duplicate identifiers in the current DOM; it does not silently remove page content.

// Save as capture.mjs
// Install with: npm install puppeteer
import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com/feed';
const itemSelector = '[data-feed-item]';
const itemIdAttribute = 'data-id';
const endSelector = '[data-feed-end]'; // Set to a real end marker, or leave absent.
const maxScrolls = 40;
const stableRoundsToStop = 3;
const settleMs = 800;

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage({
    viewport: { width: 1365, height: 900 },
  });

  await page.goto(url, { waitUntil: 'networkidle2', timeout: 60_000 });
  await page.waitForSelector(itemSelector, { timeout: 15_000 });

  let previousUniqueCount = -1;
  let stableRounds = 0;
  let reachedEnd = false;
  let reachedLimit = true;

  for (let scroll = 0; scroll < maxScrolls; scroll++) {
    const before = await page.$$eval(itemSelector, (items, idAttribute) => {
      const keys = items.map(item => {
        const id = item.getAttribute(idAttribute);
        return id || item.textContent.trim();
      }).filter(Boolean);
      return { total: items.length, unique: new Set(keys).size };
    }, itemIdAttribute);

    await page.evaluate(() => {
      window.scrollTo(0, document.documentElement.scrollHeight);
    });

    // Wait for a count increase, an end marker, or a bounded timeout.
    // A timeout here is not proof that the feed is complete.
    try {
      await page.waitForFunction(
        ({ selector, end, count }) => {
          const hasEnd = end ? Boolean(document.querySelector(end)) : false;
          return hasEnd || document.querySelectorAll(selector).length > count;
        },
        { timeout: 10_000 },
        { selector: itemSelector, end: endSelector, count: before.total },
      );
    } catch {
      // No count change before timeout: evaluate the current state below.
    }

    await new Promise(resolve => setTimeout(resolve, settleMs));

    const state = await page.$$eval(
      itemSelector,
      (items, idAttribute) => {
        const keys = items.map(item => {
          const id = item.getAttribute(idAttribute);
          return id || item.textContent.trim();
        }).filter(Boolean);
        return {
          total: items.length,
          unique: new Set(keys).size,
          duplicateKeys: keys.filter((key, index) => keys.indexOf(key) !== index),
        };
      },
      itemIdAttribute,
    );

    reachedEnd = endSelector ? await page.$(endSelector) !== null : false;
    if (state.duplicateKeys.length) {
      console.warn('Duplicate item keys are already present in the rendered DOM.');
      console.warn([...new Set(state.duplicateKeys)]);
    }

    if (state.unique === previousUniqueCount) {
      stableRounds++;
    } else {
      stableRounds = 0;
    }
    previousUniqueCount = state.unique;

    if (reachedEnd || stableRounds >= stableRoundsToStop) {
      reachedLimit = false;
      break;
    }
  }

  // Give layout and lazy-loaded assets another bounded chance to settle.
  await new Promise(resolve => setTimeout(resolve, settleMs));
  const finalState = await page.$$eval(itemSelector, items => ({
    renderedItems: items.length,
    pageHeight: document.documentElement.scrollHeight,
  }));
  console.log({ reachedEnd, reachedLimit, ...finalState });

  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with node capture.mjs https://your-site.example/feed. The sample is a control-flow pattern, not a universal or tested recipe for arbitrary sites. In particular, a stable unique count can mean the feed ended, the scroll trigger failed, a request failed, or the page virtualized its DOM. Use the page’s own end marker or known expected count when available.

Make the stopping condition stronger

  • End marker: stop when the site renders a known end-of-feed element. This is usually clearer than guessing from a pause.
  • Expected count: stop when the number of unique items reaches a known target.
  • Stable rounds: if the site exposes no completion signal, stop after several scroll rounds with no increase in unique items and enforce a hard maximum. Treat the result as potentially incomplete.
  • Request or cursor state: if you control the application, use its pagination state or API response as the completion signal rather than inferring completion from pixels.

4. Nested scroll containers

If a panel or feed element scrolls independently, window.scrollTo() will not reach its bottom. Select the actual container, scroll that element, and measure its own scroll position. For example, replace the document scroll step with:

const containerSelector = '.feed-scroll-panel';
await page.waitForSelector(containerSelector);

await page.$eval(containerSelector, element => {
  element.scrollTop = element.scrollHeight;
});

For a stronger wait, compare the item count before and after scrolling as in the main example, while querying items inside that container. If the site has multiple matching panels, scope both the item selector and end marker to the intended feed. Avoid scrolling both the window and a nested container unless the page requires it; doing so can trigger unrelated handlers and make diagnosis harder.

5. Virtualized feeds and when one full-page image cannot work

Some feeds keep only the visible items and a small buffer in the DOM. As you scroll, earlier nodes are removed and reused. In that case, fullPage: true cannot capture content that is no longer rendered as part of the page. A DOM count that stays flat also does not mean the feed stopped: the visible nodes may be changing while the number of nodes remains constant.

Check whether item identifiers change while the DOM count stays similar, and whether scrolling back up restores earlier items. If the feed is virtualized, consider these options:

  1. Use the application’s own export or data endpoint if you need a complete record of every item.
  2. Capture sequential viewport images and combine them in a workflow designed for that page, checking overlaps by stable item IDs. Sticky headers, changing content, and variable-height items require special handling.
  3. If you control the page, use a capture mode that renders all items or temporarily disables virtualization before taking a full-page screenshot.

Do not assume a single full-page screenshot preserves items the application has removed from the DOM.

6. Check for duplicates before blaming the screenshot

Compare item keys in the live DOM immediately before capture. If the same stable key appears multiple times, the duplicate is already in the rendered feed. Investigate the site’s append or pagination behavior; deleting repeated nodes in the capture script can hide real page behavior and may remove legitimate content if the key is not truly unique.

If the DOM contains unique items but the image appears to repeat a section, inspect timing, scroll-container selection, page height, fixed or sticky elements, and whether the page changed during capture. Save a DOM snapshot or log item keys after each iteration to locate the first point where duplication appears. Puppeteer documents full-page extent, but it does not define how a site manages feed records or guarantee deduplication.

7. Screenshot options and output choices

Option or approach Use Limit
fullPage: true Capture the full rendered page in one image. Does not trigger infinite loading or restore virtualized content.
path Save the screenshot to a file. The extension can determine image type. Choose an output path writable by the process.
type Choose a supported image format explicitly when needed. Do not assume format settings fix missing content.
quality Set lossy image quality where supported. Not applicable to PNG.
clip Capture a selected rectangle instead of relying on full-page extent. Coordinates and dimensions must match the desired region.
Element screenshot Capture a specific element with its screenshot method. It captures the element, not automatically the entire infinite feed.

These meanings come from Puppeteer’s ScreenshotOptions documentation and screenshot guide. Select the capture extent only after the content-loading step is complete.

8. Troubleshooting

Symptom Likely cause What to change
Screenshot contains only the first screen The feed was never scrolled, or the wrong container was scrolled. Identify the actual scroll element and drive it in bounded increments before capture.
Screenshot stops before the feed ends The loop stopped on a weak stability heuristic, a request failed, or the maximum was too low. Prefer a real end marker or expected count; log counts and whether the hard limit was reached.
Repeated cards are visible Duplicates may already be present in the DOM, or the capture ran while layout/content changed. Log stable IDs per round, inspect the DOM, and wait for a page-specific settled state before capture.
Scroll loop never sees a count increase Wrong item selector, nested scroller, failed request, or virtualized list with a constant DOM size. Inspect the selector and scroll position; track item IDs, not only node count.
waitForSelector times out The selector is incorrect, content is in a frame, or loading did not reach the expected state. Check the selector in the page, inspect frames, and wait for an element the page actually renders.
Navigation times out with networkidle2 Long-running requests can prevent the selected navigation condition from occurring. Use a navigation condition suited to the page, then wait for a concrete feed element. A weaker navigation wait does not mean the feed is ready.
Old items disappear from the screenshot The site virtualizes the feed and removed earlier nodes. Use sequential captures or an application export; one full-page capture cannot include removed nodes.
Screenshot has blank or shifting regions Images, fonts, or layout had not settled, or content changed during capture. Wait for the relevant page-specific signal and a bounded settling period; inspect the DOM and layout before capture.
Screenshot is unexpectedly huge or memory-heavy Very long pages require a large image surface and substantial browser memory. Limit the target range, capture in sections, or use a format and resolution appropriate to the output.

9. Performance, reliability, and cost

Every scroll round, page wait, and full-page image adds work. Set a maximum number of scrolls and a timeout for page-specific waits so an endless feed cannot run indefinitely. Use a stable viewport for repeatable layout, and avoid loading more content than the screenshot needs. Very tall pages can consume significant memory during rendering and image encoding; sequential viewport captures may be more practical, especially for virtualized feeds.

Reliability comes from observing the page rather than choosing a large fixed delay. A delay can give images and layout time to settle, but it cannot prove that content loaded. Log the scroll round, total item count, unique item count, end-marker state, and whether the maximum was reached. For important captures, retain enough diagnostics to distinguish a genuinely exhausted feed from a stalled or virtualized one.

With a self-hosted Puppeteer script, cost depends on the browser runtime and infrastructure you operate; this guide makes no benchmark or price claim. If you use a screenshot API instead, check its billing behavior for failed loads and bot checks, along with the features and volume limits that apply to your use case.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. For a page that can be captured with one request, call the API with the target URL. See the ScreenshotNeo API documentation for request options and setup.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses report the page verdict and billing status in headers.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card.

11. Frequently asked questions

Does fullPage: true scroll through an infinite page?

No. It sets the screenshot extent to the full page as rendered when the screenshot is taken. Trigger and verify the page’s loading behavior first.

Is networkidle2 enough to know the feed is complete?

No. A feed may wait for another scroll before requesting more items. Use a page-specific end condition or bounded unique-item check; network idle can be a settling aid.

Can Puppeteer automatically remove duplicate sections?

The documented screenshot and page APIs do not provide a universal duplicate-removal option. Check whether duplicate records are already in the DOM and diagnose the page’s feed behavior.

Will one full-page screenshot preserve a virtualized list?

Not if earlier items have been removed from the DOM. Use sequential captures or a data/export path that preserves the full item set.