ScreenshotNeo

BlogHow-to

How to Capture Lazy-Loaded Content With Puppeteer

Use Puppeteer to trigger lazy loading, wait for page-specific readiness signals, verify content, and capture complete screenshots reliably.

By the ScreenshotNeo team1 October 20268 min read

Lazy-loaded content appears only after a trigger such as scrolling an element into view. A reliable Puppeteer workflow is therefore trigger → wait → verify → extract or capture. Scroll the element that owns the content, wait for a page-specific DOM condition, verify that the expected items exist, then take an element or full-page screenshot.

A full-page screenshot, a fixed delay, or network-idle alone does not guarantee that every lazy-loaded item has appeared. The page decides what event starts loading and which DOM state proves that loading finished.

1. Install Puppeteer

mkdir lazy-capture
cd lazy-capture
npm init -y
npm install puppeteer

The examples below use modern Puppeteer Locator APIs. Puppeteer’s page-interactions guide recommends Locators for selecting and interacting with elements, and documents Locator.scroll() as using mouse-wheel events. See the page interactions guide and screenshots guide.

2. A complete incremental scrolling example

This script scrolls a feed, waits for the card count to grow, stops when a terminal marker appears or when repeated scrolls produce no growth, and saves both extracted text and a full-page screenshot. Replace the selectors and stopping condition with ones from your target page.

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({headless: true});
  const page = await browser.newPage();
  await page.setViewport({width: 1440, height: 900, deviceScaleFactor: 1});

  const url = 'https://example.com/feed';
  await page.goto(url, {waitUntil: 'domcontentloaded', timeout: 60_000});

  const cardSelector = '.card';
  const endSelector = '.feed-end';
  const maxRounds = 40;
  const scrollStep = 700;
  const noGrowthLimit = 3;
  let noGrowthRounds = 0;
  let previousCount = await page.locator(cardSelector).count();

  for (let round = 0; round < maxRounds; round++) {
    // Scroll the owner of the lazy content. Use a nested selector here
    // if the feed, rather than the window, is scrollable.
    await page.locator('body').scroll({scrollTop: scrollStep});

    try {
      await page.waitForNetworkIdle({idleTime: 500, concurrency: 2, timeout: 10_000});
    } catch (_) {
      // Some pages keep analytics or streaming requests open. Continue
      // when the DOM condition below is the useful readiness signal.
    }

    const grew = await page.waitForFunction(
      (selector, oldCount) => document.querySelectorAll(selector).length > oldCount,
      {timeout: 5_000},
      cardSelector,
      previousCount
    ).then(() => true).catch(() => false);

    const currentCount = await page.locator(cardSelector).count();
    if (currentCount > previousCount || grew) {
      noGrowthRounds = 0;
      previousCount = currentCount;
    } else {
      noGrowthRounds++;
    }

    const reachedEnd = await page.locator(endSelector).count() > 0;
    if (reachedEnd || noGrowthRounds >= noGrowthLimit) break;
  }

  // Verify the result before capture.
  const cards = await page.locator(cardSelector).map(items =>
    items.map(item => item.textContent?.trim() || '')
  ).wait();

  if (cards.length === 0) {
    throw new Error('No cards loaded; check the selector, scroll target, and trigger.');
  }

  require('fs').writeFileSync('cards.json', JSON.stringify(cards, null, 2));
  await page.screenshot({path: 'feed.png', fullPage: true});
  await browser.close();
})();

The snippet is an adaptable pattern, not a universal infinite-scroll algorithm. Confirm that .card, .feed-end, and the scroll target are stable on the actual page.

3. Choosing the correct scroll target

First determine whether the content loads from window scrolling or from a nested scroll container.

Window or document scrolling

await page.locator('body').scroll({scrollTop: 700});

You can also use browser JavaScript for pages whose scroll behavior is unusual:

await page.evaluate(() => window.scrollBy(0, window.innerHeight));

Nested scroll container

const feed = page.locator('.feed-scroll-region');
await feed.scroll({scrollTop: 700});

Scrolling the window when .feed-scroll-region owns the overflow will not fire the trigger that the feed expects. Inspect computed styles or manually scroll the page to identify the element with overflow: auto or overflow: scroll.

Scroll a specific item into view

await page.locator('.card:nth-child(10)').scroll();

This is useful when an intersection observer loads content as a particular card enters the viewport.

4. Wait for evidence that content loaded

Use a condition tied to the requested content. Good signals include an item count, a visible selector, a changed attribute, a “no more results” marker, or a page-specific status element.

Wait for an element

await page.locator('.results .card').wait();

Wait for a count to increase

const before = await page.locator('.card').count();
await page.waitForFunction(
  (selector, oldCount) => document.querySelectorAll(selector).length > oldCount,
  {timeout: 10_000},
  '.card',
  before
);

Wait for a page-specific function

await page.waitForFunction(() => {
  const status = document.querySelector('[data-loading-status]');
  return status?.getAttribute('data-loading-status') === 'ready';
});

Wait for network idle as a secondary signal

await page.waitForNetworkIdle({idleTime: 500, concurrency: 2});

Page.waitForNetworkIdle() waits for network idleness and always waits at least the configured idle period. The documented default idleTime is 500 ms. Network idle means requests are quiet; it does not prove that the desired card or image exists. Combine it with a DOM check.

5. Handle images that load lazily

Image lazy loading often uses loading="lazy", an intersection observer, or a custom data attribute. Scroll the image into view, then wait for its src or completion state.

const image = page.locator('.gallery img').first();
await image.scroll();
await page.waitForFunction(() => {
  const img = document.querySelector('.gallery img');
  return img instanceof HTMLImageElement && img.complete && img.naturalWidth > 0;
});

For pages that replace data-src with src:

await page.waitForFunction(() => {
  const images = [...document.querySelectorAll('.gallery img')];
  return images.length > 0 && images.every(img => img.getAttribute('src'));
});

Some sites use responsive srcset, CSS backgrounds, or canvas rendering. In those cases, verify the rendered result or a page-specific loaded class instead of assuming that a src attribute is sufficient.

6. Capture an element or the entire page

Capture one element

const chart = page.locator('#chart');
await chart.wait();
await chart.screenshot({path: 'chart.png'});

Puppeteer’s element screenshot API scrolls the element into view if necessary.

Capture the full page

await page.screenshot({
  path: 'full-page.png',
  fullPage: true,
  type: 'png'
});

fullPage: true changes the capture extent. It is not documented as a universal lazy-loading trigger, so complete the trigger-and-verify loop first.

Capture JPEG or WebP

await page.screenshot({path: 'page.jpg', type: 'jpeg', quality: 85});
await page.screenshot({path: 'page.webp', type: 'webp'});

7. Infinite scroll completion strategies

Strategy Use when Risk
Known item count The page or API tells you exactly how many records are required Count can include placeholders or recycled nodes
Terminal marker The page renders “end of results” or disables a load-more control Marker may appear before images finish loading
No-growth threshold No explicit end state exists Temporary network delays can look like completion
Load-more button Content appears only after a click Button may be covered, detached, or rate limited
while (true) {
  const oldCount = await page.locator('.card').count();
  const button = page.locator('button.load-more');
  if (await button.count() === 0) break;

  await button.click();
  await page.waitForFunction(
    (selector, count) => document.querySelectorAll(selector).length > count,
    {timeout: 10_000},
    '.card', oldCount
  );
}

8. Virtualized lists and replaced DOM nodes

Virtualized lists may keep only visible rows in the DOM and recycle those nodes as you scroll. In that case, the current DOM count is not the total number visited. Store each item’s stable identifier or text as you go, and stop using a page-specific total, cursor, sentinel, or no-growth rule.

const seen = new Map();
for (let round = 0; round < 50; round++) {
  const visible = await page.locator('.row').map(rows =>
    rows.map(row => ({
      id: row.getAttribute('data-id'),
      text: row.textContent?.trim() || ''
    }))
  ).wait();
  for (const item of visible) if (item.id) seen.set(item.id, item);
  await page.locator('.virtual-list').scroll({scrollTop: 600});
}
console.log([...seen.values()]);

Validate this approach against the target page because Puppeteer’s general documentation does not define one universal recipe for virtualized feeds.

9. Complete runnable variants

cURL

For a direct HTTP screenshot request, use the ScreenshotNeo API documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);

10. Troubleshooting

Symptom Likely cause Fix
Screenshot contains only the first items Wrong scroll target or no trigger Find the element with overflow scrolling and scroll it incrementally.
Network-idle wait times out Analytics, WebSockets, or polling keep requests active Use a DOM readiness condition; tune concurrency and timeout only as a secondary measure.
Count never increases Selector is wrong, content is virtualized, or loading requires a click Inspect the DOM, capture stable IDs, and reproduce the real trigger.
Images are blank Images have not entered view, failed, or are background images Scroll each region, wait for complete/naturalWidth or a page-specific loaded state, and inspect failed requests.
Full-page capture misses content Full-page extent did not trigger the site’s lazy loader Run the incremental trigger-and-verify loop before calling fullPage: true.
Element is detached Framework rerendered or recycled the node Use a Locator and reacquire it after each load; avoid holding stale element handles.
Content differs between runs Timing, viewport, cookies, geolocation, or personalized data differ Set a fixed viewport and relevant headers/cookies, and wait on deterministic page evidence.
Navigation hangs Long-running requests or a blocked resource Use a navigation timeout, then rely on a specific selector rather than waiting for every request.

11. Performance, reliability, and cost

  • Performance: Scroll in viewport-sized increments and stop as soon as the required count or terminal condition is reached. Avoid an unnecessarily large fixed delay after every scroll.
  • Reliability: Use deterministic selectors, bounded rounds, explicit timeouts, and a no-growth limit. Log the round number, item count, and final stopping reason.
  • Capture size: Element screenshots are usually smaller and faster than full-page captures. Use full-page mode only when the complete document is required.
  • Page behavior: Ads, trackers, polling, animations, and consent dialogs can change readiness. Freeze animations with custom CSS when visual stability matters.
  • Cost: Puppeteer itself is open-source software, but your browser runtime, compute, bandwidth, and any proxy or hosting service have their own costs. Set limits so an endlessly growing feed cannot consume unbounded resources.

12. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. Its capture options include full-page screenshots with lazy images loaded, custom waits, CSS and JavaScript, selectors to hide, request blocking, device presets, viewport control, and caching.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Read the API documentation and sign up free.

13. FAQ

Does fullPage: true load all lazy content?

No. It captures the full page extent, but the page’s own lazy-loading trigger may never run. Scroll and verify first.

Should I always wait for network idle?

No. Network idle is useful as a secondary signal. A selector, count, attribute, or other DOM condition tied to the requested content is stronger evidence.

How much should I scroll each time?

Use a viewport-sized or otherwise suitable increment for the page. There is no universal value; inspect how that page’s trigger works.

Why does the item count stay constant?

The feed may be virtualized, the selector may be unstable, the wrong container may be scrolling, or loading may require a click or another interaction.

When should I capture an element instead of the page?

Use an element screenshot when one chart, card, or panel is the deliverable. Use fullPage: true when the entire document is required.