How to Capture Lazy-Loaded Content With Puppeteer
Use Puppeteer to trigger lazy loading, wait for page-specific readiness signals, verify content, and capture complete screenshots reliably.
Lazy-loaded content appears only after a trigger such as scrolling an element into view. A reliable Puppeteer workflow is therefore trigger → wait → verify → extract or capture. Scroll the element that owns the content, wait for a page-specific DOM condition, verify that the expected items exist, then take an element or full-page screenshot.
A full-page screenshot, a fixed delay, or network-idle alone does not guarantee that every lazy-loaded item has appeared. The page decides what event starts loading and which DOM state proves that loading finished.
1. Install Puppeteer
mkdir lazy-capture
cd lazy-capture
npm init -y
npm install puppeteer
The examples below use modern Puppeteer Locator APIs. Puppeteer’s page-interactions guide recommends Locators for selecting and interacting with elements, and documents Locator.scroll() as using mouse-wheel events. See the page interactions guide and screenshots guide.
2. A complete incremental scrolling example
This script scrolls a feed, waits for the card count to grow, stops when a terminal marker appears or when repeated scrolls produce no growth, and saves both extracted text and a full-page screenshot. Replace the selectors and stopping condition with ones from your target page.
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({headless: true});
const page = await browser.newPage();
await page.setViewport({width: 1440, height: 900, deviceScaleFactor: 1});
const url = 'https://example.com/feed';
await page.goto(url, {waitUntil: 'domcontentloaded', timeout: 60_000});
const cardSelector = '.card';
const endSelector = '.feed-end';
const maxRounds = 40;
const scrollStep = 700;
const noGrowthLimit = 3;
let noGrowthRounds = 0;
let previousCount = await page.locator(cardSelector).count();
for (let round = 0; round < maxRounds; round++) {
// Scroll the owner of the lazy content. Use a nested selector here
// if the feed, rather than the window, is scrollable.
await page.locator('body').scroll({scrollTop: scrollStep});
try {
await page.waitForNetworkIdle({idleTime: 500, concurrency: 2, timeout: 10_000});
} catch (_) {
// Some pages keep analytics or streaming requests open. Continue
// when the DOM condition below is the useful readiness signal.
}
const grew = await page.waitForFunction(
(selector, oldCount) => document.querySelectorAll(selector).length > oldCount,
{timeout: 5_000},
cardSelector,
previousCount
).then(() => true).catch(() => false);
const currentCount = await page.locator(cardSelector).count();
if (currentCount > previousCount || grew) {
noGrowthRounds = 0;
previousCount = currentCount;
} else {
noGrowthRounds++;
}
const reachedEnd = await page.locator(endSelector).count() > 0;
if (reachedEnd || noGrowthRounds >= noGrowthLimit) break;
}
// Verify the result before capture.
const cards = await page.locator(cardSelector).map(items =>
items.map(item => item.textContent?.trim() || '')
).wait();
if (cards.length === 0) {
throw new Error('No cards loaded; check the selector, scroll target, and trigger.');
}
require('fs').writeFileSync('cards.json', JSON.stringify(cards, null, 2));
await page.screenshot({path: 'feed.png', fullPage: true});
await browser.close();
})();
The snippet is an adaptable pattern, not a universal infinite-scroll algorithm. Confirm that .card, .feed-end, and the scroll target are stable on the actual page.
3. Choosing the correct scroll target
First determine whether the content loads from window scrolling or from a nested scroll container.
Window or document scrolling
await page.locator('body').scroll({scrollTop: 700});
You can also use browser JavaScript for pages whose scroll behavior is unusual:
await page.evaluate(() => window.scrollBy(0, window.innerHeight));
Nested scroll container
const feed = page.locator('.feed-scroll-region');
await feed.scroll({scrollTop: 700});
Scrolling the window when .feed-scroll-region owns the overflow will not fire the trigger that the feed expects. Inspect computed styles or manually scroll the page to identify the element with overflow: auto or overflow: scroll.
Scroll a specific item into view
await page.locator('.card:nth-child(10)').scroll();
This is useful when an intersection observer loads content as a particular card enters the viewport.
4. Wait for evidence that content loaded
Use a condition tied to the requested content. Good signals include an item count, a visible selector, a changed attribute, a “no more results” marker, or a page-specific status element.
Wait for an element
await page.locator('.results .card').wait();
Wait for a count to increase
const before = await page.locator('.card').count();
await page.waitForFunction(
(selector, oldCount) => document.querySelectorAll(selector).length > oldCount,
{timeout: 10_000},
'.card',
before
);
Wait for a page-specific function
await page.waitForFunction(() => {
const status = document.querySelector('[data-loading-status]');
return status?.getAttribute('data-loading-status') === 'ready';
});
Wait for network idle as a secondary signal
await page.waitForNetworkIdle({idleTime: 500, concurrency: 2});
Page.waitForNetworkIdle() waits for network idleness and always waits at least the configured idle period. The documented default idleTime is 500 ms. Network idle means requests are quiet; it does not prove that the desired card or image exists. Combine it with a DOM check.
5. Handle images that load lazily
Image lazy loading often uses loading="lazy", an intersection observer, or a custom data attribute. Scroll the image into view, then wait for its src or completion state.
const image = page.locator('.gallery img').first();
await image.scroll();
await page.waitForFunction(() => {
const img = document.querySelector('.gallery img');
return img instanceof HTMLImageElement && img.complete && img.naturalWidth > 0;
});
For pages that replace data-src with src:
await page.waitForFunction(() => {
const images = [...document.querySelectorAll('.gallery img')];
return images.length > 0 && images.every(img => img.getAttribute('src'));
});
Some sites use responsive srcset, CSS backgrounds, or canvas rendering. In those cases, verify the rendered result or a page-specific loaded class instead of assuming that a src attribute is sufficient.
6. Capture an element or the entire page
Capture one element
const chart = page.locator('#chart');
await chart.wait();
await chart.screenshot({path: 'chart.png'});
Puppeteer’s element screenshot API scrolls the element into view if necessary.
Capture the full page
await page.screenshot({
path: 'full-page.png',
fullPage: true,
type: 'png'
});
fullPage: true changes the capture extent. It is not documented as a universal lazy-loading trigger, so complete the trigger-and-verify loop first.
Capture JPEG or WebP
await page.screenshot({path: 'page.jpg', type: 'jpeg', quality: 85});
await page.screenshot({path: 'page.webp', type: 'webp'});
7. Infinite scroll completion strategies
| Strategy | Use when | Risk |
|---|---|---|
| Known item count | The page or API tells you exactly how many records are required | Count can include placeholders or recycled nodes |
| Terminal marker | The page renders “end of results” or disables a load-more control | Marker may appear before images finish loading |
| No-growth threshold | No explicit end state exists | Temporary network delays can look like completion |
| Load-more button | Content appears only after a click | Button may be covered, detached, or rate limited |
while (true) {
const oldCount = await page.locator('.card').count();
const button = page.locator('button.load-more');
if (await button.count() === 0) break;
await button.click();
await page.waitForFunction(
(selector, count) => document.querySelectorAll(selector).length > count,
{timeout: 10_000},
'.card', oldCount
);
}
8. Virtualized lists and replaced DOM nodes
Virtualized lists may keep only visible rows in the DOM and recycle those nodes as you scroll. In that case, the current DOM count is not the total number visited. Store each item’s stable identifier or text as you go, and stop using a page-specific total, cursor, sentinel, or no-growth rule.
const seen = new Map();
for (let round = 0; round < 50; round++) {
const visible = await page.locator('.row').map(rows =>
rows.map(row => ({
id: row.getAttribute('data-id'),
text: row.textContent?.trim() || ''
}))
).wait();
for (const item of visible) if (item.id) seen.set(item.id, item);
await page.locator('.virtual-list').scroll({scrollTop: 600});
}
console.log([...seen.values()]);
Validate this approach against the target page because Puppeteer’s general documentation does not define one universal recipe for virtualized feeds.
9. Complete runnable variants
cURL
For a direct HTTP screenshot request, use the ScreenshotNeo API documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot contains only the first items | Wrong scroll target or no trigger | Find the element with overflow scrolling and scroll it incrementally. |
| Network-idle wait times out | Analytics, WebSockets, or polling keep requests active | Use a DOM readiness condition; tune concurrency and timeout only as a secondary measure. |
| Count never increases | Selector is wrong, content is virtualized, or loading requires a click | Inspect the DOM, capture stable IDs, and reproduce the real trigger. |
| Images are blank | Images have not entered view, failed, or are background images | Scroll each region, wait for complete/naturalWidth or a page-specific loaded state, and inspect failed requests. |
| Full-page capture misses content | Full-page extent did not trigger the site’s lazy loader | Run the incremental trigger-and-verify loop before calling fullPage: true. |
| Element is detached | Framework rerendered or recycled the node | Use a Locator and reacquire it after each load; avoid holding stale element handles. |
| Content differs between runs | Timing, viewport, cookies, geolocation, or personalized data differ | Set a fixed viewport and relevant headers/cookies, and wait on deterministic page evidence. |
| Navigation hangs | Long-running requests or a blocked resource | Use a navigation timeout, then rely on a specific selector rather than waiting for every request. |
11. Performance, reliability, and cost
- Performance: Scroll in viewport-sized increments and stop as soon as the required count or terminal condition is reached. Avoid an unnecessarily large fixed delay after every scroll.
- Reliability: Use deterministic selectors, bounded rounds, explicit timeouts, and a no-growth limit. Log the round number, item count, and final stopping reason.
- Capture size: Element screenshots are usually smaller and faster than full-page captures. Use full-page mode only when the complete document is required.
- Page behavior: Ads, trackers, polling, animations, and consent dialogs can change readiness. Freeze animations with custom CSS when visual stability matters.
- Cost: Puppeteer itself is open-source software, but your browser runtime, compute, bandwidth, and any proxy or hosting service have their own costs. Set limits so an endlessly growing feed cannot consume unbounded resources.
12. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. Its capture options include full-page screenshots with lazy images loaded, custom waits, CSS and JavaScript, selectors to hide, request blocking, device presets, viewport control, and caching.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Read the API documentation and sign up free.
13. FAQ
Does fullPage: true load all lazy content?
No. It captures the full page extent, but the page’s own lazy-loading trigger may never run. Scroll and verify first.
Should I always wait for network idle?
No. Network idle is useful as a secondary signal. A selector, count, attribute, or other DOM condition tied to the requested content is stronger evidence.
How much should I scroll each time?
Use a viewport-sized or otherwise suitable increment for the page. There is no universal value; inspect how that page’s trigger works.
Why does the item count stay constant?
The feed may be virtualized, the selector may be unstable, the wrong container may be scrolling, or loading may require a click or another interaction.
When should I capture an element instead of the page?
Use an element screenshot when one chart, card, or panel is the deliverable. Use fullPage: true when the entire document is required.


