Puppeteer: Wait for Infinite Scroll to Finish Before Taking a Screenshot
Puppeteer cannot know when infinite scroll is done by itself. Trigger loading, wait for a page-specific completion signal, then capture the settled page.
Short answer: Puppeteer has no built-in wait that can determine when a site’s infinite scroll has finished. Scroll the page to trigger loading, then wait for a page-specific signal such as an end marker or expected item count. If the site exposes no such signal, use a bounded stability heuristic, such as unchanged item count and document height across several scroll cycles. After the condition succeeds and required images are ready, capture with page.screenshot({ fullPage: true }).
fullPage: true captures the current document’s full height. It does not trigger future loads or prove that the page has reached its end. Puppeteer’s screenshot guide shows Page.screenshot(); its waitForFunction API lets you wait for a browser-side condition to become truthy.
1. Choose the right completion signal
First inspect how the target page loads content. Infinite feeds may listen to document scrolling, use a nested scroll panel, expose a “Load more” button, or add an explicit end-of-results marker. The signal should match that behavior.
| Page behavior | Best completion signal | Limitation |
|---|---|---|
| Known number of results | Rendered item count reaches the expected total | You need the correct item selector and expected count. |
| Explicit terminal marker | Marker appears or becomes visible | Confirm that it means the feed is exhausted, not merely that a batch ended. |
| Items append to the document | Item count and document height stay stable across scroll cycles | Heuristic: delayed content may arrive after stability. |
| Nested scroll panel | Panel’s own scroll position, height, and item count | Scrolling the window will not trigger this panel. |
| Load-more control | Button becomes disabled, disappears, or terminal marker appears | Handle loading and error states as well as the end state. |
| Persistent network traffic | Application state or rendered content | Network idle alone may never happen, or may occur between batches. |
Prefer an explicit application signal where possible. Use network idle as a short settling aid after triggering a batch, not as proof that no more results exist. Puppeteer documents waitForNetworkIdle() as a network-quiet wait; a quiet interval can happen between scroll-triggered requests.
2. Install Puppeteer and run a bounded document-scroll example
This example works with a page that appends results to the document when the window is scrolled. It waits for stable height and item count over multiple rounds, and it throws an error if the safety limit is reached. Replace the URL and item selector with values from the page you are capturing.
npm install puppeteer
// save as capture.mjs; run with: node capture.mjs
import puppeteer from 'puppeteer';
const url = 'https://example.com/feed';
const itemSelector = '[data-testid="feed-item"]';
async function waitForFeedToStabilize(page, {
stableRounds = 3,
maxRounds = 30,
pauseMs = 500,
} = {}) {
let previous = null;
let stable = 0;
for (let round = 0; round < maxRounds; round++) {
await page.evaluate(() => {
window.scrollTo(0, document.documentElement.scrollHeight);
});
await new Promise(resolve => setTimeout(resolve, pauseMs));
// Useful when a batch request finishes promptly. Some pages keep
// connections open, so network idle is optional and bounded.
try {
await page.waitForNetworkIdle({ idleTime: 500, timeout: 3000 });
} catch {
// The content condition below remains the stopping rule.
}
const current = await page.evaluate(selector => ({
height: document.documentElement.scrollHeight,
count: document.querySelectorAll(selector).length,
}), itemSelector);
if (previous && current.height === previous.height &&
current.count === previous.count) {
stable++;
} else {
stable = 0;
}
if (stable >= stableRounds) return current;
previous = current;
}
throw new Error(`Feed did not stabilize within ${maxRounds} scroll rounds`);
}
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 1000 });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
const result = await waitForFeedToStabilize(page);
console.log(`Captured ${result.count} items at document height ${result.height}px`);
// Wait for images that are already in the DOM to finish loading or fail.
await page.evaluate(async () => {
const images = [...document.images];
await Promise.all(images.map(img => {
if (img.complete) return Promise.resolve();
return new Promise(resolve => {
img.addEventListener('load', resolve, { once: true });
img.addEventListener('error', resolve, { once: true });
});
}));
});
await page.screenshot({ path: 'feed.png', fullPage: true });
} finally {
await browser.close();
}
The example’s pause, idle interval, and round count are tuning values, not universal settings. The image wait covers images already present in the DOM; it does not make a site load images that have not yet been requested. Since the loop scrolls to the bottom before each observation, lazy images that load on scroll may be requested as the feed is traversed.
3. Use an explicit end marker or known result count when available
An explicit terminal marker is usually a clearer signal than a stable-height heuristic. Start waiting before the event that causes the marker to appear, and set a timeout so a broken or changed page does not hang the job.
await page.waitForFunction(
selector => document.querySelector(selector) !== null,
{ timeout: 30000 },
'[data-testid="end-of-results"]',
);
await page.screenshot({ path: 'results.png', fullPage: true });
The selector is illustrative; use a marker that exists on the target site and whose meaning you have confirmed. For a known total, wait until enough result nodes are rendered:
const expectedCount = 120;
await page.waitForFunction(
({ selector, count }) =>
document.querySelectorAll(selector).length >= count,
{ timeout: 30000 },
{ selector: '[data-testid="feed-item"]', count: expectedCount },
);
await page.screenshot({ path: 'results.png', fullPage: true });
In a real infinite feed, the expected total may not be known or the site may virtualize rows by removing off-screen nodes. In that case, a rendered DOM count may never reach the total even though the page has loaded it. Use the app’s state or terminal marker when available.
4. Handle nested scroll containers and load-more buttons
If scrolling a panel triggers loading, scroll that element rather than window. Re-query the panel each round if the site can replace it during rendering.
const panelSelector = '[data-testid="results-panel"]';
const itemSelector = '[data-testid="result"]';
let previousCount = -1;
let stableRounds = 0;
for (let round = 0; round < 30 && stableRounds < 3; round++) {
await page.evaluate(selector => {
const panel = document.querySelector(selector);
if (!panel) throw new Error(`Missing scroll panel: ${selector}`);
panel.scrollTop = panel.scrollHeight;
}, panelSelector);
await new Promise(resolve => setTimeout(resolve, 500));
const count = await page.$$eval(itemSelector, nodes => nodes.length);
stableRounds = count === previousCount ? stableRounds + 1 : 0;
previousCount = count;
}
if (stableRounds < 3) throw new Error('Panel results did not stabilize');
const panel = await page.$(panelSelector);
await panel.screenshot({ path: 'panel.png' });
That last call captures the panel’s visible element area, not a stitched capture of every scrolled panel row. If the goal is the entire feed, the page must render the full content in the document or you must use a page-specific method to reveal it and then capture an appropriate full-page layout. A full-page screenshot does not automatically expand a panel with its own scrolling.
For a “Load more” button, click it repeatedly and stop on a real terminal condition. For example, wait for either the item count to increase or the button to become disabled; if neither happens before a timeout, report a load failure rather than assuming the feed is done.
const buttonSelector = '[data-testid="load-more"]';
const itemSelector = '[data-testid="feed-item"]';
for (let round = 0; round < 30; round++) {
const button = await page.$(buttonSelector);
if (!button) break; // Only treat absence as terminal if the site does.
const before = await page.$$eval(itemSelector, nodes => nodes.length);
const disabled = await button.evaluate(el =>
el.disabled || el.getAttribute('aria-disabled') === 'true'
);
if (disabled) break;
await button.click();
await page.waitForFunction(
({ selector, before }) => {
const button = document.querySelector(selector);
const count = document.querySelectorAll('[data-testid="feed-item"]').length;
return count > before || !button || button.disabled ||
button.getAttribute('aria-disabled') === 'true';
},
{ timeout: 15000 },
{ selector: buttonSelector, before },
);
}
await page.screenshot({ path: 'feed.png', fullPage: true });
Adapt the button’s loading and disabled checks to the site. If clicking navigates or replaces the page, wait for the relevant navigation or content change explicitly.
5. Capture only after content and visual assets are ready
- Full document: use
{ fullPage: true }after loading ends. It captures the current full document. - Viewport only: omit
fullPagewhen only the current viewport is required. - One element: use the element’s screenshot method when a single rendered region is the target.
- Lazy images: scroll through the content to trigger them, then wait for relevant image elements to complete.
- Fonts and layout: where the page uses web fonts or late layout updates, add a bounded, page-specific readiness check before capture.
- Very long pages: full-page capture may consume substantial memory and produce a large image. Consider capturing sections or using a PDF workflow if the deliverable permits it.
Puppeteer’s screenshot guide also documents element screenshots, which can be useful when the target is a feed panel or result card rather than the whole page. See the ScreenshotOptions API for capture settings supported by the Puppeteer version you use.
6. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot contains only the first batch | No scroll event triggered, wrong scroll target, or capture ran before the content condition | Inspect the page’s scroll container; trigger it and wait for an explicit marker or changing item count. |
networkidle times out |
Polling, analytics, streaming, or another persistent request | Use an app-specific condition; make network idle optional and bounded. |
| Network idle succeeds but results are missing | The quiet interval happened between batches | Continue scrolling and observe content state; network quiet is not terminal proof. |
| Height never stabilizes | Ads, animations, late layout shifts, or endless content keep changing dimensions | Use a terminal marker or count; track both height and item count only when appropriate; retain a strict loop and time limit. |
| Loop says stable too early | Pause is shorter than the page’s response delay, or the heuristic watches the wrong selector | Increase settling time based on observed behavior or switch to a site-specific signal. |
| Nested feed stays unchanged | The routine scrolls window, but the panel owns scrolling |
Set the panel’s scrollTop or use wheel input on the actual panel. |
| Images are blank | Lazy image requests were never triggered or have not completed | Scroll through the feed and wait for relevant image elements to load or fail. |
| Timeout waiting for item count | Selector is wrong, count is virtualized, target is unreachable, or load failed | Verify the selector in the page, inspect application state, and prefer a real terminal condition. |
| Script runs forever | No maximum rounds or outer timeout, or the page keeps adding content | Bound every loop and wait. Throw an error when the limit is reached and do not label that capture complete. |
7. Reliability, runtime, and cost tradeoffs
A completion condition is a reliability choice. A true end marker or known count is stronger than a heuristic. Stable height and count are practical when the application offers no signal, but can stop early during a delay or fail to stop on a page with continuous layout changes. Log the final item count, height, rounds, and whether the terminal condition was reached so downstream jobs can distinguish a completed capture from a partial one.
Runtime grows with the number of scroll cycles and the time spent waiting after each cycle. Reduce wasted waits by observing the specific response or DOM change the site makes, but keep a maximum duration for each batch and the overall job. Browser automation also consumes the resources needed to run Chromium and retain the rendered page; very large full-page images increase memory and output size. No universal benchmark or cost applies: it depends on the page, browser environment, capture dimensions, and workload.
For repeatable capture, keep selectors and expected terminal conditions configurable per site, record timeout failures separately from successful screenshots, and close the browser in a finally block. Avoid treating a swallowed wait timeout as successful completion.
Or skip the browser setup
If your job needs a clean capture and you do not want to maintain browser automation, ScreenshotNeo takes a screenshot with one GET request. It removes cookie and consent banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed; response headers report the page verdict and billing status. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.
For infinite-scroll pages, first confirm that the content you need is available to the capture request; a screenshot API does not remove the need for a page-specific completion rule when the target feed has not finished loading. For ordinary page captures, the one-call request is:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
See the ScreenshotNeo API documentation for request options. Sign up for 1,000 free screenshots a month with no card.
FAQ
Can I just use waitUntil: 'networkidle2'?
It can be useful during initial navigation, but it does not trigger additional scrolls or establish that an infinite feed is exhausted. Wait for the content condition that matters to your page.
Does fullPage: true scroll through the feed?
It captures the full document as currently rendered. It is not an infinite-scroll loader and will not guarantee future batches have been requested.
What if there is no end marker and the feed is genuinely endless?
Define a capture boundary, such as a target item count, maximum scroll depth, or time budget. An endless feed has no natural completion state, so your automation needs an explicit stopping policy.
Should I close the browser after a screenshot?
Yes, close it when the job is finished, including when navigation or capture fails. Use try/finally so failures do not leave browser processes running.


