How to Make Puppeteer Wait for an Infinite Scroll Feed to Stabilize
Wait for feed items or an end marker, then stop after a bounded quiet period. Learn why network idle alone cannot tell you an infinite feed is finished.
To make Puppeteer wait for an infinite-scroll feed to stabilize, define a page-specific signal—such as the number of feed items, a loading indicator, or an end marker—and wait for that signal after each scroll. Continue while content appears, and stop when the page signals completion or your bounded no-growth policy says to stop. There is no universal Puppeteer event that means every infinite feed is finished.
networkidle can be a useful secondary signal after a request, but it only describes network activity during a window. It cannot guarantee that a later scroll, observer, timer, or user interaction will not load more items.
1. Identify what “stable” means for this feed
Before writing the wait, inspect the target page and identify:
- The feed item selector, such as
.feed-item. - The scrolling element: often the document, but sometimes an inner container.
- Any loading indicator, such as
.feed-loading. - An explicit end marker, if the application exposes one, such as
.feed-end.
Prefer a signal tied to application content. A count increase proves that new items appeared; an explicit end marker can indicate completion. A quiet period with no count change is only a policy choice: it means no change was observed during that period, not that the site has promised there will never be more content.
2. Runnable Puppeteer example for a document-scrolling feed
Install Puppeteer with npm install puppeteer. Save this as feed.js and run it with node feed.js https://example.com/feed. Replace the example URL and selectors with the target site’s values. The script scrolls, waits for item growth or an end marker, and stops after an explicit end marker or a bounded number of rounds without growth.
const puppeteer = require('puppeteer');
const url = process.argv[2];
if (!url) {
throw new Error('Usage: node feed.js https://example.com/feed');
}
const ITEM_SELECTOR = '.feed-item';
const END_SELECTOR = '.feed-end';
const MAX_ROUNDS = 30;
const WAIT_FOR_CHANGE_MS = 10_000;
const QUIET_MS = 1_000;
const MAX_NO_GROWTH_ROUNDS = 3;
async function main() {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
page.setDefaultTimeout(30_000);
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
let count = await page.$$eval(ITEM_SELECTOR, nodes => nodes.length);
let noGrowthRounds = 0;
let ended = false;
for (let round = 1; round <= MAX_ROUNDS; round++) {
if (await page.$(END_SELECTOR)) {
ended = true;
break;
}
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
try {
await page.waitForFunction(
({ itemSelector, endSelector, previousCount }) =>
document.querySelectorAll(itemSelector).length > previousCount ||
Boolean(document.querySelector(endSelector)),
{ timeout: WAIT_FOR_CHANGE_MS, polling: 'mutation' },
{ itemSelector: ITEM_SELECTOR, endSelector: END_SELECTOR, previousCount: count },
);
} catch (error) {
if (error.name !== 'TimeoutError') throw error;
// A timeout means no qualifying DOM change was observed in this round.
}
// Allow short-lived rendering and layout updates to settle, then recount.
await new Promise(resolve => setTimeout(resolve, QUIET_MS));
const current = await page.$$eval(ITEM_SELECTOR, nodes => nodes.length);
ended = Boolean(await page.$(END_SELECTOR));
if (ended) break;
if (current > count) {
count = current;
noGrowthRounds = 0;
} else {
noGrowthRounds++;
if (noGrowthRounds >= MAX_NO_GROWTH_ROUNDS) break;
}
}
console.log(JSON.stringify({ itemCount: count, endMarkerFound: ended }, null, 2));
} finally {
await browser.close();
}
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
This is a bounded pattern, not a site-independent drop-in: selectors, the scroll target, wait window, quiet interval, and no-growth policy must match the site and the task. If the feed has no end marker, the result should be understood as “no new items appeared for the configured number of rounds,” not “the feed is definitely complete.”
Use an explicit loading state when available
If the site exposes a loader that appears during requests, wait for it to appear and then disappear, or wait for it to disappear after triggering the scroll. A loader transition can be more informative than item count when the page replaces items, updates existing cards, or inserts content in batches. Still impose a timeout: a loader can fail to appear or remain visible indefinitely.
For an inner scrolling container
Scroll the actual feed container instead of the window. For example, replace the scroll evaluation with the following, using the page’s real container selector:
await page.evaluate(() => {
const feed = document.querySelector('.feed-scroll-container');
if (!feed) throw new Error('Feed scroll container not found');
feed.scrollTop = feed.scrollHeight;
});
The item and end-marker predicates can still use document selectors if the feed elements are in the document. If the site renders the feed inside an iframe or shadow root, locate and query that frame or shadow root instead.
3. Choosing the wait strategy
| Signal | What it observes | Use it for | Limitation |
|---|---|---|---|
| Item count or content predicate | Application DOM state | Detecting newly loaded feed entries | Choose a predicate that reflects the page’s rendering behavior. |
| End marker | Application completion state | Stopping when the page exposes a credible end signal | Some feeds have no end marker, or only render it after scrolling further. |
| Loading indicator | Application loading state | Waiting for a known request/render cycle | May not appear for every batch, or may get stuck. |
waitForNetworkIdle() / networkidle |
Network connections during an inactivity window | A secondary pause after network-backed work | Does not prove future scrolling cannot trigger more content; persistent connections can prevent idleness. |
| Fixed delay | Elapsed time only | A short rendering grace period after a meaningful event | Can waste time or finish before slow content arrives. |
Puppeteer’s Page.waitForFunction() API evaluates a predicate in the page context until it returns a truthy value and accepts arguments for that predicate. Its documented polling choices include 'raf', 'mutation', and a numeric interval; the documented default timeout is 30 seconds. The waitForNetworkIdle() API documents a default idle time of 500 ms and a default concurrency threshold of zero. These are API defaults, not recommended feed-specific thresholds.
Puppeteer’s navigation lifecycle names networkidle0 and networkidle2 refer to network connection limits held for a 500 ms period. They describe transport activity, not feed completion. See the lifecycle event documentation and the page interactions guide for waiting and scrolling mechanics.
4. Tune the policy without hiding uncertainty
- Wait window: Set it long enough for the site’s normal response and rendering time, while keeping it finite. A timeout is an expected outcome for a feed that did not change in that round.
- Quiet interval: Use it to let DOM updates and rendering settle after a change. It is not a universal definition of stability.
- No-growth limit: Choose how many consecutive scroll rounds without growth end the job. A low limit is faster but can stop on a temporarily slow batch; a high limit can spend more time on feeds with no true end.
- Maximum rounds: Always cap total scrolls to protect against endlessly changing feeds, bugs, or adversarial pages.
- Content condition: If the page updates items in place, count alone may not detect meaningful change. Wait for a specific field, timestamp, loading state, or other application signal instead.
For a styling change that does not mutate the DOM, use polling: 'raf' or a numeric polling interval rather than relying on mutation polling. For item insertion/removal, mutation polling is often a natural fit. The exact behavior depends on the predicate and page.
5. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The wait times out although more feed entries eventually appear. | The page’s response or render time exceeds the wait window, or the predicate watches the wrong selector. | Verify the selector in the live DOM, inspect the actual loader/state transition, and adjust the finite timeout to the page’s behavior. |
| The script stops after one scroll. | The feed scrolls inside a container, or the document was already at its bottom. | Identify the scrollable element and set its scrollTop to scrollHeight. |
| The script loops until the maximum round count. | No end marker exists, the marker selector is wrong, or content continues to load. | Validate the marker and selectors. If there is no reliable end signal, report the bounded no-growth policy as the stopping rule. |
networkidle never resolves. |
The site keeps connections active, for example through streaming or polling. | Use a page-specific DOM or loading-state predicate as the primary wait; use network idle only if the page’s traffic pattern supports it. |
| The item count stays the same while the feed changes. | The app replaces or updates existing nodes rather than appending new ones. | Observe a meaningful item attribute or content value, or track a site-specific loading/completion state. |
waitForFunction fails immediately with a selector or evaluation error. |
The predicate runs in the page context, where Node-only variables and APIs are unavailable; arguments may also be passed in the wrong order. | Use browser-context APIs such as document.querySelectorAll, pass values as predicate arguments, and check the current Puppeteer API signature. |
| The page loads but has no expected items. | Navigation completed before client rendering, the page requires authentication, or a bot check blocked access. | Wait for a page-specific initial state, provide authorized session credentials where appropriate, and inspect the resulting page before starting the scroll loop. |
6. Performance, reliability, and cost
Each scroll round and wait adds latency. A predicate tied to the desired content usually avoids the waste of repeated long fixed sleeps, while bounded timeouts prevent a failed feed from hanging a job indefinitely. The browser still has to download and render the page’s content, and long feeds can consume substantial memory; stop when the requested content or policy limit is reached.
For reliability, log at least the round number, item count before and after scrolling, whether the end marker appeared, and whether the wait timed out. Keep page-specific selectors and thresholds configurable. Treat navigation, browser closure, cancellation, and unexpected evaluation errors as failures to report rather than as evidence that the feed ended.
With self-hosted Puppeteer, cost depends on the compute and browser runtime you operate; the research sources do not establish a general price or benchmark. If you need repeated production captures, consider the operational cost of browser setup, concurrency, retries, and page-specific maintenance alongside the API price.
7. Or skip the browser setup
If your task is to capture a page rather than control a custom Puppeteer feed loop, ScreenshotNeo provides a website screenshot API and MCP server. See the API documentation for parameters and response details. For example, this cURL request captures a page as WebP:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. This is a screenshot service, so it does not replace custom Puppeteer logic when you need to inspect feed growth or decide that an infinite feed has stabilized.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
8. FAQ
Can Puppeteer detect that an infinite feed is truly complete?
Only if the page exposes a credible completion signal, such as an end marker or a documented state. Otherwise, your script can apply a bounded stopping policy, but it cannot prove that more items will never appear.
Should I use waitForTimeout() after every scroll?
A short delay can allow rendering to settle, but use a content or application-state wait as the main signal. A delay alone does not indicate that the expected work finished.
Does locator stability mean the feed is stable?
No. An interaction’s stable bounding box check concerns the target element’s layout across animation frames. It does not mean the feed will not append more entries.
What should the script return if no end marker exists?
Return the collected items and the stopping reason, such as “three consecutive rounds without growth” or “maximum rounds reached.” That makes the limit explicit to downstream callers.


