How to Fix Duplicate Content in a Screenshot of an Infinite-Scroll Page in Puppeteer
Stop repeated or missing content in Puppeteer screenshots by finding the scroll owner, defining a finite capture target, and waiting for page-specific loading signals.
Short answer: Do not treat fullPage: true as an instruction to finish loading an infinite feed. First find the element that actually scrolls, define how many items or which endpoint you need, scroll in bounded steps, and wait for evidence that the expected content has rendered. Then capture the page or the relevant element. If the site recycles old rows, capture viewport segments as you go instead of expecting one full-page image to include items that are no longer in the DOM.
The exact cause depends on the page, its scroll container, its rendering strategy, and your Puppeteer and browser versions. A duplicate image alone does not establish a universal Puppeteer defect. This guide gives you a diagnostic sequence and runnable patterns you can adapt to the page’s actual item selector and loading signal.
1. Identify what is being duplicated
Start with a fixed viewport and reproduce the problem in both a normal viewport screenshot and a full-page screenshot. Record the Puppeteer version, browser version, viewport dimensions, device scale factor, URL, and whether the issue occurs only after scrolling. This separates a page-loading problem from a full-page capture symptom.
Inspect the result for repeated cards, missing cards, blank regions, repeated page sections, or a viewport reset. Historical Puppeteer issues describe different symptoms across old versions, including repeated full-page images, missing content around lazy-loaded images, and screenshot-related reload or viewport behavior. They are case reports, not proof of one current, universal cause. Check your installed versions before attributing the behavior to Puppeteer.
| Observation | What to check |
|---|---|
| Repeated cards appear after a scroll | Whether the app appended the same records again, whether a request was retried, or whether the screenshot overlaps segments you assembled. |
| Earlier cards disappear as you scroll | Whether the feed virtualizes or recycles rows. A single final DOM snapshot may not contain the full history. |
| Only the full-page screenshot looks wrong | Compare with viewport captures and inspect the document height and page layout before and after capture. |
| Content is missing near the bottom | Check lazy loading, the actual scroll owner, and whether the expected items appeared before capture. |
2. Find the scroll owner and define a finite target
Infinite feeds have no natural final height. Pick a stopping rule before writing the loop: a known item ID, a target item count, a timestamp boundary, an end-of-results marker, or a maximum number of scroll steps. A fixed cap prevents a script from scrolling forever when the site never signals completion.
Determine whether the browser window scrolls or a nested panel owns the scrollbar. In DevTools, scroll the page and observe which element’s scrollTop changes. You can also inspect candidate containers from Puppeteer:
const candidates = await page.evaluate(() => {
return [...document.querySelectorAll('*')]
.map(el => ({
tag: el.tagName,
id: el.id,
className: typeof el.className === 'string' ? el.className : '',
clientHeight: el.clientHeight,
scrollHeight: el.scrollHeight,
overflowY: getComputedStyle(el).overflowY,
scrollTop: el.scrollTop,
}))
.filter(el => el.scrollHeight > el.clientHeight + 50 &&
['auto', 'scroll'].includes(el.overflowY))
.slice(0, 20);
});
console.table(candidates);
This is a diagnostic list, not an automatic guarantee that the largest candidate is the feed. Confirm which element responds when you scroll. If it is a nested panel, scrolling window will not trigger that panel’s loading behavior.
3. Scroll in measured steps and wait for the page’s signal
Prefer a condition tied to the feed, such as the item count increasing, a known item becoming visible, or a loading indicator changing state. Puppeteer supports condition-based waits through Page.waitForFunction() and element scrolling through Locators. A generic delay can be a fallback for a page with no observable signal, but it is not evidence that the feed finished loading.
The following example is for a window-scrolling page that appends elements matching .feed-item. Replace the selector and the stopping rule with the target site’s real structure. It gathers a bounded number of items, stops when the target count is reached, an end marker appears, or the scroll limit is reached, and verifies the result before taking a screenshot.
// save as capture-feed.mjs
// Install: npm install puppeteer
// Run: TARGET_URL="https://example.com/feed" node capture-feed.mjs
import puppeteer from 'puppeteer';
const url = process.env.TARGET_URL;
if (!url) throw new Error('Set TARGET_URL to the page to capture');
const ITEM_SELECTOR = '.feed-item'; // Replace with the real item selector.
const END_SELECTOR = '[data-end-of-results]'; // Optional; replace or remove.
const TARGET_COUNT = 100; // Choose the required finite scope.
const MAX_STEPS = 40; // Safety cap for an unbounded feed.
const STEP_FRACTION = 0.75; // Scroll less than one viewport.
const WAIT_MS = 10_000;
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector(ITEM_SELECTOR, { timeout: WAIT_MS });
let previousCount = await page.$$eval(ITEM_SELECTOR, els => els.length);
let stableRounds = 0;
for (let step = 0; step < MAX_STEPS && previousCount < TARGET_COUNT; step++) {
const ended = END_SELECTOR
? await page.$(END_SELECTOR).then(Boolean)
: false;
if (ended) break;
await page.evaluate(fraction => {
window.scrollBy(0, Math.floor(window.innerHeight * fraction));
}, STEP_FRACTION);
try {
await page.waitForFunction(
({ selector, before }) => document.querySelectorAll(selector).length > before,
{ timeout: WAIT_MS },
{ selector: ITEM_SELECTOR, before: previousCount },
);
} catch {
// No new item appeared before the deadline. Recheck the end marker and
// count; do not assume that a timeout means the feed is complete.
}
const count = await page.$$eval(ITEM_SELECTOR, els => els.length);
if (count > previousCount) {
previousCount = count;
stableRounds = 0;
} else {
stableRounds++;
if (stableRounds >= 2) break;
}
}
const summary = await page.evaluate(selector => {
const items = [...document.querySelectorAll(selector)];
return {
count: items.length,
firstText: items[0]?.innerText?.slice(0, 120) ?? null,
lastText: items.at(-1)?.innerText?.slice(0, 120) ?? null,
documentHeight: document.documentElement.scrollHeight,
scrollY: window.scrollY,
};
}, ITEM_SELECTOR);
console.log('Capture summary:', summary);
if (summary.count === 0) throw new Error('No feed items found at capture time');
await page.screenshot({ path: 'feed.png', fullPage: true });
} finally {
await browser.close();
}
The loop deliberately checks for new items after each scroll and has a hard limit. Its two unchanged rounds are a practical stopping fallback, not proof that no more results exist. If a reliable end marker or expected count exists, use that as the authoritative condition and report when the cap is reached before it.
4. Adapt the loop for a nested scrolling panel
For a panel that scrolls independently, target that element. The selector must identify the scrollable feed container, and the item count or readiness condition must still match the site.
const feedSelector = '.feed-scroll-panel';
const itemSelector = '.feed-item';
const before = await page.$$eval(itemSelector, els => els.length);
await page.locator(feedSelector).scroll({ scrollTop: 650 });
await page.waitForFunction(
({ selector, count }) => document.querySelectorAll(selector).length > count,
{ timeout: 10_000 },
{ selector: itemSelector, count: before },
);
const countAfter = await page.$$eval(itemSelector, els => els.length);
console.log({ before, countAfter });
Repeat this in a bounded loop using the panel’s actual dimensions and an explicit cap. Puppeteer’s Locator scroll action uses mouse-wheel behavior and checks element readiness; it does not know the application’s definition of “all feed results loaded.”
5. Decide between one full-page image and viewport segments
Puppeteer’s ScreenshotOptions describes fullPage: true as capturing the full page. That controls capture geometry; it does not define how far an infinite feed should load or ensure items removed from the DOM are preserved.
- Use one full-page capture when the required content is present in the document at once, the document height is bounded, and a single tall image is appropriate.
- Use viewport captures when the page virtualizes rows, the full-page image is too tall, or you need a sequence of page states. Capture each viewport after the expected content is present.
- Use an application export or data endpoint when the actual requirement is a complete data record rather than a visual representation. A screenshot cannot recover records that the page never rendered.
When stitching viewport images, track the scroll position and overlap intentionally. Do not concatenate screenshots by assuming every scroll step moved exactly one viewport: sticky headers, smooth scrolling, dynamic card heights, and late image loads can shift the seam. Save a manifest of each segment’s scroll position and observed item IDs so repeated or skipped sections can be diagnosed.
const segments = [];
for (let index = 0; index < 10; index++) {
const state = await page.evaluate(() => ({
y: window.scrollY,
ids: [...document.querySelectorAll('.feed-item')]
.map(el => el.getAttribute('data-id'))
.filter(Boolean),
}));
const path = `segment-${String(index).padStart(2, '0')}.png`;
await page.screenshot({ path, fullPage: false });
segments.push({ path, ...state });
const atBottom = await page.evaluate(() =>
window.scrollY + window.innerHeight >= document.documentElement.scrollHeight - 2,
);
if (atBottom) break;
await page.evaluate(() => window.scrollBy(0, Math.floor(window.innerHeight * 0.75)));
// Replace this delay with a wait for the page's item count or loading signal.
await new Promise(resolve => setTimeout(resolve, 500));
}
console.log(JSON.stringify(segments, null, 2));
The short delay in this segment example is only a placeholder where the site has no usable signal. For production, wait for a new item, a changed loading state, or a known item ID before capturing the next viewport.
6. Wait for the right readiness condition
page.goto(url, { waitUntil: 'networkidle2' }) can be useful during initial navigation, but an infinite feed may request content only after scrolling. Some pages also keep connections open. Network idle is therefore not a general completion test for scroll-triggered content. Prefer a site-specific condition.
Useful conditions include:
- The item count is at least the target count.
- A known item ID or end marker is present.
- A loading indicator disappears after the item count increases.
- The last visible item’s stable identifier changes after a scroll.
If the site exposes a loading marker, wait for its state as well as content. For example, wait for the marker to become visible after triggering a request and then hidden, while confirming that the item count or final item ID changed. Avoid waiting only for “spinner hidden” if the spinner was never shown.
7. Troubleshooting
| Symptom or error | Likely cause | Fix |
|---|---|---|
| The screenshot repeats a large section | The site may have duplicated data, the script may be capturing overlapping segments, or the full-page rendering may behave differently from viewport capture. | Compare viewport and full-page output; log item IDs and scroll positions; verify the DOM before capture; reproduce with fixed versions and viewport dimensions. |
| The item count never increases | Wrong item selector, wrong scroll owner, blocked request, or a feed that has reached its end. | Inspect the live DOM and network activity, test the actual scroll container, and check for an end marker or request error. |
TimeoutError from waitForSelector or waitForFunction |
The selector or condition does not match the page, content is gated, or loading exceeded the deadline. | Validate selectors in the browser, increase the timeout only when justified, and log the observed count and page state at timeout. |
| Only the latest rows appear | The app may virtualize or recycle old rows. | Capture viewport segments as rows appear, or obtain the records from an application export or supported data endpoint. |
| Images are blank or late | Image loading is independent of the text/item condition or uses lazy loading. | Wait for relevant image elements to report complete, or scroll them into view before capture. Verify image readiness rather than adding an arbitrary long sleep. |
| The page scrolls but the feed does not load | The window is moving while a nested panel owns the feed, or the page requires a different interaction. | Scroll the identified container using a Locator or mouse wheel and verify the expected signal changes. |
| Page height changes during capture | Late content, fonts, images, sticky elements, or layout shifts changed geometry. | Wait for the desired items and critical assets, capture after the layout settles, and prefer segments for highly dynamic pages. |
| Screenshot output differs across runs | Viewport, scale factor, browser version, animations, or asynchronous page data differ. | Fix viewport and device scale factor, record versions, and use a deterministic test page or page state where possible. |
8. Performance, reliability, and cost
Capture only the number of items needed. Loading an unbounded feed increases navigation time, memory use, and image size, and makes a timeout or layout change more likely before capture finishes. Use a maximum step count, item target, and overall deadline. Set viewport and device scale factor explicitly so output dimensions are predictable.
Full-page screenshots of very tall documents can consume substantial browser memory and produce large files. Viewport segments bound each image’s dimensions and make retries local: if one segment fails, recapture that segment after restoring its expected state. Segment capture takes more orchestration and can introduce seams, so record scroll positions and item identifiers.
For reliability, close the browser in a finally block, log the stopping reason, item count, first and last item identifiers, and final scroll position, and distinguish “target reached” from “stopped at safety cap.” Retry transient navigation or loading failures with a finite retry count; do not silently retry forever or capture a partial result as though complete.
Self-hosted Puppeteer has no per-screenshot API charge, but it uses compute and engineering time to run and maintain a browser. Resource use depends on the site, browser, viewport, loaded assets, and capture strategy; there is no universal timing or memory figure for this fix.
9. Or skip the browser setup
If the goal is a screenshot of the page rather than collecting every record in an infinite feed, ScreenshotNeo provides a one-request screenshot API. Its full-page option loads lazy images; it also supports selector-based element capture, waits, custom JavaScript and CSS, and other capture settings. Review the ScreenshotNeo API documentation for parameters and response behavior.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. These captures still represent what the page renders: if the feed virtualizes old items, use a bounded capture strategy or the application’s data export for a complete record.
Sign up for 1,000 free screenshots a month, with no card required.
10. FAQ
Does fullPage: true load an infinite feed to the end?
No. It requests a full-page capture of the page’s current rendered document. Your script must decide what content to load and when it is ready.
Is a fixed sleep ever acceptable?
It can serve as a fallback on a page with no observable readiness signal, but it is less reliable than waiting for a page-specific change. Keep a deadline and verify the resulting items.
Can one screenshot include content removed by virtualization?
Usually a screenshot can only show what the page has rendered for that capture. If old rows are removed or recycled, capture segments while they are present or use an application export.
What information is needed to diagnose a specific duplicate?
The page URL, Puppeteer and browser versions, viewport settings, scroll owner, relevant selectors, and examples of the expected and actual output narrow down the cause.


