Capture Infinite Scroll Search Results Without Duplicate Items in Puppeteer
Build a bounded Puppeteer loop that collects infinite-scroll results, deduplicates stable keys, waits for real progress, and handles virtualized lists.
To capture infinite-scroll search results in Puppeteer without duplicates, keep a persistent Set of stable result keys in Node.js. On each round, read the currently rendered cards, add only unseen records, scroll the element that actually owns scrolling, then wait for a page-specific progress signal. Stop when the site signals the end, a bounded number of rounds adds nothing new, or a maximum round/time limit is reached.
The selectors and progress condition must match the target site. Infinite-scroll pages often retain earlier cards, use an inner scrolling panel, or virtualize rows by removing them from the DOM. Inspect the page before choosing selectors or an end condition.
1. Inspect the result list and choose a stable key
Identify four things before writing the loop:
- Result card selector: a selector matching each search result, such as
.result-card. - Stable key: preferably a site-provided result ID. If none exists, use a normalized canonical URL and check whether distinct results can share that URL.
- Scroll target: the document viewport or the nested panel whose
scrollTopchanges. - Progress and end signals: a result count, cursor, last-result key, loading indicator, or explicit end marker.
Do not use a title alone as the key unless the site guarantees titles are unique. URLs may include tracking parameters; remove only parameters known to be irrelevant for that site. If results with the same URL can be distinct, combine the URL with another stable field or use the result ID.
2. Install Puppeteer
npm install puppeteer
The example below uses Puppeteer’s current Node.js API pattern. It assumes a page with result cards, a document-level scroll, and a result count that increases as more cards load. Replace the example URL, selectors, key logic, and progress signal to fit the site. Puppeteer recommends locators for selection and interaction; locators wait for presence and readiness when performing actions. For reading all currently matching cards, page.$$eval() is suitable. See the official Page interactions guide, Page.evaluate reference, Page API, and waitForFunction options.
3. Run a bounded collection loop
const puppeteer = require('puppeteer');
const fs = require('node:fs/promises');
const SEARCH_URL = 'https://example.com/search?q=puppeteer';
const CARD = '.result-card';
const END_MARKER = '[data-results-end="true"]';
const MAX_ROUNDS = 100;
const MAX_NO_PROGRESS = 3;
const MAX_RESULTS = 10_000;
const WAIT_TIMEOUT = 15_000;
function normalizeUrl(raw) {
if (!raw) return '';
try {
const url = new URL(raw);
url.hash = '';
// Remove only tracking parameters known to be irrelevant on this site.
for (const name of ['utm_source', 'utm_medium', 'utm_campaign']) {
url.searchParams.delete(name);
}
url.searchParams.sort();
return url.toString();
} catch {
return '';
}
}
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto(SEARCH_URL, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.waitForSelector(CARD, { timeout: 15_000 });
const seen = new Set();
const results = [];
let noProgressRounds = 0;
for (let round = 0; round < MAX_ROUNDS; round++) {
// Snapshot the current DOM into plain data. Do not keep ElementHandles
// across scrolls: cards can be replaced or virtualized.
const batch = await page.$$eval(CARD, cards => cards.map(card => {
const link = card.querySelector('a[href]');
return {
id: card.getAttribute('data-result-id') || '',
title: card.querySelector('.result-title')?.textContent?.trim() || '',
url: link?.href || '',
snippet: card.querySelector('.result-snippet')?.textContent?.trim() || '',
};
}));
let added = 0;
for (const item of batch) {
const normalizedUrl = normalizeUrl(item.url);
// Prefer an explicit ID. Fall back to URL; skip records with no key.
const key = item.id ? `id:${item.id}` : normalizedUrl ? `url:${normalizedUrl}` : '';
if (!key || seen.has(key)) continue;
seen.add(key);
results.push({ ...item, url: normalizedUrl });
added++;
if (results.length >= MAX_RESULTS) break;
}
if (results.length >= MAX_RESULTS || await page.$(END_MARKER)) break;
noProgressRounds = added ? 0 : noProgressRounds + 1;
if (noProgressRounds >= MAX_NO_PROGRESS) break;
// Capture a count before scrolling, then wait for it to increase.
const beforeCount = await page.$$eval(CARD, cards => cards.length);
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
try {
await page.waitForFunction(
({ selector, before }) => document.querySelectorAll(selector).length > before,
{ polling: 'mutation', timeout: WAIT_TIMEOUT },
{ selector: CARD, before: beforeCount },
);
} catch (error) {
if (error.name !== 'TimeoutError') throw error;
// A timeout is one bounded no-progress observation, not proof of end.
noProgressRounds++;
if (noProgressRounds >= MAX_NO_PROGRESS) break;
}
}
await fs.writeFile('results.json', JSON.stringify(results, null, 2));
console.log(`Saved ${results.length} unique results to results.json`);
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Save this as capture.js and run node capture.js. The script is runnable after replacing the example site’s selectors and URL. It stores each snapshot immediately in memory, so it can collect rows from virtualized lists even when earlier rows later disappear from the DOM. For very large collections, write batches incrementally to a file or database instead of retaining every record in memory.
4. Adapt progress, scrolling, and stopping to the site
Prefer a meaningful progress condition
waitForFunction() resolves when its predicate becomes truthy. Use the strongest observable signal the page provides:
- Count increased: simple when older cards remain in the DOM and new cards are appended.
- Cursor changed: useful when the page exposes a pagination cursor or next-page token in DOM state.
- Last visible key changed: useful when row count stays constant because the list is virtualized.
- Loading indicator cleared: helpful as a secondary condition, but a cleared spinner alone does not prove new results arrived.
- End marker appeared: stop immediately when the site exposes a reliable completion marker.
Puppeteer supports condition polling, including raf, mutation, and numeric polling options. A mutation wait is useful when the relevant DOM changes; it may not help if progress is represented only in JavaScript state or a fixed-height virtual list. Choose polling based on the signal.
Use the right scroll container
If scrolling the window does nothing, inspect the page for an element with its own overflow and changing scrollTop. For a results panel, use its selector and scroll it directly:
await page.$eval('.results-panel', panel => {
panel.scrollTop = panel.scrollHeight;
});
For a virtualized list, row count may remain constant. Capture the visible batch, scroll the panel by a portion of its client height, and wait until the last visible row key changes. Keep the persistent seen set outside the loop and add each batch before scrolling again.
Use locators for controls
If the page uses a “Load more” button instead of automatic scrolling, use a locator to click it and let Puppeteer wait for the control’s actionable state. Then wait on the site-specific progress signal before extracting the next batch. waitForSelector() is a lower-level presence wait; it does not automatically retry a failed action.
Bound the work
Use both a maximum round count and a no-progress limit. A single slow request should not falsely end a capture, but a page that never advances must not loop forever. Treat the no-progress threshold as a site-specific policy, and log why the loop stopped: explicit end, result cap, repeated no progress, or maximum rounds.
5. Use a stable deduplication key
Deduplicate in Node.js, outside the page evaluation. A persistent set covers overlap between successive snapshots, and plain objects are easier to serialize than browser element handles.
| Key choice | When it fits | Risks |
|---|---|---|
| Result ID | The site exposes an ID stable across renders. | IDs may only be unique within a query or session; scope the key if needed. |
| Canonical URL | Each result points to a distinct destination. | Tracking parameters, fragments, redirects, or duplicate destination URLs can affect identity. |
| Composite key | URLs repeat but another field distinguishes results. | Mutable titles or snippets can make a composite key unstable. |
Normalize consistently before inserting and comparing. Do not remove all query parameters blindly: some sites use them to identify different content. If a record has no trustworthy key, either skip it with a diagnostic or use a documented fallback and accept the risk of collisions. The correct choice depends on the target site; Puppeteer does not define result identity.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The same results appear repeatedly. | The page retains earlier cards between loads, or the key is not normalized consistently. | Keep one persistent set across the full run; prefer a stable ID and normalize fallback URLs the same way every time. |
| Scrolling has no effect. | A nested panel owns scrolling, or the page has not reached the trigger threshold. | Inspect scrollable elements and scroll the actual results container. Confirm its scrollTop changes. |
| The wait times out, but more results eventually appear. | The timeout is shorter than the site’s load time, or the predicate watches the wrong signal. | Increase the bounded wait window and observe a count, cursor, or visible-key change. Keep a maximum duration and retry limit. |
| The count never increases. | The page virtualizes rows, replacing old rows instead of appending; or count measures a different element. | Wait for a last-row key or cursor to change, and capture every visible batch before scrolling again. |
| Some results are missing from the final DOM. | Virtualization removed earlier rows after scrolling. | Persist each extracted batch immediately. Do not scrape only once at the end. |
| The run stops too early. | A temporary slow response was mistaken for completion, or the no-progress threshold is too low. | Check for an explicit end marker and loading state; allow several bounded no-progress observations suited to the site. |
| The run never finishes. | There is no end marker and each scroll triggers content or layout mutations without new results. | Base progress on unique keys or a cursor, not arbitrary DOM mutations. Enforce round, time, and result limits. |
waitForSelector() succeeds but clicking still fails. |
Presence does not guarantee visibility, stability, or actionability. | Use a locator for interaction or explicitly wait for the condition required by the control. |
| Duplicate pages survive URL cleanup. | The site redirects aliases to one destination, or URL parameters distinguish content. | Inspect canonical links and result IDs. Adjust normalization only for known irrelevant parameters. |
7. Performance, reliability, and cost
Each full-page DOM scan costs work proportional to the number of currently rendered cards. On pages that retain every result, rescanning all cards on every round can repeat work; if that becomes material, extract only newly appended nodes where the site’s DOM structure makes that safe, while retaining deduplication as a correctness check. Virtualized lists usually keep each scan bounded, but require immediate persistence of every batch.
Wait on progress rather than using a long fixed sleep after each scroll. This can return promptly on fast responses while preserving a timeout for slow ones. Use a single browser session for a capture, close it in a finally block, and save checkpoints for long runs so a browser failure does not discard all collected data. Retries should be bounded; repeated navigation or scrolling without observing state can create duplicate work.
Cost depends on where Puppeteer runs and the target site’s access rules, browser runtime, network use, and storage. This guide makes no benchmark or universal runtime claim. For public sites, follow their access requirements and avoid sending a request rate that harms the service. If the task is only to capture a page image or PDF rather than extract every search result into structured data, a screenshot API can remove browser setup from the job.
8. Or skip the browser setup
For a visual capture of a search page, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. It does not extract each result into structured records or replace the deduplication loop above. See the ScreenshotNeo API documentation for its request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/search?q=puppeteer -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/search?q=puppeteer"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/search?q=puppeteer',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
- Cookie banners, popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, and failed loads are never billed; response headers report the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
9. FAQ
Can I use the loop to scrape any search engine?
The collection pattern is general, but selectors, scroll behavior, authentication, and access rules are site-specific. Inspect the target page and follow its terms and access requirements.
Should the Set live in the browser or in Node.js?
Keep it in Node.js for this pattern. It survives DOM replacement and makes it straightforward to write checkpoints or persist results.
What if the same result appears under different URLs?
Use a stable result identifier if available. Otherwise, define site-specific canonicalization and validate that it does not merge distinct results.
Does a screenshot capture replace extracting results?
No. A screenshot or PDF preserves visual appearance; structured extraction requires reading result fields and applying a deduplication key.


