ScreenshotNeo

BlogHow-to

How to Scroll a Website While Crawling with Node.js

Use Playwright or Puppeteer to load infinite-scroll pages, detect progress, stop safely, deduplicate records, and handle failures.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: a JavaScript-rendered infinite-scroll page usually needs a real browser. With Node.js, launch Playwright or Puppeteer, navigate to the page, scroll a bottom sentinel or the correct scroll container, wait for a measurable progress signal, extract each batch, deduplicate records, and stop after a bounded number of rounds or repeated rounds without progress. A plain HTTP request may return only the initial HTML.

Choose the right approach

Situation Best approach
Items appear after scrolling Playwright or Puppeteer
The page calls a JSON endpoint for more items Find and call that endpoint directly when permitted
You need a screenshot after content loads Run the browser workflow, then capture the rendered page
The list scrolls inside a panel Scroll that element’s scrollTop, not the window

Playwright: complete bounded crawler

Install Playwright and its browser, then save this as crawl.mjs.

npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
  viewport: { width: 1440, height: 900 },
  userAgent: 'ExampleCrawler/1.0 (contact: you@example.com)'
});

const url = 'https://example.com/list';
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator('.item').first().waitFor({ state: 'attached', timeout: 15_000 }).catch(() => {});

const seen = new Set();
const rows = [];
const maxRounds = 40;
const maxMinutes = 3;
const started = Date.now();
let stagnantRounds = 0;
let terminationReason = 'maximum rounds reached';

for (let round = 0; round < maxRounds; round++) {
  if (Date.now() - started > maxMinutes * 60_000) {
    terminationReason = 'time limit reached';
    break;
  }

  const beforeCount = await page.locator('.item').count();
  const beforeHeight = await page.evaluate(() => document.documentElement.scrollHeight);
  const sentinel = page.locator('.list-end, footer').last();

  if (await sentinel.count()) {
    await sentinel.scrollIntoViewIfNeeded();
  } else {
    await page.mouse.wheel(0, 1200);
  }

  await page.waitForTimeout(500);
  const spinner = page.locator('.loading, [aria-busy="true"]').first();
  if (await spinner.count()) {
    await spinner.waitFor({ state: 'hidden', timeout: 10_000 }).catch(() => {});
  }

  const afterCount = await page.locator('.item').count();
  const afterHeight = await page.evaluate(() => document.documentElement.scrollHeight);
  const batch = await page.locator('.item').evaluateAll(nodes => nodes.map(node => ({
    id: node.getAttribute('data-id') || node.querySelector('a')?.href || null,
    text: node.textContent?.trim() || ''
  })));

  for (const row of batch) {
    if (row.id && !seen.has(row.id)) {
      seen.add(row.id);
      rows.push(row);
    }
  }

  const progressed = afterCount > beforeCount || afterHeight > beforeHeight;
  stagnantRounds = progressed ? 0 : stagnantRounds + 1;

  const endMarker = page.locator('.end-of-results, [data-end="true"]').first();
  if (await endMarker.count() && await endMarker.isVisible().catch(() => false)) {
    terminationReason = 'end marker visible';
    break;
  }
  if (stagnantRounds >= 3) {
    terminationReason = 'no measurable progress for three rounds';
    break;
  }
}

await browser.close();
console.log(JSON.stringify({ rows, count: rows.length, terminationReason }));

The loop uses three independent safeguards: a maximum round count, a wall-clock deadline, and repeated rounds with no increase in item count or document height. Keep all three because a site can report an unchanged item count while still changing height, or change height because of layout shifts without loading records.

Scrolling a nested container

Many dashboards keep the page fixed and scroll a nested div. Scrolling window then has no effect. Set the container’s scroll position and measure its own height.

const list = page.locator('.results-panel');
await list.waitFor();

for (let round = 0; round < 40; round++) {
  const before = await list.locator('.item').count();
  await list.evaluate(el => { el.scrollTop = el.scrollHeight; });
  await page.waitForTimeout(500);
  const after = await list.locator('.item').count();
  if (after === before) break;
}

Use scrollIntoViewIfNeeded() for a bottom sentinel when the site exposes one. Use page.mouse.wheel() when incremental wheel events are required. These are documented Playwright scrolling patterns; locator actions also wait for actionable elements and retry transient failures.

Wait for progress instead of sleeping blindly

A fixed delay alone is unreliable. Tie the wait to one or more signals:

  • the number of item locators increases;
  • a loading spinner becomes hidden;
  • a “load more” control disappears or becomes disabled;
  • document or container height changes;
  • a known network response arrives;
  • an end-of-list marker appears.
const responsePromise = page.waitForResponse(
  response => response.url().includes('/api/items') && response.ok(),
  { timeout: 10_000 }
).catch(() => null);
await page.mouse.wheel(0, 1400);
const response = await responsePromise;
if (response) {
  console.log('Loaded batch from', response.url());
}

Use bounded waits and catch timeout errors when a site legitimately loads from cache or does not make a request on every scroll.

Extract rendered data safely

Extract after the page settles, and deduplicate with a stable record ID or canonical URL. Virtualized lists recycle DOM nodes, so the same elements can represent different records over time.

const batch = await page.locator('.item').evaluateAll(nodes => nodes.map(node => {
  const link = node.querySelector('a');
  return {
    id: node.getAttribute('data-id') || link?.href || null,
    title: node.querySelector('.title')?.textContent?.trim() || '',
    href: link?.href || null
  };
}));

for (const item of batch) {
  if (!item.id || seen.has(item.id)) continue;
  seen.add(item.id);
  rows.push(item);
}

Persist raw HTML or structured records after each successful round when the crawl matters. Log the round number, item count, last height, and termination reason so a later run can explain what happened.

Puppeteer variant

Puppeteer locators can scroll targets into view before interaction. The following crawler stops when item counts stop changing twice in a row, then saves the rendered HTML.

npm install puppeteer
import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto('https://example.com/list', {
  waitUntil: 'domcontentloaded',
  timeout: 45_000
});

let previousCount = -1;
let stagnant = 0;
const seen = new Set();
const rows = [];

for (let round = 0; round < 40 && stagnant < 3; round++) {
  const count = await page.locator('.item').count();
  const end = page.locator('.list-end, footer').last();
  if (await end.count()) {
    await end.scroll({ scrollTop: 1000 });
  } else {
    await page.mouse.wheel({ deltaY: 1200 });
  }
  await new Promise(resolve => setTimeout(resolve, 500));

  const current = await page.locator('.item').count();
  stagnant = current === count && current === previousCount ? stagnant + 1 : 0;
  previousCount = current;

  const batch = await page.locator('.item').evaluateAll(nodes => nodes.map(node => ({
    id: node.getAttribute('data-id') || node.querySelector('a')?.href || null,
    text: node.textContent?.trim() || ''
  })));
  for (const item of batch) {
    if (item.id && !seen.has(item.id)) {
      seen.add(item.id);
      rows.push(item);
    }
  }
}

const html = await page.content();
await browser.close();
console.log(JSON.stringify({ rows, htmlLength: html.length }));

Puppeteer’s page.content() returns the current rendered HTML. Prefer locators for extraction and interaction so visibility and retry behavior are handled consistently.

Find a direct data endpoint

Before automating dozens of scroll events, inspect requests made after one scroll. If the page calls a JSON endpoint containing the records, calling that endpoint can be faster and less fragile than rendering every batch. Confirm authentication, rate limits, terms, and permission to collect the data.

page.on('response', async response => {
  const type = response.headers()['content-type'] || '';
  if (type.includes('application/json') && response.url().includes('/api/')) {
    console.log(response.status(), response.url());
  }
});
await page.mouse.wheel(0, 1200);

Use the browser to establish cookies or tokens when required, then use the documented endpoint with the same authorization. Do not assume a private endpoint is stable or permitted simply because it is visible in developer tools.

Stopping rules that work

Rule What it detects Typical use
Maximum rounds Runaway pages and unexpected loops Always
Time limit Slow or stalled servers Production crawlers
Repeated no-progress rounds End of list or failed loading Most infinite lists
End marker Explicit terminal state Sites that expose one
Stable height and count No new DOM content Simple feeds
Disabled “load more” Finite pagination Button-driven lists

Use a small tolerance for layout shifts. A height change alone is not proof of new records, and a count change can be caused by advertisements or placeholders. Combine signals and record why the loop stopped.

Retries, timeouts, and reliability

  • Set navigation and selector timeouts explicitly.
  • Retry navigation or a failed batch with exponential backoff, but cap attempts.
  • After a retry, re-check the last stable ID to avoid duplicate output.
  • Save partial results after every round.
  • Use a realistic user agent and identify your crawler where appropriate.
  • Close the browser in a finally block in long-running services.
async function gotoWithRetry(page, url, attempts = 3) {
  let lastError;
  for (let attempt = 0; attempt < attempts; attempt++) {
    try {
      return await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
    } catch (error) {
      lastError = error;
      await new Promise(resolve => setTimeout(resolve, 2 ** attempt * 1000));
    }
  }
  throw lastError;
}

Performance and cost considerations

  • Reuse one browser process and create separate pages for related URLs.
  • Block images, fonts, analytics, and advertisements only when they are not needed for the content or layout you are extracting.
  • Prefer a direct JSON endpoint when it is documented and permitted.
  • Use a sentinel or container jump instead of hundreds of tiny wheel events.
  • Keep concurrency below the site’s limits and add delays between pages.
  • Bound every crawl so a broken “load more” implementation cannot consume unbounded CPU, memory, or bandwidth.

Browser crawling costs CPU and memory for each active page. Direct endpoint requests usually use fewer resources, while browser rendering is necessary when JavaScript, cookies, interaction, or layout determines what appears.

Troubleshooting

Symptom Cause Fix
Item count never changes You scrolled the window but the list uses a nested container Set that container’s scrollTop or scroll its sentinel
Only the first batch is extracted Extraction runs before loading settles Wait for count, spinner, response, or height progress
Loop never ends No terminal marker or progress signal is reliable Add maximum rounds, a deadline, and a stagnant-round limit
Duplicate records Virtualized DOM nodes are recycled Deduplicate by stable ID or canonical URL
Timeout on a response wait Content came from cache or an alternative request Catch the timeout and use a count or spinner signal
Headless browser sees a challenge Bot detection or authentication is blocking the page Respect access rules; authenticate only with permission and do not attempt to bypass a challenge
Blank content in CI Browser dependency, sandbox, or missing wait condition Install the browser, inspect logs, and wait for a real selector

Compliance checklist

  • Read robots.txt before crawling.
  • Check the site’s terms, authentication requirements, and rate limits.
  • Collect only data you are permitted to collect.
  • Protect personal data and respect copyright obligations.
  • Throttle requests and identify your crawler when appropriate.

robots.txt communicates which URLs crawlers may access and helps manage traffic; it is not a security control. A disallowed URL may still be indexed if linked elsewhere.

Or skip the browser setup

If your goal is a clean visual capture after the page is ready, ScreenshotNeo provides a one-request screenshot API. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/list -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/list"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/list' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response reports the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can I crawl an infinite-scroll page with fetch?

Only when the records are present in the initial HTML or exposed through a callable endpoint. Otherwise use browser automation or identify the page’s data request.

Should I scroll by pixels or scroll a sentinel?

Prefer a bottom sentinel when available. Pixel scrolling is useful when the page loads content from wheel events or has no stable marker.

How many stagnant rounds should I allow?

Three is a practical starting point, but tune it to the site’s latency and loading behavior while keeping a hard maximum and deadline.

Why are stable IDs important?

Infinite lists often virtualize rows and reuse the same DOM nodes. A stable ID or canonical URL lets you retain each record exactly once.