ScreenshotNeo

BlogHow-to

How to Screenshot a Website with Infinite Scroll in Playwright

Load the content you need before capturing it. Learn how to scroll infinite lists, detect when loading is complete, and take reliable Playwright screenshots.

By the ScreenshotNeo team4 October 20269 min read

Short answer: trigger the infinite-scroll behavior first, wait until the content you need has loaded, then capture it with page.screenshot({ fullPage: true }). A full-page screenshot captures the page’s current scrollable extent; it does not guarantee that content which has not loaded yet will appear. If the list scrolls inside a nested container, scroll that container rather than the document.

Playwright’s scrolling guide describes manual scrolling as a way to force an infinite list to load more elements. The selector, loading signal, and stopping rule below must match the target site; there is no universal end-of-list detector.

1. Install Playwright and prepare a page

This runnable Node.js example uses Playwright’s test runner. It opens a page, scrolls a results container until its item count stays unchanged for several passes, and saves a full-page screenshot. Replace the URL, selectors, and stopping condition with values from the site you are capturing.

npm init -y
npm install --save-dev @playwright/test
npx playwright install chromium

Save this as infinite-scroll.spec.js:

const { test, expect } = require('@playwright/test');

test('capture a loaded infinite-scroll page', async ({ page }) => {
  await page.goto('https://example.com/feed', { waitUntil: 'domcontentloaded' });

  const results = page.getByTestId('results');
  const items = results.locator('.item');
  const targetCount = 100; // Set this to the number of items you need.
  const maxPasses = 80; // Safety bound; tune for the site's load rate.

  for (let pass = 0; pass < maxPasses; pass++) {
    const count = await items.count();
    if (count >= targetCount) break;

    const previousCount = count;
    await results.evaluate(element => {
      element.scrollTop = element.scrollHeight;
    });

    // Prefer a site-specific signal when available. This timeout is a fallback.
    await page.waitForTimeout(500);
    const nextCount = await items.count();

    if (nextCount === previousCount) {
      // A single unchanged pass may mean a request is still in flight.
      // The final check below decides whether the requested content arrived.
      await page.waitForTimeout(500);
      if (await items.count() === nextCount) break;
    }
  }

  const finalCount = await items.count();
  if (finalCount < targetCount) {
    throw new Error(`Expected at least ${targetCount} items, found ${finalCount}`);
  }

  await page.screenshot({ path: 'full.png', fullPage: true });
  expect(finalCount).toBeGreaterThanOrEqual(targetCount);
});

Run it with:

npx playwright test infinite-scroll.spec.js

The example uses a bounded count-based stopping rule so a broken or never-ending feed cannot make the script loop forever. For a site with a reliable end marker, loading indicator, or known number of results, use that site-specific condition instead. The short waits are illustrative fallback waits, not a guarantee that a request has completed.

2. Identify what actually scrolls

Before automating, inspect the page and determine whether the document or an inner element owns the scrollbar. A nested list may keep its own scrollTop; scrolling window in that case will not trigger more results.

Scroll the document

await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));

Scroll a nested container

const results = page.getByTestId('results');
await results.evaluate(element => {
  element.scrollTop = element.scrollHeight;
});

Programmatic scrolling is convenient when the application reacts to scroll position. Some pages respond only to real wheel input or require the container to be hovered first:

const results = page.getByTestId('results');
await results.hover();
await page.mouse.wheel(0, 800);

You can also bring a known bottom item or footer into view to trigger loading:

await page.getByText('Footer text').scrollIntoViewIfNeeded();

These approaches are documented in Playwright’s scrolling guide. Use locators that describe the page’s user-facing content or an explicit test contract. Playwright locators resolve the matching element again when used, which helps when a list’s DOM changes as new entries arrive; see the locator documentation.

3. Choose a completion condition

Do not treat a fixed delay as proof that the list is finished. Prefer a signal tied to the application’s behavior, then keep a maximum number of scroll passes as a safety bound.

Signal Useful when Implementation idea
Known item count You need a defined number of results Continue until locator.count() reaches the target
End marker The site renders an explicit end-of-results element Wait for that marker to become visible
Loading indicator The page shows a spinner or loading state Wait for loading to appear or disappear, as appropriate
Stable item count No terminal marker exists Stop after multiple scroll passes produce no new items, with a bounded loop
Network response The feed uses a known request Wait for the matching response, then assert that the DOM updated

For a feed with an explicit terminal marker, the loop can be structured around it:

const results = page.getByTestId('results');
const endMarker = page.getByTestId('end-of-results');
const maxPasses = 80;

for (let pass = 0; pass < maxPasses; pass++) {
  if (await endMarker.isVisible().catch(() => false)) break;

  const before = await results.locator('.item').count();
  await results.evaluate(element => {
    element.scrollTop = element.scrollHeight;
  });

  // Replace with the site's loading indicator or request signal if possible.
  await page.waitForTimeout(500);
  const after = await results.locator('.item').count();
  if (after === before && !await endMarker.isVisible().catch(() => false)) {
    throw new Error('No new items appeared and the end marker is not visible');
  }
}

Adapt the error handling if temporary network delays are expected. A stable count alone can mistake a slow request for the end of the list; waiting for a known request or loading state is more reliable when the site exposes one.

4. Capture the loaded content

After the desired content is present, choose a capture shape:

  • One tall image: use page.screenshot({ fullPage: true }) to capture the currently scrollable page extent.
  • Separate viewport images: capture at each scroll position when you need readable sections, manageable image dimensions, or a sequence for review.
  • One element: use a locator screenshot when only a component matters. For a scrollable container, an element screenshot shows the content currently visible in that container; it does not expand the container to include all its off-screen items.
// Full page after loading the desired entries
await page.screenshot({ path: 'full.png', fullPage: true });

// Current viewport only
await page.screenshot({ path: 'viewport.png' });

// The matched element's visible content
await page.getByTestId('results').screenshot({ path: 'results-visible.png' });

Playwright documents the fullPage and locator screenshot behavior in its screenshots guide. A very long page can produce a very large image; capturing viewport-sized sections is often more practical when the full result is unwieldy.

5. Python and cURL alternatives

The loading logic is browser automation, so Python needs Playwright too. cURL alone cannot execute page JavaScript or trigger a browser’s scroll events; it can only make HTTP requests. Use it to fetch a screenshot from a service that performs the browser capture, such as the ScreenshotNeo example below.

Python with Playwright

from playwright.sync_api import sync_playwright

URL = 'https://example.com/feed'
TARGET_COUNT = 100
MAX_PASSES = 80

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(URL, wait_until='domcontentloaded')

    results = page.get_by_test_id('results')
    items = results.locator('.item')

    for _ in range(MAX_PASSES):
        count = items.count()
        if count >= TARGET_COUNT:
            break

        results.evaluate('(element) => { element.scrollTop = element.scrollHeight; }')
        page.wait_for_timeout(500)  # Prefer a site-specific load signal.
        if items.count() == count:
            page.wait_for_timeout(500)
            if items.count() == count:
                break

    final_count = items.count()
    if final_count < TARGET_COUNT:
        raise RuntimeError(f'Expected {TARGET_COUNT} items, found {final_count}')

    page.screenshot(path='full.png', full_page=True)
    browser.close()

Install the package and browser with pip install playwright and playwright install chromium. The Python locator and screenshot APIs follow the same browser workflow; change selectors and waits for the target site.

6. Make screenshots more repeatable

Dynamic content can change between captures even when the code is unchanged. For useful visual comparisons:

  • Keep the Playwright version, browser version, operating system, viewport, device scale factor, and headless setting consistent.
  • Wait for the intended content and images to appear before capture; do not assume navigation completion means a lazy-loaded feed is complete.
  • Where animation causes unstable output, use Playwright’s screenshot animation controls or a stylesheet to hide or alter dynamic elements. See the screenshot API.
  • Capture the same number of entries or the same terminal state on each run.
  • Be aware that rendering can vary with host OS, browser version, settings, hardware, power source, and headless mode; Playwright’s visual comparisons guidance discusses keeping the environment consistent.

7. Troubleshooting

Symptom Likely cause Fix
The screenshot contains only the first few results The page was captured before scrolling triggered more loads Scroll in a loop, wait for a site-specific load signal, and verify the item count before capture
Scrolling has no effect The list uses a nested scroll container, or the wrong selector matched Inspect the page, target the actual container, and check that its scroll height exceeds its client height
The list loads in a real browser but not in automation The site may require wheel input, hover, or a specific interaction Hover the container and use page.mouse.wheel(), or scroll a bottom item into view
The loop exits while a request is still running A short fixed wait and one unchanged count were treated as completion Wait for the matching response or loading indicator; require multiple stable passes and retain a maximum pass count
The script loops indefinitely The feed has no detectable end or continues generating items Set a target count or explicit maximum passes and report when the stopping condition was not met
The full-page image is too tall or large The feed contains many entries or unbounded content Set a deliberate content limit and capture viewport sections instead of a single full-page image
Visual snapshots differ across runs Content, animations, fonts, browser, or host environment changed Stabilize the environment and page state; disable or mask dynamic elements where appropriate

8. Performance, reliability, and cost

Every extra scroll pass can trigger more application work and network requests. Choose the smallest content target that serves the task, use a meaningful stopping condition, and avoid repeatedly scrolling a feed whose terminal state is already known. A bounded loop limits wasted work and makes failures visible.

For reliability, check that the desired count or end marker was reached before writing the screenshot. If it was not, fail with a useful message instead of silently saving an incomplete capture. Keep the browser and host configuration stable for visual comparisons. The exact runtime and image size depend on the site’s behavior and the amount of content; the Playwright documentation provides no universal benchmark for infinite feeds.

With self-hosted Playwright, account for the compute and browser time your own environment uses. There is no per-shot Playwright API price in the cited documentation; infrastructure costs depend on your deployment. If you use a hosted screenshot API, check that provider’s billing rules and whether unsuccessful captures are charged.

Or skip the browser setup

If you need a screenshot of a page without writing and maintaining browser automation, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Example request for a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. Response headers identify the page verdict and billing status.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Does fullPage: true scroll the page to load everything?

No. It captures the full currently scrollable page extent. Trigger and verify the infinite-scroll loading first.

Can one screenshot include every item in an endless feed?

Only if you define a finite stopping point, such as a target count or terminal marker. A feed that keeps producing content has no natural finite full-page capture.

Should I scroll by a fixed pixel amount or jump to the bottom?

Either can work depending on the site’s event handling. If jumping to the bottom skips a trigger, use repeated wheel input or bring successive items into view.

Can I screenshot only the list?

Yes, with a locator screenshot, but it captures the element’s visible content. Load and capture individual portions if you need off-screen items from a scrollable list.

Sources