Fix Cropped Website Screenshots on Pages with Infinite Scroll
Full-page capture can miss content that appears only after scrolling. Load the sections you need, verify them, then capture the page with Playwright.
A full-page screenshot captures the page’s scrollable height, but that does not guarantee that an infinite-scroll page has loaded everything below the fold. Scroll through the page in controlled steps, wait for the next expected content to appear, and repeat until the portion you need is present. Then capture the full page and inspect the image from top to bottom.
For a static page, Playwright’s full-page option is often enough. For a page that adds content in response to scrolling, trigger that behavior first. The exact trigger and wait condition depend on the site; there is no universal delay that works for every page.
1. Check what kind of crop you have
First distinguish a viewport screenshot from a full-page screenshot. A viewport screenshot contains only what is visible at the current scroll position. A full-page screenshot asks the browser to capture the scrollable page as one tall image. Playwright documents this with fullPage: true in JavaScript and full_page=True in Python. Playwright screenshot documentation
If the image shows the top viewport and ends there, enable full-page capture. If it is full-page but sections are missing, the likely problem is that the site has not loaded or rendered those sections yet. A full-page capture request does not, by itself, prove that scroll-triggered application code has run. A Playwright issue report describes lazy or scroll-triggered content being missed when full-page capture does not move the visual viewport; treat this as a reported limitation, not a guarantee about every browser or page. Playwright issue report
2. Load the content before capturing
- Open the page and identify the section or final item that must appear in the screenshot.
- Scroll down by a controlled amount, such as a fraction of the viewport, rather than jumping straight to the bottom.
- At each loading boundary, wait until the expected new item or section is visible.
- Repeat until the needed content has appeared. Stop when the target section is present, or when the page reaches a clear end.
- Capture the full page and inspect the output, including the transitions between loaded batches and the final section.
Infinite-scroll pages commonly fetch or render more items when scrolling approaches a boundary. The steps above are practical guidance based on that behavior; the precise scroll distance, trigger, and readiness condition are site-specific. A fixed sleep can be useful as a short settling pause, but it is not evidence that the intended content loaded.
3. Automate the workflow with Playwright
Install Playwright and its Chromium browser in a project that supports Node.js:
npm install playwright
npx playwright install chromium
Save this as capture.mjs. Replace the target URL, the item selector, and the readiness condition with selectors that match the site. The loop scrolls in increments, waits for the page’s item count to grow, and stops when the target count appears or no new items arrive within the per-step timeout.
import { chromium } from 'playwright';
const url = 'https://example.com/feed';
const itemSelector = 'article';
const targetCount = 80;
const maxSteps = 40;
const growthTimeoutMs = 8000;
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.locator(itemSelector).first().waitFor({ state: 'visible', timeout: 15000 });
let previousCount = 0;
let stagnantSteps = 0;
for (let step = 0; step < maxSteps; step++) {
const countBefore = await page.locator(itemSelector).count();
if (countBefore >= targetCount) break;
await page.evaluate(() => window.scrollBy(0, Math.floor(window.innerHeight * 0.75)));
try {
await page.waitForFunction(
({ selector, count }) => document.querySelectorAll(selector).length > count,
{ selector: itemSelector, count: countBefore },
{ timeout: growthTimeoutMs }
);
} catch {
// No additional matching items appeared during this step.
}
const countAfter = await page.locator(itemSelector).count();
if (countAfter <= previousCount) stagnantSteps++;
else stagnantSteps = 0;
previousCount = countAfter;
if (countAfter >= targetCount || stagnantSteps >= 2) break;
}
// Return to the top before asking for a single full-page image.
await page.evaluate(() => window.scrollTo(0, 0));
await page.screenshot({ path: 'screenshot.png', fullPage: true });
} finally {
await browser.close();
}
Run it with node capture.mjs. The selector and target count are examples, not universal settings. Some sites append items to the document; others replace off-screen items, use a nested scrolling panel, or expose a “load more” button. Adapt the loop to the page’s actual behavior. For a robust capture, wait for a specific expected item, stable identifier, or visible section rather than relying only on total item count.
Playwright discourages using networkidle as a generic readiness condition for tests and recommends checking readiness with assertions. A page may keep analytics or other requests open, or it may finish network activity before the target content has rendered. Use a condition tied to the content you need. Playwright Page API
4. Python version
Install Playwright and Chromium:
python -m pip install playwright
python -m playwright install chromium
Save as capture.py. As with the JavaScript example, replace the URL and selector with ones from the target page.
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
url = "https://example.com/feed"
item_selector = "article"
target_count = 80
max_steps = 40
growth_timeout_ms = 8000
with sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
page = browser.new_page(viewport={"width": 1365, "height": 900})
try:
page.goto(url, wait_until="domcontentloaded", timeout=30000)
page.locator(item_selector).first.wait_for(state="visible", timeout=15000)
stagnant_steps = 0
previous_count = 0
for _ in range(max_steps):
count_before = page.locator(item_selector).count()
if count_before >= target_count:
break
page.evaluate("window.scrollBy(0, Math.floor(window.innerHeight * 0.75))")
try:
page.wait_for_function(
"({selector, count}) => document.querySelectorAll(selector).length > count",
arg={"selector": item_selector, "count": count_before},
timeout=growth_timeout_ms,
)
except PlaywrightTimeoutError:
pass
count_after = page.locator(item_selector).count()
if count_after <= previous_count:
stagnant_steps += 1
else:
stagnant_steps = 0
previous_count = count_after
if count_after >= target_count or stagnant_steps >= 2:
break
page.evaluate("window.scrollTo(0, 0)")
page.screenshot(path="screenshot.png", full_page=True)
finally:
browser.close()
5. Handle nested scrolling and virtualized lists
Sometimes the page itself does not scroll: a feed, table, or results panel has its own scrollbar. In that case, inspect the page for the element whose scrollTop changes and scroll that container in steps. Then wait for the relevant content to appear before capture. Playwright supports screenshots of individual elements, which can help when the desired output is one panel rather than the entire document. Playwright screenshot documentation
Virtualized interfaces may remove off-screen items from the DOM and render only the visible window of a long list. A single full-page image may therefore omit items even after you have scrolled through them. There is no universal fix established for every virtualized interface: capture separate sections, use the site’s export function if available, or use a page-specific method that assembles the desired content.
6. Verify the screenshot
- Check the top and bottom of the image against the intended start and end points.
- Look for missing batches or abrupt gaps where new items should appear.
- Confirm that the final expected item or section is present.
- Check whether a fixed header, overlay, consent banner, or chat widget obscures content.
- If the page uses a nested panel, verify that the panel content—not just the outer page—was loaded and captured.
- If content is virtualized, compare the image with the actual required items; scrolling through the list does not guarantee they all remain in one document image.
7. Troubleshooting
| Symptom | Likely cause | What to try |
|---|---|---|
| Screenshot ends at the viewport height | The capture used viewport mode. | Enable full-page capture: fullPage: true in JavaScript or full_page=True in Python. |
| Full-height image still misses lower sections | Scroll-triggered content was never loaded. | Scroll in increments, wait for the expected content at each boundary, then capture again. |
| The outer page scrolls but the feed does not load | The feed may be an independently scrolling container. | Find and scroll the inner container; wait for its items to change. |
| The item count grows, but items are missing in the final image | The site may virtualize the list and replace off-screen nodes. | Use a site export, capture separate sections, or build a page-specific output method. |
| The script times out waiting for content | The selector is wrong, the page is blocked, or no new batch appeared within the chosen timeout. | Confirm the selector against the live DOM, inspect whether the page loaded, and adjust the condition or timeout to the site’s behavior. |
networkidle never occurs |
Long-lived requests or background activity keep the network busy. | Wait for a content-specific locator or assertion instead of generic network quiet. |
| The capture is blank or only partially rendered | Navigation or rendering had not reached the required state. | Wait for a visible page-specific element and inspect navigation errors before capturing. |
| The bottom content is cut off after returning to the top | Returning to the top may prompt a site to unload items or alter the document. | Check the item count and final section after scrolling back; for virtualized pages, capture sections separately. |
8. Performance, reliability, and cost
Each scroll step adds waiting and browser work, so target only the amount of content needed. Use a finite maximum step count and a clear stop condition to avoid an unbounded loop on a page that keeps generating items. Reuse one browser process for multiple captures when appropriate, while giving each page its own navigation, selectors, and readiness checks.
Reliability depends on observing the page’s actual loading signal. Item counts can be misleading when content is replaced, duplicated, or rendered in a nested panel. A specific expected item or section is a stronger completion condition. Save the image and inspect it as an output artifact; a successful screenshot call only means the browser produced an image, not that the image contains every intended item.
Self-hosted Playwright has no per-screenshot API charge, but it does require a running browser environment and the engineering work to maintain it. Resource use and elapsed time vary with page length, assets, scroll behavior, and wait conditions; there is no universal timing or cost figure for this workflow.
Or skip the browser setup
ScreenshotNeo provides a one-request website screenshot API. Its full-page capture loads lazy images, but a site that requires actual scrolling to fetch more feed entries may still need a page-specific approach. It accepts common screenshot API parameter names, which can make switching straightforward. See the ScreenshotNeo API documentation for options and request details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/feed -o shot.webp
The API can capture images or PDFs and offers options such as viewport presets, custom waits, selectors, headers, cookies, and JavaScript. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.
FAQ
Does full-page mode scroll the page for me?
It requests a capture of the scrollable page. Do not assume it triggers every site’s scroll-dependent loading logic; load and verify the content first.
How far should I scroll between checks?
Use increments small enough to cross loading boundaries without skipping the opportunity to observe them. The right distance depends on the page’s layout and trigger behavior.
Can one screenshot include every item in a virtualized feed?
Not necessarily. A virtualized feed may keep only visible items rendered. Use a site export or capture separate sections if the items do not coexist in the document.


