Puppeteer Infinite Scroll Screenshot Has Blank Images: How to Fix It
Full-page screenshots do not trigger infinite-scroll loading. Scroll the page deliberately, wait for the right content and images, then capture with bounded checks.
A Puppeteer full-page screenshot captures the page’s rendered state; it does not make an infinite-scroll application load content that has not yet been requested. Scroll through the page to trigger lazy loading, wait for the intended items and their images to become ready, and only then capture. Use a site-specific completion signal where possible, and bound every loop and wait with a timeout.
fullPage: true changes screenshot geometry. It is not an instruction to scroll indefinitely or fetch every item from an infinite feed. Puppeteer’s screenshot guide documents full-page capture; its interaction guide covers scrolling and interacting with page content.
Why infinite-scroll screenshots have blank images
Many pages defer work until an element is near the viewport. The browser may not request an image until scrolling triggers an intersection observer or another site-specific handler. Taking a full-page screenshot does not reliably trigger those handlers for every item.
Even after content appears, several separate states can look like “loaded”:
- The item exists in the DOM, but its image request has not finished.
- The image request finished with an error. An image can be
completewhile still having no usable image data. - The page is quiet on the network, but JavaScript has not updated the DOM or painted the image yet.
- The visual is a CSS background, canvas, or content inside a nested scroller rather than a regular document
<img>.
Puppeteer’s network-idle wait reports a network condition, not proof that every visual is successfully rendered. The documented default idle interval is 500 ms; a quiet interval is a useful signal, but not a completion guarantee.
Use bounded scrolling and explicit readiness checks
The following Node.js script is a runnable starting point for a page that appends items as the document scrolls. It uses a maximum number of scrolls, an overall deadline, a stable-height condition, and checks HTML image success before capture. Replace ITEM_SELECTOR and, ideally, EXPECTED_COUNT with signals that match the target site. Install Puppeteer with npm install puppeteer, save this as screenshot.mjs, then run node screenshot.mjs https://example.com.
import puppeteer from 'puppeteer';
const url = process.argv[2];
if (!url) throw new Error('Usage: node screenshot.mjs <url>');
const MAX_SCROLLS = 40;
const REQUIRED_STABLE_PASSES = 3;
const EXPECTED_COUNT = 0; // Set to a known target count if the page has one.
const ITEM_SELECTOR = 'article'; // Replace with the site's repeated item selector.
const OVERALL_TIMEOUT_MS = 90_000;
const IMAGE_TIMEOUT_MS = 15_000;
const deadline = Date.now() + OVERALL_TIMEOUT_MS;
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
let stablePasses = 0;
let previousHeight = -1;
let reachedTarget = EXPECTED_COUNT === 0;
for (let i = 0; i < MAX_SCROLLS && Date.now() < deadline; i++) {
const state = await page.evaluate((selector) => ({
height: document.documentElement.scrollHeight,
count: document.querySelectorAll(selector).length,
atBottom: window.scrollY + window.innerHeight >= document.documentElement.scrollHeight - 2
}), ITEM_SELECTOR);
if (state.height === previousHeight) stablePasses++;
else stablePasses = 0;
previousHeight = state.height;
if (EXPECTED_COUNT > 0 && state.count >= EXPECTED_COUNT) reachedTarget = true;
if (reachedTarget && stablePasses >= REQUIRED_STABLE_PASSES) break;
await page.evaluate(() => window.scrollBy(0, Math.max(1, window.innerHeight * 0.8)));
// Prefer a site's new-item or loading-indicator condition here when available.
// This short pause gives scroll-triggered handlers a chance to run; it is not proof of readiness.
await new Promise(resolve => setTimeout(resolve, 250));
try {
await page.waitForNetworkIdle({ idleTime: 500, timeout: 2_000 });
} catch {
// Persistent requests can prevent network idle. The bounded loop continues;
// record this condition in production and require a site-specific signal.
}
}
const finalState = await page.evaluate((selector) => ({
count: document.querySelectorAll(selector).length,
height: document.documentElement.scrollHeight
}), ITEM_SELECTOR);
if (Date.now() >= deadline) throw new Error('Overall page readiness deadline reached.');
if (EXPECTED_COUNT > 0 && finalState.count < EXPECTED_COUNT) {
throw new Error(`Only ${finalState.count} of ${EXPECTED_COUNT} expected items appeared.`);
}
const images = await page.evaluate(async (timeoutMs) => {
const imgs = [...document.images];
const waitForOne = img => img.complete
? Promise.resolve()
: new Promise(resolve => {
img.addEventListener('load', resolve, { once: true });
img.addEventListener('error', resolve, { once: true });
});
await Promise.race([
Promise.all(imgs.map(waitForOne)),
new Promise(resolve => setTimeout(resolve, timeoutMs))
]);
return imgs.map(img => ({
src: img.currentSrc || img.src,
complete: img.complete,
naturalWidth: img.naturalWidth
}));
}, IMAGE_TIMEOUT_MS);
const failedImages = images.filter(img => !img.complete || img.naturalWidth === 0);
if (failedImages.length) {
console.error(`Image readiness check: ${failedImages.length} image(s) incomplete or failed.`);
console.error(failedImages);
}
await page.evaluate(() => window.scrollTo(0, 0));
await page.screenshot({ path: 'page.png', fullPage: true });
console.log(`Saved page.png; ${finalState.count} items, ${failedImages.length} incomplete/failed HTML images.`);
} finally {
await browser.close();
}
The script deliberately reports incomplete or failed images instead of treating them as successful. Adapt the stopping rule to the page: if the feed has a known number of records, wait for that count; if it shows a loading indicator, wait for the indicator to disappear after new items appear. A stable document height alone can mean the feed is finished, or simply that a request failed or the site requires a button click. Do not silently interpret it as proof of completeness.
Make the completion condition fit the page
- Choose the actual content signal. Count a repeated card selector, wait for a known last item, or observe a site-specific end marker. If the site has a “Load more” control, click it as needed and wait for the next batch.
- Trigger lazy loading. Scroll in viewport-sized or smaller increments. Some sites require an element to enter the viewport, so one jump to the bottom may skip triggers.
- Wait for the batch. Prefer a new-item count or loading state. Network idle can supplement that signal, but analytics, polling, streaming, or other persistent requests can make it time out.
- Check image outcomes. For HTML images, inspect
completeandnaturalWidth. Decide whether failed images should fail the job, be retried, or be recorded and captured as-is. - Return to the capture position. Scroll back to the top if the desired output begins there, then take the full-page screenshot.
Puppeteer’s page.evaluate() API can run and await page-context promises, which is useful for bounded checks. For advanced cases, use a MutationObserver to detect appended items or wait for the site’s own loading marker; keep the observer bounded by a timeout.
Handle backgrounds, nested scrollers, and other edge cases
CSS background images
document.images only covers HTML image elements. It does not validate URLs in background-image. Identify the relevant elements and inspect their computed backgroundImage, then wait for the corresponding resource or a site-specific ready marker. An older Puppeteer issue describes a background image not appearing after the load event; it is a historical report, not evidence that every current version has the same defect.
Nested scroll containers
Some feeds scroll inside a panel with overflow: auto, not the window. Scrolling window will not trigger that panel’s lazy content. Find the scrollable container, scroll that element, and measure its scrollHeight, clientHeight, and scrollTop. A full-page screenshot expands the document viewport; it may not expand an internal panel into a complete feed. Consider capturing the panel separately or using the site’s data endpoint when appropriate.
Other rendering cases
<picture>and responsive sources: inspectcurrentSrc, since the browser may select a source different fromsrc.- Canvas and WebGL: image readiness checks do not apply. Wait for the application’s render-complete signal or a known canvas state.
- Shadow DOM:
document.imagesdoes not automatically traverse every shadow root. Query relevant roots or use an application-level readiness signal. - Animations and transitions: they can change pixels after DOM readiness. Disable them with test CSS if that is acceptable, or wait for a stable visual state.
- Fonts: wait for
document.fonts.readyif font swaps alter layout or cause a capture before text settles. - Cross-origin resources: the page can often display them normally, but browser security may restrict reading their pixels through canvas. Do not use canvas pixel inspection as a universal readiness test.
- Consent dialogs or overlays: if the page blocks content behind a dialog, handle the site’s intended consent flow before waiting for the feed.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Top items appear, lower images are blank | Scroll-triggered loading never ran for lower items | Scroll incrementally through the relevant range, then wait for each appended batch. |
| Items appear but images remain blank | Capture started before image completion, or image requests failed | Wait for image events with a timeout; report naturalWidth === 0 and inspect the failed URL and browser console. |
| Network-idle wait times out | Polling, analytics, streaming, or long requests keep activity open | Use a site-specific item or loading-state condition. Keep network idle as a bounded supplementary check. |
| Scroll loop stops before the feed ends | Height stayed stable temporarily, a request failed, or a “Load more” action is needed | Require a known count or end marker; inspect loading errors and click the control when the page requires it. |
| HTML image checks pass but visual is blank | The visual is a CSS background, canvas, shadow-DOM image, or internal frame | Use the readiness signal that matches that rendering mechanism. |
| Screenshot is clipped despite full-page mode | Content lives in a nested scroller or the page changes while capture runs | Scroll the correct container; stabilize the page before capture or capture the container separately. |
| Script reports expected count not reached | Selector is wrong, the site has fewer results, or content is blocked | Verify the selector in DevTools, confirm the expected count, and check authentication, consent, and bot-check states. |
| Images fail only in automation | Authentication, rate limits, bot checks, or request headers differ from a normal browser session | Inspect response status and console/network errors; provide the required session state or headers where authorized. |
Historical reports include missing content after lazy-loaded elements, partial full-page screenshots, and questions about waiting for images. These reports date from 2017–2018 and do not establish a current bug in every Puppeteer release. Reproduce against your Puppeteer and Chromium versions and the target site’s actual loading behavior.
Performance, reliability, and cost
- Bound the work: cap scroll count, per-wait timeouts, and overall job duration. Infinite feeds can otherwise run indefinitely.
- Avoid arbitrary long sleeps: short pauses can let scroll handlers start, but explicit content conditions reduce wasted time and false readiness.
- Limit what you load: stop at the needed item count or page boundary. Scrolling an entire feed increases runtime, memory use, network traffic, and screenshot size.
- Keep failure visible: log the stopping reason, item count, failed image URLs, timeouts, and browser console errors. A screenshot file existing does not mean it is complete.
- Expect visual variance: changing content, ads, animation, fonts, and timing can change pixels between captures. Disable or wait for these only when the use case allows it.
- Cost depends on the environment: a local Puppeteer run has no ScreenshotNeo API charge, but consumes browser compute, memory, bandwidth, and engineering time. Hosted browser runtime may add provider charges.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. For a straightforward page capture, call its API once; see the ScreenshotNeo API documentation for options and parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
- An MCP server lets Claude, Cursor, and other MCP clients use
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Does fullPage: true scroll the page for me?
No. It requests a screenshot covering the document’s full page dimensions. Drive scroll-dependent loading before capture.
Is networkidle2 enough?
It can help, but it only describes network activity. Pair it with a check that the target content appeared and, where relevant, that its images succeeded.
How many scrolls should I use?
There is no universal number. Stop at a known item count or end marker when possible; otherwise use a maximum scroll count and a stable-state rule, then treat reaching the cap as an incomplete result.
Why do some images remain blank after img.complete is true?
complete also covers failed loads. Check naturalWidth and inspect the request outcome; use another method for CSS backgrounds, canvas, or other non-HTML-image content.


