Puppeteer Screenshot of a Lazy-Loaded Page: How to Capture Every Image
A full-page screenshot does not trigger every lazy image. Scroll the page, wait for images, validate what loaded, then capture.
page.screenshot({fullPage: true}) captures the full document, but it does not guarantee that scrolling has triggered every lazy-loaded image. For a more complete capture, set the viewport before navigation, scroll through the page in steps, wait for newly triggered content, check image loading, and then take the full-page screenshot. The exact readiness check depends on the site: an infinite feed, nested scroller, or interaction-triggered gallery needs a page-specific condition.
This guide uses Puppeteer with JavaScript. The scroll loop is practical implementation guidance, not a guarantee that every site’s lazy-loading mechanism will behave the same way. Puppeteer documents the screenshot, viewport, and waiting APIs in its screenshots guide and API references.
1. Install Puppeteer and choose a viewport
In a new project, install Puppeteer:
npm install puppeteer
Pick the viewport before navigating. Responsive layouts may request different images or use different lazy-loading thresholds at different widths. Puppeteer also notes that changing the viewport can reload a page in certain mobile or touch configurations, so set it up front when possible. See Page.setViewport.
const puppeteer = require('puppeteer');
const url = 'https://example.com/article';
const outputPath = 'page.png';
const viewport = { width: 1280, height: 900, deviceScaleFactor: 1 };
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport(viewport);
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
// Continue with the scroll, readiness checks, and screenshot below.
} finally {
await browser.close();
}
})();
domcontentloaded is a navigation milestone, not proof that images or application content are ready. The rest of the workflow deliberately triggers content after navigation.
2. Scroll to trigger lazy-loaded content
Many pages load images when they approach the viewport. Scroll in increments and allow the page to react after each step. Re-read the document height as you go because loading content may lengthen the page. Bound the loop so a feed that continually appends content cannot run forever.
async function scrollThroughPage(page, {
maxSteps = 100,
settleMs = 500,
stableBottomPasses = 3,
} = {}) {
let stablePasses = 0;
let lastHeight = 0;
for (let step = 0; step < maxSteps; step++) {
const state = await page.evaluate(() => ({
height: document.documentElement.scrollHeight,
y: window.scrollY,
viewport: window.innerHeight,
}));
if (state.height === lastHeight && state.y + state.viewport >= state.height) {
stablePasses++;
} else {
stablePasses = 0;
}
if (stablePasses >= stableBottomPasses) return;
lastHeight = state.height;
await page.evaluate(() => window.scrollBy(0, Math.max(1, window.innerHeight * 0.8)));
await new Promise(resolve => setTimeout(resolve, settleMs));
}
throw new Error(`Reached maxSteps (${maxSteps}) before the page bottom stabilized`);
}
The overlap between scroll steps helps avoid skipping content whose trigger zone is smaller than a viewport. The number of steps, delay, and bottom-stability count are tuning parameters, not universal values. For a known site, prefer a documented “loaded” signal or a selector that appears after the last item.
3. Wait for images and validate them
After scrolling, wait for image elements currently in the document to finish, up to a deadline. Then collect a report. An image can complete with no usable natural dimensions, so treat complete alone as insufficient. Some sites use CSS background images, canvases, or virtualized lists; those will not be fully covered by checking img elements.
async function waitForImages(page, timeoutMs = 15000) {
return page.evaluate(async (timeout) => {
const images = Array.from(document.images);
const pending = images.filter(img => !img.complete);
const deadline = new Promise(resolve => setTimeout(resolve, timeout));
await Promise.race([
Promise.all(pending.map(img => new Promise(resolve => {
img.addEventListener('load', resolve, { once: true });
img.addEventListener('error', resolve, { once: true });
}))),
deadline,
]);
return images.map(img => ({
src: img.currentSrc || img.src,
complete: img.complete,
naturalWidth: img.naturalWidth,
naturalHeight: img.naturalHeight,
}));
}, timeoutMs);
}
function findProblemImages(report) {
return report.filter(img => !img.complete || img.naturalWidth === 0 || img.naturalHeight === 0);
}
For targets that keep adding images while scrolling, run the scroll and image checks again after the final append signal. If a report contains broken images, decide whether to retry, wait for a site-specific state, or capture with the failure documented; a screenshot cannot restore an image the page never fetched.
4. Capture the full page
Once the page is ready, take the screenshot with fullPage: true. Puppeteer’s screenshot options also include output type, quality for JPEG or WebP, clipping, and other capture settings. Consult the ScreenshotOptions API for the version installed in your project.
await page.screenshot({ path: outputPath, fullPage: true, type: 'png' });
Here is the complete runnable script combining navigation, scrolling, image checks, and capture:
const puppeteer = require('puppeteer');
const url = process.argv[2] || 'https://example.com/article';
const outputPath = process.argv[3] || 'page.png';
async function scrollThroughPage(page, {
maxSteps = 100, settleMs = 500, stableBottomPasses = 3,
} = {}) {
let stablePasses = 0;
let lastHeight = 0;
for (let step = 0; step < maxSteps; step++) {
const state = await page.evaluate(() => ({
height: document.documentElement.scrollHeight,
y: window.scrollY,
viewport: window.innerHeight,
}));
if (state.height === lastHeight && state.y + state.viewport >= state.height) {
stablePasses++;
} else {
stablePasses = 0;
}
if (stablePasses >= stableBottomPasses) return;
lastHeight = state.height;
await page.evaluate(() => window.scrollBy(0, Math.max(1, window.innerHeight * 0.8)));
await new Promise(resolve => setTimeout(resolve, settleMs));
}
throw new Error(`Reached maxSteps (${maxSteps}) before the page bottom stabilized`);
}
async function waitForImages(page, timeoutMs = 15000) {
return page.evaluate(async (timeout) => {
const images = Array.from(document.images);
const pending = images.filter(img => !img.complete);
const deadline = new Promise(resolve => setTimeout(resolve, timeout));
await Promise.race([
Promise.all(pending.map(img => new Promise(resolve => {
img.addEventListener('load', resolve, { once: true });
img.addEventListener('error', resolve, { once: true });
}))),
deadline,
]);
return images.map(img => ({
src: img.currentSrc || img.src,
complete: img.complete,
naturalWidth: img.naturalWidth,
naturalHeight: img.naturalHeight,
}));
}, timeoutMs);
}
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1280, height: 900, deviceScaleFactor: 1 });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await scrollThroughPage(page);
const imageReport = await waitForImages(page);
const problems = imageReport.filter(img =>
!img.complete || img.naturalWidth === 0 || img.naturalHeight === 0
);
if (problems.length) {
console.warn(`${problems.length} image(s) may be missing or broken:`);
console.warn(problems);
}
await page.screenshot({ path: outputPath, fullPage: true, type: 'png' });
console.log(`Saved ${outputPath}; inspected ${imageReport.length} img elements.`);
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Run it with node capture.js https://example.com/article article.png. The script has bounded navigation, scroll, and image waits. Adapt it for the target site and inspect the resulting image; the algorithm cannot infer every application-specific loading rule.
5. Adapt the workflow to page behavior
Infinite scroll or content appended after a delay
A stable document height for a few passes is a useful stopping heuristic, but it does not prove a feed is complete. Set a maximum number of items or wait for a known end marker. If the page intentionally never ends, decide the capture boundary, such as the first 30 items, and stop there.
Nested scroll containers
window.scrollBy() only moves the main document. If the images live in an element with overflow: auto, identify that container and scroll it directly. Also inspect whether the element virtualizes content by removing off-screen items; a full-page document screenshot cannot include elements that the app has removed from the DOM.
Interaction-triggered images
Carousels, accordions, and galleries may require a click or keyboard action before an image is attached or loaded. Trigger the intended state before capture, and use a selector or application signal to confirm it. Puppeteer’s ElementHandle screenshot API scrolls an element into view when needed for that element’s screenshot; it does not mean a full-page screenshot scrolls every element into view.
Network idle
page.waitForNetworkIdle() can help settle requests after a scroll, but it is not a completeness signal: a page may defer requests until later scrolling or keep long-lived connections open. Puppeteer documents a default idleTime of 500 ms and concurrency of zero for its network-idle options. Use it after triggering content and keep a timeout or site-specific condition. See WaitForNetworkIdleOptions.
6. Choose full-page, clipped, or tiled captures
| Need | Approach | Tradeoff |
|---|---|---|
| One image of the document | fullPage: true |
Simple output, but tall output dimensions and memory use can grow with page height. |
| One section or region | clip with x, y, width, height |
Controls the output region; ensure the desired content has loaded first. |
| Very long page or fixed viewport states | Capture viewport-sized segments and stitch if appropriate | Requires overlap, seam checks, and care with sticky or animated elements. |
The API documents fullPage, clipping, and image format options; it does not establish a universal maximum page height. Test the actual target and output dimensions. See Page.screenshot.
7. Output options and capture details
- PNG: lossless and useful for sharp text or later inspection; files can be larger.
- JPEG: lossy and usually suited to photographic content; quality is configurable.
- WebP: available as an output type in supported Puppeteer/Chrome versions; quality can be configured. Check the API for your installed version.
- Clip: capture a specified rectangle rather than the full document.
- Viewport and scale: viewport dimensions determine responsive layout; device scale factor affects pixel density and output size.
Use a stable viewport and keep animations or changing content in mind if repeatable captures matter. If the page changes while capture runs, the image may represent a mix of states; wait for the application’s settled state or hide/disable motion with page-specific CSS where appropriate.
Or skip the browser setup
With ScreenshotNeo, one GET request returns a website screenshot. Its API offers full-page capture with lazy images loaded; see the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, and failed loads are not billed, and cache hits are not billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Images are missing below the fold | The capture started before scroll-triggered loading. | Scroll in overlapping steps, wait after each step, and verify image dimensions before capture. |
| The loop stops too early | The page appends content after a longer delay or uses a late end condition. | Increase the settling window, use a known end marker, and keep a bounded maximum. |
| Wait for network idle times out | Persistent connections or continuous requests prevent idle. | Use a selector or application readiness condition; treat network idle as a helpful signal, not a requirement. |
| Images report complete but appear broken | A failed image can still be complete with zero natural dimensions. | Check naturalWidth and naturalHeight, inspect the image URL and browser console, then retry only if the site can recover. |
| Some images never appear in the report | They may be CSS backgrounds, canvas content, or in a virtualized list. | Inspect the relevant elements and application behavior; scroll the correct container or capture a deliberate viewport sequence. |
| The screenshot is unexpectedly huge | A tall document, high device scale factor, or lossless format increases dimensions or file size. | Use a lower scale factor, JPEG/WebP where suitable, a clip, or viewport-sized segments. |
| Page layout differs from the normal browser view | The viewport, mobile emulation, or device scale factor changed responsive behavior. | Set the intended viewport before navigation and match the target device settings. |
Performance, reliability, and cost
Scrolling and waiting add time roughly in proportion to the number of steps and each step’s settling delay. Large full-page images also take more memory and produce larger files, particularly at high device scale factors. Keep the viewport and output format matched to the downstream use, and use a selector or completion event when one is available instead of an unnecessarily long fixed sleep.
Reliability comes from page-specific conditions plus validation: bound navigation and scroll work, detect broken image dimensions, record which URLs failed, and inspect the final capture. Retry only transient failures and cap retries so an unavailable page does not create an endless job. A browser screenshot has no per-capture Puppeteer service fee, but running Chrome consumes compute and may require browser dependencies and operational maintenance in your environment.
FAQ
Does fullPage: true load lazy images?
It requests a full-page image. It does not document a promise to scroll the page and trigger each website’s lazy-loading behavior, so scroll and validate first.
Should I use networkidle0 or networkidle2?
Use the condition that fits the site as a settling aid, with a timeout. Neither proves that content scheduled for a later scroll has been requested.
Can I capture every image on an infinite page?
Only after defining what “every” means for an unbounded feed, such as a known item count or end marker. Set a deliberate boundary and verify it.
Why does scrolling not reveal all images?
The page may use a nested scroller, require an interaction, defer loading beyond the chosen wait, or virtualize off-screen content. Inspect the page’s own loading behavior and adapt the trigger and completion check.


