How to Bulk Screenshot Pages with Lazy-Loaded Images Fully Rendered
Scroll each page to trigger lazy loading, verify its images, then capture and log a full-page screenshot. Runnable Playwright and Puppeteer batch examples included.
To bulk screenshot long pages with lazy-loaded images rendered, navigate to each page, scroll through it in increments to trigger deferred content, verify the images and sections you need, then take a full-page screenshot. In Playwright, fullPage: true captures the scrollable page, but does not itself trigger lazy loading. Network-idle is only a network condition, not proof that the page is visually complete. [Playwright Page API]
The examples below use site-specific readiness checks. No one scroll distance, delay, or image check works for every site: responsive images, CSS backgrounds, interaction-gated content, and custom loading logic may need additional checks.
1. Choose a browser tool and define readiness
Playwright and Puppeteer both provide screenshot APIs. Choose based on the browser engines, runtime, and interactions your project already needs; the available documentation does not establish one as universally faster or more reliable. Playwright supports Chromium, Firefox, and WebKit examples and a full-page screenshot option. Puppeteer documents page and element screenshots and locator interactions. [Playwright] [Puppeteer screenshots]
Before coding the batch, decide what “ready” means for your pages. A useful baseline is: the document navigated, a scroll pass completed, required images loaded with nonzero natural dimensions, and any important site-specific content appeared. Treat this as a starting point: image readiness does not detect a missing CSS background, a canvas that has not drawn, or content hidden behind an interaction.
2. Bulk capture with Playwright
Install Playwright and its Chromium browser in your project:
npm install playwright
npx playwright install chromium
Save this as bulk-screenshots.mjs and run node bulk-screenshots.mjs. It processes URLs sequentially, scrolls in viewport-sized steps, checks ordinary img elements, writes a full-page PNG, and records each outcome. Edit urls, outputDir, and the readiness policy for your workload.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const urls = [
'https://example.com/article-one',
'https://example.com/article-two',
];
const outputDir = './screenshots';
const resultsPath = path.join(outputDir, 'results.json');
const navigationTimeoutMs = 45_000;
const imageTimeoutMs = 20_000;
const maxScrollSteps = 80;
const pauseBetweenScrollsMs = 150;
function safeName(url, index) {
const parsed = new URL(url);
const base = `${parsed.hostname}${parsed.pathname}`
.replace(/[^a-z0-9]+/gi, '-')
.replace(/^-|-$/g, '')
.slice(0, 100);
return `${String(index + 1).padStart(3, '0')}-${base || 'page'}.png`;
}
async function scrollThrough(page) {
let previousHeight = 0;
let stableRounds = 0;
for (let step = 0; step < maxScrollSteps; step += 1) {
const state = await page.evaluate(() => ({
y: window.scrollY,
viewport: window.innerHeight,
height: document.documentElement.scrollHeight,
}));
const nextY = Math.min(state.y + Math.max(state.viewport, 400), state.height);
await page.evaluate(y => window.scrollTo(0, y), nextY);
await page.waitForTimeout(pauseBetweenScrollsMs);
const newHeight = await page.evaluate(() => document.documentElement.scrollHeight);
if (nextY >= newHeight - 2 && newHeight === previousHeight) stableRounds += 1;
else stableRounds = 0;
previousHeight = newHeight;
if (stableRounds >= 2) break;
}
await page.evaluate(() => window.scrollTo(0, 0));
}
async function checkImages(page) {
return page.evaluate(() => {
const images = [...document.images];
return {
total: images.length,
pending: images.filter(img => !img.complete).length,
failed: images
.filter(img => img.complete && img.naturalWidth === 0)
.map(img => ({ src: img.currentSrc || img.src, alt: img.alt })),
};
});
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
const page = await context.newPage();
page.setDefaultNavigationTimeout(navigationTimeoutMs);
for (const [index, url] of urls.entries()) {
const result = { url, status: 'started', startedAt: new Date().toISOString() };
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
result.httpStatus = response?.status() ?? null;
await scrollThrough(page);
// Wait up to the configured bound for image elements to settle.
await page.waitForFunction(
() => [...document.images].every(img => img.complete),
{ timeout: imageTimeoutMs },
).catch(() => {});
result.images = await checkImages(page);
const filename = safeName(url, index);
await page.screenshot({ path: path.join(outputDir, filename), fullPage: true });
result.file = filename;
result.status = 'captured';
if (result.images.failed.length) result.warning = 'Some img elements completed without a usable naturalWidth.';
} catch (error) {
result.status = 'failed';
result.error = error instanceof Error ? error.message : String(error);
}
result.finishedAt = new Date().toISOString();
results.push(result);
await writeFile(resultsPath, JSON.stringify(results, null, 2));
console.log(`${result.status}: ${url}${result.error ? ` — ${result.error}` : ''}`);
}
await context.close();
} finally {
await browser.close();
}
This script deliberately records failed image elements as warnings rather than claiming every site image has loaded. For strict capture jobs, change that policy so failed or still-pending required images mark the URL as failed and can be retried. If the page appends content as you approach the bottom, the scroll loop stops only after repeated stable bottom observations, subject to the maximum step count.
Playwright options to adapt
waitUntil:domcontentloadedis an initial navigation milestone;loadwaits for the load event.networkidlemeans no network connections for at least 500 ms, and Playwright discourages relying on it alone to assess readiness.commitwaits for the response to start arriving. Pick the milestone that fits the site, then apply content checks. [Playwright Page API]fullPage: true: expands the capture to the full scrollable document. It controls capture extent, not whether lazy resources were requested.viewport: use the viewport that matches the desired responsive layout. Mobile layouts can expose different content and image sources.deviceScaleFactor: configure the browser context if you need retina-scale output; higher scale increases image dimensions and memory use.- Readiness: add selectors for page-specific sections, or check a list of required images rather than every decorative image. Avoid treating a fixed sleep as a universal completion signal.
3. Puppeteer alternative
Puppeteer offers a similar workflow. Its documentation includes Page.screenshot() and locator scrolling interactions. [Screenshots] [Page interactions]
npm install puppeteer
import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
const urls = ['https://example.com/article-one', 'https://example.com/article-two'];
await mkdir('./screenshots', { recursive: true });
const browser = await puppeteer.launch({ headless: true });
const results = [];
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900 });
page.setDefaultNavigationTimeout(45_000);
for (const [index, url] of urls.entries()) {
const result = { url, status: 'started' };
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
result.httpStatus = response?.status() ?? null;
let lastHeight = 0;
let stable = 0;
for (let step = 0; step < 80; step += 1) {
const state = await page.evaluate(() => ({
y: scrollY, viewport: innerHeight, height: document.documentElement.scrollHeight,
}));
await page.evaluate(y => scrollTo(0, y), Math.min(state.y + Math.max(state.viewport, 400), state.height));
await new Promise(resolve => setTimeout(resolve, 150));
const height = await page.evaluate(() => document.documentElement.scrollHeight);
if (height === lastHeight && state.y + state.viewport >= height - 2) stable += 1;
else stable = 0;
lastHeight = height;
if (stable >= 2) break;
}
await page.evaluate(() => scrollTo(0, 0));
await page.waitForFunction(() => [...document.images].every(img => img.complete), { timeout: 20_000 }).catch(() => {});
result.images = await page.evaluate(() => {
const imgs = [...document.images];
return {
total: imgs.length,
pending: imgs.filter(img => !img.complete).length,
failed: imgs.filter(img => img.complete && img.naturalWidth === 0).map(img => img.currentSrc || img.src),
};
});
const filename = `./screenshots/${String(index + 1).padStart(3, '0')}.png`;
await page.screenshot({ path: filename, fullPage: true });
result.file = filename;
result.status = 'captured';
} catch (error) {
result.status = 'failed';
result.error = error instanceof Error ? error.message : String(error);
}
results.push(result);
console.log(JSON.stringify(result));
}
} finally {
await browser.close();
}
Puppeteer also exposes Page.waitForNetworkIdle(), which waits for an idle network condition. Like Playwright’s network-idle navigation state, that condition alone does not verify visual completeness. [Puppeteer waitForNetworkIdle]
4. Make the batch reliable
- Keep a result per URL. Store the URL, start and finish time, response status, image counts, output path, and error. The examples write or print this information so one bad page does not erase the rest of the batch.
- Separate failure types. Distinguish navigation timeout, non-success HTTP response, missing required image, screenshot failure, and output write failure. Retrying all of them identically can waste time or repeat a permanent failure.
- Retry selectively. Retry transient navigation or server failures with a small bounded retry count and backoff. Do not retry indefinitely, and do not assume a retry fixes access controls or a broken asset URL.
- Limit concurrency. Start sequentially as shown. If you add workers, cap concurrent browser pages based on available memory and the size of your pages. Each page can consume substantial memory, especially with large full-page images.
- Use deterministic names. Include an index or stable identifier because different URLs can normalize to the same filename. Keep a manifest mapping filenames to original URLs.
- Protect the run from partial output. Write results incrementally, and consider capturing to a temporary file before renaming it into place. Keep enough metadata to resume only failed URLs.
5. Readiness edge cases
Responsive and cross-origin images
document.images only covers image elements. currentSrc reports the chosen responsive source; naturalWidth can help identify an image that failed to produce usable dimensions. Cross-origin display can still render in a screenshot even when a script cannot inspect the image’s bytes. Record what the browser can observe and avoid treating the check as a universal visual validation.
CSS backgrounds, canvas, and embedded content
CSS background images are not in document.images. If they matter, identify their selectors and inspect computed styles or assert a site-specific loaded state. Canvas and embedded frames may need page-specific readiness signals; a generic image check will not prove they are complete.
Infinite scroll and very long pages
An infinite-scroll page may never reach a stable bottom. Set a maximum scroll step or total time, as the examples do, and define which portion is required. Full-page capture of an extremely tall page can consume significant memory and produce a very large image; split the page into sections if a single image is impractical.
Sticky headers and scroll-triggered effects
Scrolling can activate animations, sticky elements, or content that changes position. Return to the top before capture when the desired screenshot should start at the document beginning. If a sticky element appears repeatedly in a full-page image or an animation is mid-transition, use a site-specific CSS override or wait for a stable state. Check the output visually for important pages.
Interaction-gated content
Some sites load images only after a click, consent choice, or tab selection. Scrolling cannot trigger those states. Add the necessary interaction before checking readiness, and follow the site’s access and usage rules.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Images below the fold are blank | The screenshot was taken before those elements entered the viewport and triggered loading. | Scroll through the page before capture, pause according to the site’s behavior, and verify required images. |
| Navigation succeeded but images are missing | Navigation completion only marks a browser lifecycle event; it does not establish visual readiness. | Add image or section assertions after scrolling; do not use navigation success as the only check. |
networkidle times out or still yields missing content |
Background requests may continue, or the page may be visually incomplete despite a quiet network. | Choose a suitable navigation milestone and use explicit content checks. Network idle is not a completeness test. |
| Scroll loop stops before the bottom | The page height changes after a delay, or the step limit is too small. | Increase the bounded step limit, adapt the pause, and log the final scroll position and height. |
| Image check says loaded, but image looks wrong | complete can be true for failed images; responsive source selection or CSS backgrounds may be involved. |
Check naturalWidth, expected source or section state, and inspect background-image cases separately. |
| Screenshot is clipped or too large | Full-page extent can create a very tall bitmap with substantial memory use. | Capture sections, reduce viewport scale where acceptable, or use PDF when the deliverable is a document. |
| Browser crashes on a large batch | Too many pages or large screenshots are resident at once. | Use sequential or capped concurrency, close pages/contexts between groups, and reduce capture dimensions if appropriate. |
| Output file is missing or corrupt | Screenshot or filesystem write failed, or the process stopped mid-write. | Record write errors separately, capture to a temporary path, and verify file existence before marking success. |
7. Performance, reliability, and cost
Capture time depends on the target pages, network, browser, readiness conditions, image dimensions, and concurrency. The cited documentation does not provide a universal throughput or timeout recommendation. Start with a modest sequential run, measure your own pages, and increase concurrency only while memory and failure rates remain acceptable.
Use timeouts as bounds, not as readiness signals. A timeout that is too short can reject slow legitimate pages; an unbounded wait can stall the batch. Keep separate limits for navigation and content readiness, and preserve partial results so a failed URL does not force a full rerun.
Browser automation has no per-screenshot API charge in these examples, but you bear the compute, storage, network, and maintenance costs of running browsers. Ensure you have permission to capture and store the target content for your intended use; permissions depend on the site and context.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server by ScreenshotNeo. One GET request accepts a URL and returns an image or PDF. Its full-page capture option loads lazy images; it also supports bulk capture of up to 100 URLs per call. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie banners and removes known consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Does full-page screenshot mode trigger lazy loading?
Do not assume so. It expands the captured area; scroll the page first and verify the content you need.
Is network-idle enough to know the page is ready?
No. It describes network activity, not whether required images or visual sections rendered.
Should I use Playwright or Puppeteer?
Use the library that fits your runtime, browser requirements, and existing project. Both support screenshot workflows; readiness checks still need to match the site.
Can I use this for an infinite-scroll page?
Yes, with an explicit maximum extent or time budget. Define how much content to include because an endless feed has no final bottom.


