Why PageCrawl.io Screenshots Miss Lazy-Loaded Images and How to Fix It
PageCrawl can capture a long page before its lazy images load. Learn how scrolling, readiness checks, and site fixes can help.
PageCrawl screenshots can miss lazy-loaded images when the capture reaches the page before those images have been requested and rendered. A full-page screenshot is not necessarily the same as scrolling through the page: many sites load off-screen images only when they approach the viewport or when JavaScript responds to a scroll or visibility event.
To troubleshoot, reproduce the missing area, scroll the page in viewport-sized steps, let image requests settle, and capture again. PageCrawl publicly describes JavaScript-enabled browser capture and page actions, but its reviewed public materials do not document a dedicated “scroll before screenshot” or “wait for lazy images” switch. Check the controls available to your account or ask PageCrawl support about the specific page; do not assume a named setting exists.
1. Why lazy-loaded images are missing
Lazy loading postpones loading an image or other content until it is likely to be needed. A page may use the browser’s native loading="lazy" behavior, an IntersectionObserver, or a JavaScript library that reacts when content enters or approaches the viewport. The browser can then fetch the image and replace a placeholder.
A screenshot service may render a page and produce a full-page image without generating the same sequence of real viewport changes as a person scrolling. If the page’s loading logic never sees the relevant image become visible, the image request may not begin before capture. Even after scrolling triggers a request, the screenshot can still run before the response and rendering finish.
This explains a common failure mode; it does not establish PageCrawl’s internal capture sequence for a particular request. Google’s guidance describes the site-side mechanism: relevant lazy-loaded content should load when it is visible in the viewport. See Google Search Central’s lazy-loading guidance.
2. Diagnose which content is missing
- Open the affected URL in a normal browser and note the missing images or sections. Record whether they are below the initial viewport.
- Reload the page and scroll down through it in viewport-sized steps. Pause briefly after each movement so the site can respond.
- Watch whether the missing images appear after scrolling. If they do, the capture workflow likely needs to trigger the page’s loading behavior and wait for the resulting requests.
- If they remain absent in a normal browser, investigate the page implementation, network failures, consent requirements, or other interactions before blaming the screenshot capture.
- Repeat the screenshot after the page’s normal loading behavior has run, then compare the same regions.
Use a real browser to inspect the page’s rendered DOM and network activity when you control the environment. Check whether the image has a final URL in src (or a source selected from srcset) and whether its request succeeds. A placeholder or a URL stored only in a data attribute may indicate that site JavaScript has not promoted it to an actual image source yet.
3. Fix the capture workflow by scrolling before capture
The general remedy is to trigger the same visibility events a visitor would, then wait for the relevant images to finish loading before taking the screenshot. A single full-page operation may not do that. If you have browser automation available, the following Playwright example scrolls in viewport increments and waits for images to finish or fail before saving a full-page screenshot.
Runnable JavaScript example with Playwright
Install Playwright and its Chromium browser in a Node.js project:
npm install playwright
npx playwright install chromium
Save this as capture-lazy.js and run it with node capture-lazy.js https://example.com. Replace the example URL with the page you need to capture.
const { chromium } = require('playwright');
async function main() {
const url = process.argv[2];
if (!url) throw new Error('Usage: node capture-lazy.js https://example.com');
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
// Scroll through the document to trigger viewport-based lazy loading.
await page.evaluate(async () => {
const pause = (ms) => new Promise(resolve => setTimeout(resolve, ms));
const step = Math.max(300, Math.floor(window.innerHeight * 0.8));
let previousHeight = 0;
let stableRounds = 0;
for (let round = 0; round < 100; round++) {
const height = document.documentElement.scrollHeight;
for (let y = 0; y < height; y += step) {
window.scrollTo(0, y);
await pause(250);
}
window.scrollTo(0, document.documentElement.scrollHeight);
await pause(500);
const newHeight = document.documentElement.scrollHeight;
stableRounds = newHeight === previousHeight ? stableRounds + 1 : 0;
previousHeight = newHeight;
if (stableRounds >= 2) break;
}
});
// Wait for currently discovered images to finish loading. Failed images
// are reported below rather than blocking the capture indefinitely.
await page.evaluate(async () => {
const images = Array.from(document.images);
await Promise.all(images.map(img => {
if (img.complete) return Promise.resolve();
return new Promise(resolve => {
img.addEventListener('load', resolve, { once: true });
img.addEventListener('error', resolve, { once: true });
});
}));
});
const imageReport = await page.locator('img').evaluateAll(images =>
images.map(img => ({
src: img.currentSrc || img.src,
loaded: img.complete && img.naturalWidth > 0,
naturalWidth: img.naturalWidth
}))
);
const failed = imageReport.filter(image => !image.loaded);
if (failed.length) console.warn('Images not loaded:', failed);
await page.screenshot({ path: 'page.png', fullPage: true });
console.log(`Saved page.png; checked ${imageReport.length} images.`);
} finally {
await browser.close();
}
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
The scroll loop handles pages that grow as more content loads and stops after the document height remains stable for two rounds, with a 100-round safety limit. The image wait listens for both success and failure so one broken image cannot hang the run forever. Sites that append content only after a particular button click, require login, or use a custom interaction need a page-specific action; generic scrolling cannot infer those behaviors.
Use a page-specific readiness condition when possible
Scrolling and a fixed pause are practical fallbacks. For a stable production capture, wait for a selector that appears when the required content is ready, or check the particular image elements that matter. For example, after scrolling, Playwright can wait until a known image has a nonzero natural width:
await page.locator('#article img.hero').waitFor({ state: 'visible', timeout: 15000 });
await page.waitForFunction(() => {
const img = document.querySelector('#article img.hero');
return img && img.complete && img.naturalWidth > 0;
}, { timeout: 15000 });
Use selectors that match the target site. A visibility check alone does not guarantee the image request succeeded; checking complete and naturalWidth helps distinguish a rendered image from a broken or still-loading one. Some pages replace image elements as you scroll, so query them after the relevant section has appeared.
4. Fix the page if you own it
- Do not defer images that should appear immediately in the opening viewport. Lazy loading is intended for content that is initially off-screen.
- Make deferred content load when it becomes visible. Test by opening the page and scrolling through it, not only by checking the initial viewport.
- Ensure the real image URL is present in rendered markup when the image should be available. Google recommends checking that image or video URLs appear in the rendered HTML’s
srcattribute. - Check that JavaScript errors, blocked image hosts, or failed requests are not preventing the lazy-loading code from assigning the source.
- For content that must be available without scrolling or interaction, consider whether eager loading is more appropriate for that content.
These changes improve the page’s behavior for visitors and automated renderers. They do not guarantee that every capture service will wait for every resource, so capture readiness still matters.
5. What to check in PageCrawl
PageCrawl’s public product page describes a JavaScript-enabled browser, page actions such as waiting for text or clicking, and full-page screenshots. The reviewed material does not document a dedicated control named for scrolling before capture or waiting for lazy images. Check the current product controls and documentation for your account, and ask support whether a page-specific sequence can scroll the page and wait for its content.
If PageCrawl supports a wait-for-text or click action for your case, use it to wait for a reliable page-specific signal or perform an interaction the site requires. Do not treat a generic fixed delay as proof that all images loaded. If the page needs incremental viewport scrolling, confirm that the capture workflow can perform that behavior.
6. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. Its full-page capture loads lazy images. Here is the one-call API pattern; replace the target URL with the page you want to capture. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
7. Performance, reliability, and cost
Scrolling every viewport and waiting for resources takes longer than capturing immediately. Keep the viewport and scroll step appropriate for the page, and avoid an unbounded wait: use timeouts and report images that failed to load. Very long or dynamically growing pages may need a maximum scroll count, a page-specific end condition, or a cap on how many images you inspect.
Network idle can be a useful signal on pages with finite requests, but analytics, polling, and streaming connections can prevent it from occurring. A fixed delay is simple but can be too short on a slow connection and unnecessarily long on a fast one. Prefer a page-specific selector or image-complete condition for important content, with a timeout and a useful failure report.
For self-hosted browser automation, account for the browser runtime, network transfer, and any retries in your own infrastructure costs. Retry only transient failures, and avoid repeating a full long-page capture without checking whether the missing requests actually failed. ScreenshotNeo bills only clean shots; bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Its plans include 1,000 free monthly shots, then paid options from $5 for 3,000 up to $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan.
8. Troubleshooting common failures
| Symptom | Likely cause | What to try |
|---|---|---|
| Images below the fold are placeholders | The page has not received visibility or scroll events | Scroll in viewport-sized steps, pause for the page to react, then capture. |
| Images appear after manual scrolling but not in the capture | The capture may create a full-page image without equivalent incremental scrolling | Use a workflow that scrolls before capture, or confirm whether PageCrawl can perform the required action. |
| Some images appear, others do not | Different sections may load at different thresholds or require distinct site logic | Inspect each missing image’s rendered source and request status; wait for the specific elements that matter. |
| The screenshot is taken while images are still appearing | Capture starts before requests or rendering settle | Wait on a relevant selector or image-complete condition; use a bounded timeout. |
| Images remain broken after scrolling | The image URL may be invalid, blocked, or failing, or page JavaScript may have errored | Inspect the rendered src, browser console, and network response; fix the page or access issue. |
| Automation never reaches the bottom | The page grows as it loads more items, or the scroll loop has no stopping rule | Stop after document height stabilizes, detect a page-specific end marker, and set a maximum number of rounds. |
| Waiting for network idle times out | Persistent analytics or background requests keep the network active | Wait for the needed content selector or image state instead of global network idle. |
| Only content behind a control is missing | The page requires a click, consent choice, or other interaction | Identify the required interaction and add it explicitly; scrolling alone cannot trigger it. |
9. FAQ
Does a full-page screenshot always scroll the page?
No. Full-page output describes the captured area, but does not by itself establish that the browser triggered every viewport event used by a site’s lazy-loading code.
Is a longer delay enough?
Only if the image requests have already started and finish within that delay. If the page never received the event that starts loading, waiting longer may not help.
Does this prove PageCrawl has a bug?
No. The available public materials do not establish its internal capture sequence. The cause may be the site’s lazy-loading implementation, a required interaction, request failure, or capture timing.
How can I tell whether the image loaded?
In browser automation, inspect the rendered image URL and check that the image is complete with a nonzero natural width. Also check the browser’s network panel for a successful image response.


