ScreenshotNeo

BlogHow-to

How to Capture Lazy-Loaded Product Images with Selenium Python

Scroll product images into view, wait for their real pixels to load, and capture them with Selenium Python. Includes robust waits, troubleshooting, and an API option.

By the ScreenshotNeo team4 October 20268 min read

To capture lazy-loaded product images with Selenium Python, scroll each image into or near the viewport, wait until the image has a real source and has finished loading, then take a screenshot. A successful driver.get() or a complete document state does not mean a JavaScript-driven gallery or its images are ready.

The exact selector and readiness condition depend on the retailer’s markup. The example below uses product images under .product-gallery; inspect the page and replace that selector and the placeholder checks to match the site.

1. Install Selenium and a browser

Install Selenium and make a supported browser available. Recent Selenium releases can manage a compatible driver through Selenium Manager when the browser is installed.

python -m pip install selenium

Save the following as capture_products.py. It writes a PNG screenshot of each gallery image element after scrolling it into view and waiting for its image data to load.

2. Scroll images into view, wait, and capture

from pathlib import Path
import re

from selenium import webdriver
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait

URL = "https://example.com/products/example-item"
IMAGE_SELECTOR = ".product-gallery img"
OUTPUT_DIR = Path("product-images")
OUTPUT_DIR.mkdir(exist_ok=True)

options = webdriver.ChromeOptions()
# Use headless mode for unattended runs. Remove this line to see the browser.
options.add_argument("--headless")

# 'normal' waits for the page load event; it does not guarantee that
# JavaScript-rendered content or lazy images are ready.
options.page_load_strategy = "normal"

driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 20)

try:
    driver.get(URL)

    # Wait for the gallery markup to exist before collecting image elements.
    wait.until(lambda d: d.find_elements(By.CSS_SELECTOR, IMAGE_SELECTOR))
    images = driver.find_elements(By.CSS_SELECTOR, IMAGE_SELECTOR)

    for index, image in enumerate(images, start=1):
        # Scrolling is the trigger for viewport-based lazy loading.
        driver.execute_script(
            "arguments[0].scrollIntoView({block: 'center'});", image
        )

        # Wait for a non-placeholder source and decoded image dimensions.
        # Adapt this condition for sites using data-src, srcset, CSS backgrounds,
        # or a custom gallery loader.
        try:
            wait.until(lambda d: d.execute_script("""
                const img = arguments[0];
                const src = img.currentSrc || img.src || '';
                const placeholder = /placeholder|spacer|blank/i.test(src);
                return !placeholder && img.complete &&
                       img.naturalWidth > 0 && img.naturalHeight > 0;
            """, image))
        except TimeoutException:
            src = image.get_attribute("currentSrc") or image.get_attribute("src")
            print(f"Skipping image {index}: image did not load ({src})")
            continue

        safe_name = re.sub(r"[^a-zA-Z0-9_-]+", "_", f"product_{index}")
        image.screenshot(str(OUTPUT_DIR / f"{safe_name}.png"))
        print(f"Saved {OUTPUT_DIR / f'{safe_name}.png'}")
finally:
    driver.quit()

Run it with:

python capture_products.py

This captures each matched image element. If you need a viewport screenshot showing the surrounding product page, use driver.save_screenshot("page.png") after scrolling to the desired position and waiting for the relevant images.

3. Choose a wait that matches the page

There are separate states to consider: navigation completion, gallery element presence, image source assignment, image fetch completion, and visibility. Wait for the state your capture needs. Selenium’s explicit waits poll a condition until it succeeds or times out; this is more reliable than guessing with a fixed sleep.

Page behavior Useful action What to verify
Native loading="lazy" images Scroll each image into or near view complete and nonzero naturalWidth
Placeholder replaced by JavaScript Scroll, then wait for source change Actual source is no longer the placeholder
Images added while scrolling Scroll incrementally and re-query the DOM New gallery elements appear and load
Gallery changes image on selection Trigger the relevant thumbnail or control Wait for the displayed image source or dimensions to change

Do not assume every site uses src. An image may use data-src, srcset, a CSS background, or a custom JavaScript loader. Inspect the live element in browser developer tools and adjust the locator and wait condition accordingly. A browser’s own lazy-load threshold is heuristic and can vary; no single scroll distance guarantees that all images have loaded.

For pages that reveal more products as you scroll

Re-query elements after each scroll because galleries may append nodes after the initial page load. A basic incremental scroll loop can trigger loading across the page:

import time

previous_height = 0
while True:
    height = driver.execute_script("return document.body.scrollHeight")
    if height == previous_height:
        break
    previous_height = height
    driver.execute_script("window.scrollTo(0, arguments[0]);", height)
    # Prefer a condition for newly appended content when you can identify it.
    # This short pause only gives the page a chance to react; it is not proof
    # that images are loaded.
    time.sleep(0.3)

# Re-query after scrolling; the DOM may have changed.
images = driver.find_elements(By.CSS_SELECTOR, IMAGE_SELECTOR)

For a dynamic infinite-scroll page, a fixed number of passes or a site-specific stop condition is safer than waiting for the document height to stop once. Stop when the target product set is present, a “load more” control disappears, or the page’s own completion signal is reached. Then scroll each target image into view and apply its image-loaded wait before capturing.

4. Capture image elements or the current viewport

Selenium’s WebDriver screenshot methods save the current window as PNG or return PNG data. The element screenshot method is convenient when each image is the desired output. For a browser viewport screenshot, scroll to the relevant section and call:

driver.save_screenshot("product-page-viewport.png")

Element screenshots do not include the surrounding page. A viewport screenshot includes whatever is currently visible. The cited Selenium API documents these screenshot outputs; do not rely on a universal full-page screenshot method across browsers. If you need a full-page image, confirm that your chosen browser and capture approach support it, or capture sections and assemble them with a method appropriate to your use case.

5. Page-load options and reliability

Selenium’s page-load strategy controls how long navigation waits:

Strategy Navigation behavior Use with lazy images
normal Waits for the load event Good default, but still wait for the gallery and image conditions
eager Returns when the DOM is ready while some resources may still load Can reduce navigation waiting; explicit image waits remain necessary
none Returns without waiting for page readiness Use only when the script manages all required readiness conditions itself

Single-page applications can continue changing after any navigation readiness state. For dependable runs, set explicit timeouts, capture diagnostics when a wait expires, and always close the driver in a finally block. If a site returns different markup to automated browsers, verify the page and selector in a normal browser session; do not treat a missing image as proof that scrolling failed.

6. Troubleshooting

Symptom Likely cause Fix
No images found Selector does not match, gallery has not rendered, or content is inside a frame Inspect the DOM, wait for the gallery, and switch into the correct frame if needed
Saved image is blank or tiny Screenshot happened before loading, or the element is a placeholder Scroll it into view and wait for complete plus nonzero natural dimensions; check the actual source
Wait times out despite a visible image The page uses a different source attribute, CSS background, or image replacement pattern Inspect the element and wait on its actual loaded state or the gallery’s own signal
Only the first few images are captured Remaining images have not entered the viewport, or more nodes are added during scroll Scroll incrementally, re-query after page changes, then visit each target image
Image source is valid but request fails Image request may depend on cookies, headers, authentication, or a page interaction Check browser network errors and reproduce the required session or interaction before capture
Driver or browser fails to start Browser installation, driver compatibility, or runtime configuration issue Install a supported browser, update Selenium, and review Selenium Manager or driver startup output

7. Performance and cost considerations

Scrolling and waiting for each image is more controlled than a long arbitrary delay, but it adds browser work and network requests. Keep waits bounded, capture only the selectors and image sizes you need, and avoid repeatedly loading the same page when a cached result is suitable. For parallel jobs, account for the memory and browser-process cost of each session and limit concurrency to what the host can sustain. Site behavior and network speed determine completion time; there is no universal delay that guarantees readiness.

A local Selenium workflow uses your own browser and compute resources. If you run it as a recurring service, include browser maintenance, retries, storage, and failed-navigation handling in your operating cost. The screenshot API option below is a per-shot service with a free tier and stated paid-plan quotas.

8. Or skip the browser setup

If you need the screenshot rather than browser automation, ScreenshotNeo accepts a URL in one GET request and returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for options and parameters. For example, cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page info, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Learn about ScreenshotNeo, then sign up for 1,000 free screenshots a month with no card.

9. FAQ

Will this work on every product site?

The workflow applies broadly, but selectors and readiness conditions are site-specific. Inspect the page’s markup and loading behavior before choosing them.

Does document readiness mean every product image is loaded?

No. Lazy images may not be requested until they approach the viewport, and JavaScript may continue changing the page after navigation returns.

Can Selenium save an image element directly?

Yes. Selenium’s element screenshot method saves that element as PNG. Use a window screenshot instead when you want the visible browser viewport.

Why does scrolling by a fixed distance sometimes miss images?

Browser loading distance is heuristic, and pages can use custom loading code. Scroll target images into view and wait for the condition that confirms each image is usable.

Sources