ScreenshotNeo

BlogHow-to

Take Bulk Website Screenshots in Java with Selenium for an Indian Catalogue

Capture a controlled list of catalogue pages with Java and Selenium, save each result safely, and choose a verified method for full-page screenshots.

By the ScreenshotNeo team4 October 202614 min read

To take bulk website screenshots in Java with Selenium, put an explicit list of URLs in your application, open each URL with a WebDriver session, wait for the page state you need, capture a screenshot, and save it under a unique filename. Selenium provides screenshot operations for a browser context and for an individual element; the batch loop, retries, output naming, and result log are application logic. A normal screenshot should not be assumed to include the entire document: decide whether you need the viewport, one element, or a verified full-page capture.

This guide uses an unnamed Indian catalogue as the workload context. No specific catalogue, selector, access policy, or site behavior is assumed or claimed to have been tested. Confirm authorization and site rules for your actual targets, and choose a request rate appropriate for them.

1. Set up Java, Selenium, and Chrome

The example uses Maven, Selenium Java, and ChromeDriver. Use a Java version and Selenium release supported by your environment; keep the Selenium dependency current according to its official documentation. Selenium Manager can manage browser drivers for standard setups, while locked-down or pinned environments may need an explicitly managed driver installation.

<!-- pom.xml -->
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd">
  <modelVersion>4.0.0</modelVersion>
  <groupId>example</groupId>
  <artifactId>catalogue-capture</artifactId>
  <version>1.0.0</version>
  <properties>
    <maven.compiler.release>17</maven.compiler.release>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
    <selenium.version>4.27.0</selenium.version>
  </properties>
  <dependencies>
    <dependency>
      <groupId>org.seleniumhq.selenium</groupId>
      <artifactId>selenium-java</artifactId>
      <version>${selenium.version}</version>
    </dependency>
  </dependencies>
</project>

The version shown makes the sample reproducible, but check the Selenium project documentation for a suitable supported release in your environment. Official Selenium screenshot examples use TakesScreenshot, OutputType.FILE, and copy the resulting file to the desired location. See Selenium’s WebDriver documentation.

2. Capture a list of pages and keep per-URL results

This runnable program reads one URL per line from urls.txt, opens each page, waits for the document readiness state, takes a viewport screenshot, and writes one PNG per URL. It continues after individual failures and writes a CSV outcome log. The explicit wait is a baseline: for dynamic pages, replace or extend it with an observable page-specific condition.

// src/main/java/example/CatalogueCapture.java
package example;

import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.time.Duration;
import java.time.Instant;
import java.util.ArrayList;
import java.util.List;
import java.util.Locale;
import java.util.concurrent.TimeUnit;
import java.util.stream.Collectors;

import org.openqa.selenium.Dimension;
import org.openqa.selenium.JavascriptExecutor;
import org.openqa.selenium.OutputType;
import org.openqa.selenium.TakesScreenshot;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.support.ui.WebDriverWait;

public class CatalogueCapture {
    private static String safeName(int index, String url) {
        String host;
        try {
            host = java.net.URI.create(url).getHost();
        } catch (Exception e) {
            host = "page";
        }
        if (host == null || host.isBlank()) host = "page";
        String slug = host.toLowerCase(Locale.ROOT).replaceAll("[^a-z0-9.-]+", "-");
        return String.format(Locale.ROOT, "%04d-%s.png", index, slug);
    }

    private static void waitForDocument(WebDriver driver) {
        new WebDriverWait(driver, Duration.ofSeconds(30)).until(d ->
            "complete".equals(((JavascriptExecutor) d)
                .executeScript("return document.readyState")));
    }

    public static void main(String[] args) throws IOException {
        Path input = Path.of(args.length > 0 ? args[0] : "urls.txt");
        Path output = Path.of(args.length > 1 ? args[1] : "screenshots");
        Files.createDirectories(output);
        List<String> urls = Files.readAllLines(input, StandardCharsets.UTF_8).stream()
            .map(String::trim)
            .filter(line -> !line.isEmpty() && !line.startsWith("#"))
            .collect(Collectors.toList());

        ChromeOptions options = new ChromeOptions();
        // For a headless batch in a server environment, enable headless mode.
        options.addArguments("--headless=new", "--window-size=1440,1000");

        List<String> report = new ArrayList<>();
        report.add("index,url,timestamp,status,file,error");
        WebDriver driver = new ChromeDriver(options);
        try {
            driver.manage().timeouts().pageLoadTimeout(Duration.ofSeconds(60));
            driver.manage().window().setSize(new Dimension(1440, 1000));
            for (int i = 0; i < urls.size(); i++) {
                String url = urls.get(i);
                String fileName = safeName(i + 1, url);
                String timestamp = Instant.now().toString();
                try {
                    driver.get(url);
                    waitForDocument(driver);
                    Path target = output.resolve(fileName);
                    Files.copy(((TakesScreenshot) driver).getScreenshotAs(OutputType.FILE), target,
                        java.nio.file.StandardCopyOption.REPLACE_EXISTING);
                    report.add((i + 1) + ",\"" + url.replace("\"", "\"\"") + "\",\"" + timestamp
                        + "\",OK,\"" + fileName + "\",");
                    System.out.println("Saved " + target);
                } catch (Exception e) {
                    String message = e.getClass().getSimpleName() + ": " + String.valueOf(e.getMessage());
                    report.add((i + 1) + ",\"" + url.replace("\"", "\"\"") + "\",\"" + timestamp
                        + "\",ERROR,\"",\"" + message.replace("\"", "\"\"") + "\"");
                    System.err.println("Failed " + url + " — " + message);
                    // Recover from a stuck navigation before attempting the next URL.
                    try { driver.manage().timeouts().pageLoadTimeout(Duration.ofSeconds(60)); }
                    catch (Exception ignored) { }
                }
            }
        } finally {
            driver.quit();
        }
        Files.write(output.resolve("results.csv"), report, StandardCharsets.UTF_8);
    }
}

Provide an input file such as this, one URL on each line:

# urls.txt
https://example.com/catalogue/category-a
https://example.com/catalogue/category-b
https://example.com/catalogue/item-123

Run it with Maven after placing the Java source in the matching package path:

mvn compile exec:java -Dexec.mainClass=example.CatalogueCapture -Dexec.args="urls.txt screenshots"

The Maven Exec plugin is not declared above. Either add org.codehaus.mojo:exec-maven-plugin to the POM, or compile and run through your IDE. To make the command above directly runnable, add this plugin configuration under <build><plugins>:

<plugin>
  <groupId>org.codehaus.mojo</groupId>
  <artifactId>exec-maven-plugin</artifactId>
  <version>3.5.0</version>
</plugin>

Make the capture deterministic

  1. Use a controlled URL list. Validate each URL before starting, and assign stable IDs if the same host can appear more than once.
  2. Keep viewport dimensions, browser version, device scale, locale, timezone, and authentication state consistent when images will be compared over time.
  3. Wait for the state that matters. document.readyState == complete means the document load lifecycle completed; it does not prove that client-side catalogue data, lazy images, or fonts have finished updating.
  4. Save each image immediately and record timestamp, URL, viewport, status, and error. A per-URL record makes partial success visible and allows failed pages to be retried.
  5. Use collision-safe filenames. A host-only name is readable, but a stable catalogue item ID or URL hash is safer when multiple paths share a host.

3. Choose viewport, element, or full-page capture

A driver screenshot captures the browser’s screenshot context; the exact extent depends on the browser and chosen method. Selenium’s documented API also supports an element screenshot. Do not infer that a regular driver screenshot always captures the full document.

Viewport screenshot

The main program uses a 1440 × 1000 browser window. Configure the dimensions before navigation or capture and keep them fixed across a batch. Headless and headed browser rendering can differ in available fonts, device scale, and environment, so use the same mode for captures that you intend to compare.

Element screenshot

Use a selector for the specific catalogue region and save that element’s image:

import org.openqa.selenium.By;
import org.openqa.selenium.OutputType;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.TakesScreenshot;
import java.nio.file.Files;
import java.nio.file.Path;

WebElement card = driver.findElement(By.cssSelector(".product-card"));
Files.copy(card.getScreenshotAs(OutputType.FILE).toPath(), Path.of("product-card.png"));

Element screenshot output can be affected by clipping, overflow, and whether the element is in view. Wait until the element is present and visible, and ensure the selector identifies the intended instance.

Full-page screenshot

For a long page, use a browser-specific full-page feature, a verified DevTools route, or a scrolling and stitching method that you have validated with the exact browser and Selenium versions. A secondary approach expands the browser window to document.documentElement.scrollWidth and scrollHeight and captures the expanded viewport. That changes the viewport and can trigger responsive breakpoints, so it is not a universal full-page API. Very tall pages can also produce large images and high memory use.

FirefoxDriver versions may expose a full-page screenshot method, and Chrome can be driven through browser-specific DevTools functionality; check the installed versions and validate on representative pages before relying on either. The Selenium API material does not promise identical full-document behavior across browsers for ordinary screenshot calls.

For catalogue pages with below-the-fold images, scroll the page in measured increments to trigger lazy loading before capture, then wait for the image elements you require to load. Scrolling can affect sticky headers and animations; decide whether those visual effects belong in the capture. Do not assume the unnamed target uses lazy loading or any particular selector.

4. Add waiting, retries, and failure isolation

Fixed sleeps are simple but waste time on fast pages and may still be too short on slow ones. Prefer explicit waits for conditions tied to the page. For example, if a known catalogue container should appear:

new WebDriverWait(driver, Duration.ofSeconds(30))
    .until(d -> d.findElement(By.cssSelector("main .catalogue-grid")).isDisplayed());

Use a selector that actually exists on your target. For a batch with unreliable pages, retry only transient failures, with a small bounded attempt count and a delay between attempts. Do not retry indefinitely or immediately hammer a site after repeated timeouts. Keep each URL’s failure separate from successful captures, and capture the exception class and URL in the report.

A shared WebDriver should be used sequentially. If you need parallel workers, create a separate driver session and distinct output path per worker; measure resource use and failure rate before increasing concurrency. The available sources do not establish a safe worker count or throughput figure.

5. Or skip the browser setup

ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. Its clean-shot processing accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers to show the outcome.

For this illustrative catalogue URL, use a real target URL you are authorized to capture. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/catalogue -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/catalogue"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/catalogue' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace the placeholder key and URL with your credentials and target. The API also supports full-page capture, element selection, device presets, PDF, custom CSS and JavaScript, click and wait actions, request blocking, headers and cookies, caching, async jobs, bulk capture of up to 100 URLs per call, and more. Its MCP tools include take_screenshot, get_page_info, and capture_pdf. ScreenshotNeo is available at screenshotneo.com. One thousand screenshots per month are free without a card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

6. cURL, Python, and Node.js alternatives for running Selenium

Selenium WebDriver is controlled through a language binding, so cURL does not take a local Selenium screenshot by itself. cURL can call Selenium Grid’s WebDriver HTTP interface when a remote session is already running. The exact session endpoint and capabilities depend on your Grid version and configuration; this minimal illustration creates a Chrome session, navigates, and asks the remote driver to return a screenshot as Base64:

curl -X POST http://localhost:4444/session \
  -H 'Content-Type: application/json' \
  -d '{"capabilities":{"alwaysMatch":{"browserName":"chrome"}}}'

Read the returned session ID, then replace SESSION_ID in the version-appropriate session endpoints below. WebDriver protocol routes can differ by implementation/version; consult your Grid’s documentation.

curl -X POST http://localhost:4444/session/SESSION_ID/url \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com/catalogue"}'

curl http://localhost:4444/session/SESSION_ID/screenshot

The screenshot response is JSON containing a Base64 image value; decode that value to a file. The Java client is usually more convenient for a Java application because it handles session details and file output.

Python Selenium follows the same per-URL loop. Install the binding with python -m pip install selenium; Selenium Manager can resolve standard drivers in supported setups.

from pathlib import Path
from urllib.parse import urlparse
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait

urls = [line.strip() for line in Path("urls.txt").read_text().splitlines()
        if line.strip() and not line.lstrip().startswith("#")]
out = Path("screenshots")
out.mkdir(exist_ok=True)
opts = Options()
opts.add_argument("--headless=new")
opts.add_argument("--window-size=1440,1000")
driver = webdriver.Chrome(options=opts)
try:
    for i, url in enumerate(urls, 1):
        try:
            driver.get(url)
            WebDriverWait(driver, 30).until(
                lambda d: d.execute_script("return document.readyState") == "complete")
            host = urlparse(url).hostname or "page"
            name = f"{i:04d}-{host.replace('.', '-')}.png"
            driver.save_screenshot(str(out / name))
            print("OK", url, name)
        except Exception as exc:
            print("ERROR", url, type(exc).__name__, str(exc))
finally:
    driver.quit()

Node.js Selenium uses the JavaScript WebDriver package. Install it with npm install selenium-webdriver and configure a matching browser/driver environment:

const fs = require('node:fs/promises');
const { Builder, By, until } = require('selenium-webdriver');
const chrome = require('selenium-webdriver/chrome');

async function main() {
  const urls = (await fs.readFile('urls.txt', 'utf8'))
    .split(/\r?\n/).map(s => s.trim()).filter(s => s && !s.startsWith('#'));
  await fs.mkdir('screenshots', { recursive: true });
  const options = new chrome.Options().addArguments('--headless=new', '--window-size=1440,1000');
  const driver = await new Builder().forBrowser('chrome').setChromeOptions(options).build();
  try {
    for (let i = 0; i < urls.length; i++) {
      try {
        await driver.get(urls[i]);
        await driver.wait(async d => await d.executeScript('return document.readyState') === 'complete', 30000);
        const png = await driver.takeScreenshot();
        await fs.writeFile(`screenshots/${String(i + 1).padStart(4, '0')}.png`, png, 'base64');
        console.log('OK', urls[i]);
      } catch (error) {
        console.error('ERROR', urls[i], error.name, error.message);
      }
    }
  } finally {
    await driver.quit();
  }
}
main().catch(error => { console.error(error); process.exitCode = 1; });

7. Troubleshooting

Symptom Likely cause Fix
Driver or browser session fails to start Browser and driver mismatch, missing browser binary, or restricted environment Check installed browser version and Selenium setup; configure an explicit browser binary or driver when automatic management cannot access the environment.
Screenshot is blank or the wrong page Navigation failed, redirect is still in progress, the target returned an error page, or the page renders after document readiness Record the final URL and page state, wait for a page-specific signal, and inspect the saved outcome before retrying.
Capture is only the visible viewport Ordinary screenshot extent is browser/method dependent Use a verified full-page technique for the selected browser and validate it on a long representative page.
Images or product data are missing Lazy loading or client-side rendering has not completed Wait for relevant elements and, where appropriate, scroll through the page before capture; avoid assuming one fixed delay works for all URLs.
Element not found or not visible Selector does not match, element is delayed, or the page structure differs Inspect the actual page structure, use a target-specific selector, and wait for presence/visibility before capture.
Some URLs fail and the rest stop Exception escapes the loop or a navigation stalls Catch failures per URL, set a page-load timeout, save a result row for every item, and restart a stuck browser session if needed.
Files overwrite one another Names are based only on a repeated host or generic label Include a stable input ID or index and a sanitized slug or hash in each filename.
Full-page output has altered layout Window expansion changed responsive breakpoints Use a browser-native full-page route or capture under the intended viewport; compare the output to the page as viewed at normal dimensions.
Batch is slow or consumes too much memory Large pages, high-resolution screenshots, excessive parallel browsers, or sessions retained too long Save and release each result promptly, keep concurrency conservative, and measure resource use before scaling.

8. Performance, reliability, and cost

Sequential navigation is a good starting point because it is easy to trace and keeps browser resource use straightforward. Each URL still incurs page load time plus capture and file I/O, and there is no reliable throughput figure without measuring the actual pages and environment. Parallel runs need isolated WebDriver sessions and collision-safe outputs; additional sessions consume browser and machine resources and may increase target-site load.

For reliability, use bounded timeouts, page-specific readiness checks, per-URL logs, and bounded retries only for transient errors. Reuse a driver for a sequential list when it remains healthy, but recreate it if a browser session becomes unusable. Preserve the input URL and capture conditions with each output so later comparison is meaningful.

Selenium’s direct software cost is not a screenshot API charge, but the job still uses compute, browser runtime, storage, and maintenance. If a managed API is preferable, ScreenshotNeo offers 1,000 free screenshots per month with no card; plans are $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000, with two months free on yearly billing. Every feature is on every plan. Its response headers indicate page verdict and billing status, including no charge for bot checks, blank pages, failed loads, timeouts, or cache hits.

9. Other Java options when requirements change

If an existing Java Selenium application is already the right fit, its driver and element screenshot primitives are the direct route. Percy’s Selenium integration adds a visual-testing workflow and documents an optional fullPage setting, with that behavior tied to Percy CLI 1.27.6 or later; check current terms and availability before adopting it. If changing browser automation frameworks is acceptable, Playwright Java documents full-page capture through setFullPage(true) and locator screenshots. These are different operational choices: Percy adds a visual-testing service, while Playwright replaces Selenium for the capture code.

10. Frequently asked questions

Does Selenium provide a bulk screenshot command?

The cited API material documents screenshot primitives, not a separate bulk command. The application loops over URLs and manages each output and failure.

Can the browser keep me logged in for all URLs?

Yes, if the URLs are captured in a session where you have established the required authentication state and are authorized to access the pages. Keep credentials out of source code and logs.

Will the same screenshot look identical on every run?

Not necessarily. Dynamic content, fonts, animations, viewport, browser version, locale, device scale, and page state can change pixels. Fix those conditions when doing visual comparison.

Is this guide specific to a particular Indian catalogue?

No. The title’s catalogue context does not identify a site. Select and validate the URL list, selectors, access rules, and pacing for the actual catalogue you use.