How to Take Bulk Screenshots of Web Pages with Selenium in Java
Capture a list of URLs with one Selenium WebDriver session, save each screenshot safely, handle per-page failures, and understand full-page limits.
To take bulk screenshots with Selenium in Java, load a list of URLs, reuse one WebDriver session, navigate to each URL, wait for the page condition you care about, save a screenshot under a unique filename, and record success or failure for every URL. Selenium’s standard driver screenshot captures the current browsing context; full-page capture is browser-specific, so choose and verify the browser and driver before relying on it.
The example below reads one URL per line from urls.txt, captures the current view with Chrome, writes PNG files and a CSV manifest, and continues when an individual URL fails. It requires Java 17 or later, Selenium Java, and a compatible Chrome browser and ChromeDriver available to Selenium. Selenium’s official Java example uses TakesScreenshot with OutputType.FILE and copies the resulting file to its destination. Selenium screenshot documentation.
1. Set up the Java project
For Maven, add Selenium Java to your pom.xml. Replace the version placeholder with the Selenium version your project has selected; keep the browser and driver versions compatible. This avoids claiming a version that may not match your environment.
<dependency>
<groupId>org.seleniumhq.selenium</groupId>
<artifactId>selenium-java</artifactId>
<version>YOUR_SELENIUM_VERSION</version>
</dependency>
Install Chrome and ensure Selenium can locate a compatible ChromeDriver, then create urls.txt with one absolute HTTP or HTTPS URL per line:
https://example.com/
https://www.selenium.dev/
https://developer.mozilla.org/
2. Run a bulk capture
Save this as BulkScreenshots.java. It accepts an optional input file and output directory. It sets the viewport, page-load timeout, and explicit wait; sanitizes URL-derived names and adds a digest to avoid collisions; and writes a manifest row for each valid input URL.
import java.io.BufferedWriter;
import java.io.IOException;
import java.net.URI;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.StandardCopyOption;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import java.time.Duration;
import java.util.HexFormat;
import java.util.List;
import java.util.Locale;
import org.openqa.selenium.By;
import org.openqa.selenium.OutputType;
import org.openqa.selenium.TakesScreenshot;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
public class BulkScreenshots {
private static final Duration PAGE_LOAD_TIMEOUT = Duration.ofSeconds(45);
private static final Duration CONDITION_TIMEOUT = Duration.ofSeconds(20);
public static void main(String[] args) throws Exception {
Path input = args.length > 0 ? Path.of(args[0]) : Path.of("urls.txt");
Path output = args.length > 1 ? Path.of(args[1]) : Path.of("screenshots");
Files.createDirectories(output);
List<String> lines = Files.readAllLines(input, StandardCharsets.UTF_8);
ChromeOptions options = new ChromeOptions();
// For a machine without a display, uncomment this line:
// options.addArguments("--headless=new");
options.addArguments("--window-size=1440,1000");
Path manifestPath = output.resolve("manifest.csv");
WebDriver driver = null;
try (BufferedWriter manifest = Files.newBufferedWriter(
manifestPath, StandardCharsets.UTF_8)) {
manifest.write("input_url,status,file,error");
manifest.newLine();
driver = new ChromeDriver(options);
driver.manage().timeouts().pageLoadTimeout(PAGE_LOAD_TIMEOUT);
WebDriverWait wait = new WebDriverWait(driver, CONDITION_TIMEOUT);
for (String raw : lines) {
String url = raw.trim();
if (url.isEmpty() || url.startsWith("#")) continue;
String status = "SUCCESS";
String filename = "";
String error = "";
try {
validateHttpUrl(url);
driver.get(url);
// Replace with a page-specific condition when possible.
wait.until(d -> !d.getTitle().isBlank());
filename = fileNameFor(url) + ".png";
Path temporary = Files.createTempFile(output, "capture-", ".png");
try {
Files.copy(((TakesScreenshot) driver)
.getScreenshotAs(OutputType.FILE).toPath(),
temporary, StandardCopyOption.REPLACE_EXISTING);
Files.move(temporary, output.resolve(filename),
StandardCopyOption.REPLACE_EXISTING);
} finally {
Files.deleteIfExists(temporary);
}
} catch (Exception e) {
status = "FAILED";
error = e.getClass().getSimpleName() + ": " + e.getMessage();
}
manifest.write(csv(url) + "," + status + "," + csv(filename)
+ "," + csv(error));
manifest.newLine();
manifest.flush();
}
} finally {
if (driver != null) driver.quit();
}
}
private static void validateHttpUrl(String value) {
URI uri = URI.create(value);
String scheme = uri.getScheme();
if (uri.getHost() == null || scheme == null
|| !(scheme.equalsIgnoreCase("http")
|| scheme.equalsIgnoreCase("https"))) {
throw new IllegalArgumentException("Expected an absolute HTTP(S) URL");
}
}
private static String fileNameFor(String url)
throws NoSuchAlgorithmException {
URI uri = URI.create(url);
String path = uri.getPath() == null ? "" : uri.getPath();
String base = path.isBlank() ? "page" : path.substring(path.lastIndexOf('/') + 1);
if (base.isBlank()) base = "page";
base = base.replaceAll("[^A-Za-z0-9._-]", "_");
if (base.length() > 48) base = base.substring(0, 48);
byte[] digest = MessageDigest.getInstance("SHA-256")
.digest(url.getBytes(StandardCharsets.UTF_8));
return base + "-" + HexFormat.of().formatHex(digest).substring(0, 12);
}
private static String csv(String value) {
return "\"" + value.replace("\"", "\"\"").replace("\n", " ")
.replace("\r", " ") + "\"";
}
}
Run it from the project directory:
mvn compile
mvn exec:java -Dexec.mainClass=BulkScreenshots -Dexec.args="urls.txt screenshots"
The last command assumes the Maven Exec Plugin is configured. Alternatively, run the class through your IDE or add that plugin to the project. If the title is not populated on a target site, the example’s readiness wait times out and marks that URL failed. Replace the title check with a condition meaningful to the application, such as visibility of a known content element:
wait.until(ExpectedConditions.visibilityOfElementLocated(
By.cssSelector("main article")));
3. Choose screenshot scope and readiness
Viewport or full page
TakesScreenshot captures the current browsing context. Do not assume the generic driver call produces a complete, consistent full-page image in every browser. Selenium’s RemoteWebDriver reference describes best-effort behavior for non-W3C-conformant drivers; W3C-conformant behavior follows the WebDriver specification. Check the documentation and behavior for the browser and driver you actually deploy. RemoteWebDriver Java API.
For Firefox, Selenium documents HasFullPageScreenshot and getFullPageScreenshotAs(OutputType). The API labels this capability beta. It is not a portable method to assume for every browser:
import org.openqa.selenium.OutputType;
import org.openqa.selenium.firefox.FirefoxDriver;
import org.openqa.selenium.firefox.HasFullPageScreenshot;
FirefoxDriver driver = new FirefoxDriver();
try {
driver.get("https://example.com/");
byte[] png = ((HasFullPageScreenshot) driver)
.getFullPageScreenshotAs(OutputType.BYTES);
Files.write(Path.of("full-page.png"), png);
} finally {
driver.quit();
}
Use a Firefox driver that implements the interface and confirm its compatibility in your chosen Selenium release. Selenium HasFullPageScreenshot Java API.
Wait for the page you need
- Prefer an explicit condition tied to the page, such as visibility of the main content or a known application-ready state.
- A navigation load event may happen before client-rendered content, images, or data requests have finished.
- A fixed sleep is simple but wastes time on fast pages and may still be too short on slow ones.
- For image-heavy pages, wait for the relevant images to load if that matters to the result. Define that condition for the site; there is no universal readiness signal.
- Set page-load and condition timeouts for your workload. The example’s values are starting points, not Selenium-wide recommendations.
4. Make the batch dependable
| Concern | Recommended handling |
|---|---|
| Input | Validate absolute HTTP(S) URLs and skip blank lines or comments before navigation. |
| File names | Use sanitized names plus a sequence number or stable URL hash; URL basenames alone can collide. |
| Failures | Catch errors per URL when later captures should continue, and record status and error in a manifest. |
| Cleanup | Always call quit() in a finally block so browser processes close after errors. |
| Reproducibility | Record browser, driver, viewport, capture time, and URL with the output when results need to be compared later. |
| Parallelism | Start serially. If needed, use bounded parallel workers with a separate browser session per worker and sufficient host capacity. There is no universal safe concurrency limit or throughput figure. |
The example uses one session and processes inputs in order. This reduces browser startup overhead compared with starting a browser for every URL. A single session can still accumulate state such as cookies, local storage, service workers, or dialogs between sites. If pages must be isolated, use separate sessions or clear relevant browser state between captures, accounting for the added startup cost.
5. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Session cannot start or driver error | ChromeDriver is missing or mismatched with Chrome. | Install compatible browser and driver versions and verify Selenium can locate the driver. |
| Timeout during navigation | The page is slow, stuck, or never reaches the configured page-load condition. | Set a suitable page-load timeout; consider Selenium’s page-load strategy for your application, then use an explicit readiness condition. Log and continue if the URL is allowed to fail. |
| Wait times out after navigation | The example’s title condition is unsuitable, or the page failed to render expected content. | Use a condition for a stable, visible element or application state and inspect the page failure separately. |
| Screenshot is blank or incomplete | The capture occurred before content rendered, the site returned an error/interstitial, or the requested full-page scope is unsupported. | Wait for the content you need, inspect the resulting page, and verify the selected browser’s screenshot behavior. |
| Only the visible area appears | The generic screenshot call captured the current browsing context viewport. | Use a supported browser-specific full-page capability, such as Selenium’s documented Firefox API, and verify its beta status and compatibility. |
| Files overwrite or are missing | Names collide, path is unwritable, or the process stopped before file output. | Use unique names, create the destination directory, check permissions and available disk space, and inspect the manifest. |
| Later URLs show logged-in or personalized content | Browser state persisted in the reused session. | Use isolated sessions or explicitly clear state according to the capture requirements. |
| Headless run differs from local desktop | Viewport, browser mode, fonts, GPU behavior, or environment differs. | Set viewport deliberately and compare captures in the same browser environment used for the run. |
6. Performance, reliability, and cost
Serial capture time is dominated by navigation and the readiness condition for each page; no source here establishes a throughput benchmark. Reuse one driver for a serial batch, avoid arbitrary long sleeps, and tune waits to the pages being captured. Parallelism can reduce elapsed time but consumes more memory and CPU and introduces more browser processes to manage. Use a measured, bounded worker count for your own environment.
Reliability comes from validating inputs, isolating per-URL failures, recording a manifest, choosing deterministic filenames, and closing sessions in all cases. For long jobs, flushing the manifest after each row helps preserve progress information if a later capture fails. Check output disk capacity for large batches and images.
Selenium’s direct software cost depends on where you run the browser and the compute, storage, and maintenance you provide. This workflow makes no claim about a fixed cost per screenshot. If the actual deliverable is printable archival rather than raster images, Selenium’s Java print-page functionality can produce PDF with options for orientation, page ranges, page size, margins, scale, backgrounds, and shrink-to-fit; confirm PDF satisfies the downstream need. Selenium print page documentation.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. A GET request captures a URL as PNG, JPEG, WebP, or PDF. See the ScreenshotNeo site and API documentation.
For a single capture, the Java workflow can be replaced with this cURL call (adapt the URL and output extension as needed):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also supports bulk capture of up to 100 URLs per call; consult the docs for request details. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture, and each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; all features are on every plan. Sign up for 1,000 free screenshots a month, with no card.
FAQ
Can I take bulk screenshots without opening a new browser for every URL?
Yes. The main example reuses one WebDriver session and navigates it through the URL list.
Does a Selenium screenshot include browser chrome?
The WebDriver screenshot API is for the current browsing context; it is not a general desktop screenshot of browser controls.
Can Selenium save screenshots as JPEG?
The sample requests a PNG file. If another image format is required, check the selected browser and driver support and convert the saved image with an image library.
Should I use PDF instead of a screenshot?
Use PDF when the consumer needs printable page output and the browser’s print representation is acceptable. Use raster captures when the consumer needs image pixels.


