Web Scraping with Selenium and Java: A Practical Guide
Learn a maintainable Selenium Java workflow for permitted web scraping: set up WebDriver, wait for dynamic content, extract and validate data, and handle common errors.
Selenium lets a Java program control a real browser, which can help collect data from pages that depend on browser interaction or client-side rendering. The basic workflow is to start a WebDriver session, navigate to a page you are permitted to access, wait for the specific content you need, locate and validate it, and close the session in a finally block.
Before collecting data, check the target site’s terms and applicable access rules. Some sites prohibit scraping, and some block Selenium. Do not bypass access controls, rate limits, authentication boundaries, or anti-bot protections. If a documented API or permitted export meets the need, consider that simpler route first. Selenium’s scraping guidance explains the project’s warning.
1. Add Selenium to a Java project
Use the official Selenium Java binding through your build tool. The Maven dependency below uses a version placeholder: select a current Selenium version from the official installation documentation instead of copying an old pinned version.
<dependencies>
<dependency>
<groupId>org.seleniumhq.selenium</groupId>
<artifactId>selenium-java</artifactId>
<version>YOUR_CURRENT_SELENIUM_VERSION</version>
</dependency>
</dependencies>
Selenium Manager is included with Selenium releases and can discover and manage missing browser drivers. Selenium documents it as available starting with 4.6; automated browser management is described as of 4.11.0. Consult the current Selenium Manager documentation for version-specific behavior.
2. Scrape a permitted page with explicit waits
This complete example targets Selenium’s sample page so you can see the mechanics without assuming a third-party site’s terms or DOM structure. It waits for the result field, reads it, validates that it is non-empty, and always closes the browser session. Replace the sample URL and locator only after checking that you are permitted to collect the target data.
import java.time.Duration;
import java.util.List;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
public class SeleniumScrape {
public static void main(String[] args) {
WebDriver driver = new ChromeDriver();
try {
driver.get("https://www.selenium.dev/selenium/web/web-form.html");
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10));
WebElement textBox = wait.until(
ExpectedConditions.visibilityOfElementLocated(By.name("my-text")));
textBox.sendKeys("Selenium Java");
wait.until(ExpectedConditions.elementToBeClickable(By.cssSelector("button"))).click();
WebElement result = wait.until(
ExpectedConditions.visibilityOfElementLocated(By.id("message")));
String value = result.getText().trim();
if (value.isEmpty()) {
throw new IllegalStateException("The result was empty");
}
System.out.println("Title: " + driver.getTitle());
System.out.println("Result: " + value);
} finally {
driver.quit();
}
}
}
The sample demonstrates browser navigation, locators, interaction, an explicit condition wait, text extraction, a basic sanity check, and cleanup. To collect a list of repeated records, locate the records container or repeated elements and iterate over them, validating required fields before saving each record. Do not assume a selector will remain stable if the page structure changes.
3. Choose locators that match the page
Selenium supports locator strategies including ID, name, CSS selector, class name, and link text. Use a clear, stable identifier when the page provides one, and verify that the locator selects the intended element.
| Locator | Java example | Useful when |
|---|---|---|
| ID | By.id("price") |
The element has a unique, stable ID. |
| Name | By.name("query") |
A form field has a meaningful name. |
| CSS selector | By.cssSelector("article h2") |
You can express the target through a concise DOM relationship or attribute. |
| Class name | By.className("product-title") |
A single class identifies the intended element. This locator accepts one class name. |
| Link text | By.linkText("Next") |
A link’s visible text is stable and specific. |
For multiple results, use findElements; it returns an empty list when there are no matches rather than throwing the missing-element exception used by findElement.
List<WebElement> cards = driver.findElements(By.cssSelector("article.product"));
for (WebElement card : cards) {
String title = card.findElement(By.cssSelector("h2")).getText().trim();
String href = card.findElement(By.cssSelector("a")).getAttribute("href");
if (!title.isEmpty() && href != null && !href.isBlank()) {
System.out.println(title + "\t" + href);
}
}
This selector is illustrative: change it to reflect the target page’s actual DOM. Prefer semantic attributes and specific relationships over selectors based on position, generated class names, or incidental nesting.
4. Wait for the content you need
A browser can finish navigation before a JavaScript application has rendered or updated the content you want. Fixed sleeps add avoidable delay when a page is fast and may still be too short when it is slow. Use an explicit wait for the state that makes extraction safe: an element becoming visible, a result count appearing, or a loading marker disappearing.
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15));
WebElement results = wait.until(
ExpectedConditions.visibilityOfElementLocated(By.cssSelector("main .results")));
String text = results.getText();
Selenium warns that mixing implicit and explicit waits can produce unpredictable timeout durations. Keep the implicit wait at its default (zero) when using explicit waits, and make each wait describe the page condition needed by the next step.
- Wait for visibility before reading text that is rendered asynchronously.
- Wait for clickability before interacting with a control that may still be loading or covered.
- Wait for staleness or a changed value when an action replaces existing results.
- Set a finite timeout and report which condition timed out so failures are diagnosable.
5. Extract, validate, and store data carefully
Read only the fields needed for the task, then validate them before writing output. For example, normalize whitespace, check required fields, and handle optional attributes explicitly. Keep extraction separate from storage so malformed or incomplete records can be logged or skipped according to your requirements.
String name = element.getText().trim();
String url = element.getAttribute("href");
if (name.isBlank() || url == null || url.isBlank()) {
throw new IllegalStateException("Record is missing a required name or URL");
}
If a page uses pagination, identify the site’s permitted next-page mechanism and place a clear stopping condition in the loop. Avoid unbounded loops: stop when there are no results, no next page, or a configured page limit is reached. Handle duplicate records if the source can return overlapping pages.
6. Run locally or use a remote browser
For a small job, a local browser is usually the simplest setup. Selenium WebDriver controls browser implementations through language bindings; Selenium also supports driving a browser on a remote machine through Selenium Server. Remote execution adds infrastructure and is useful when the browser needs to run in a separate environment or as part of a distributed setup. Choose based on where the browser must run and how you will manage sessions; the dossier provides no measured speed comparison.
7. Troubleshoot common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Driver or browser cannot start | Browser installation, driver discovery, or version setup is unavailable. | Confirm the browser is installed and consult current Selenium Manager guidance. Check Selenium and browser versions and the startup exception details. |
NoSuchElementException |
The locator is wrong, the page has not rendered the element, or the element is in a different browsing context. | Inspect the current DOM, verify the locator, wait for the required condition, and account for frames or shadow DOM where applicable. |
TimeoutException |
The expected page condition did not occur before the configured deadline. | Check whether the page loaded, whether the selector is correct, and whether the condition matches the actual UI state. Do not mask a persistent failure by increasing the timeout without diagnosis. |
| Stale element reference | The page replaced or refreshed the element after it was located. | Wait for the update to finish, then locate the element again instead of reusing the old reference. |
| Empty or partial values | Extraction ran before rendering completed, or the selector matched a wrapper or the wrong node. | Wait for the specific content, verify the matched node, and validate required fields before storing. |
| Browser session remains open after an exception | Cleanup was not protected against failure. | Put driver.quit() in a finally block so the session is closed on both success and error paths. |
| Access denied or CAPTCHA appears | The site may block automated access or require a different permitted access path. | Stop and check the site’s terms and access rules. Use an authorized API or request permission; do not bypass the control. |
8. Performance, reliability, and cost
A Selenium run starts and controls a browser, so it involves more setup and moving parts than a direct request. Use a browser when browser execution is actually needed, keep sessions scoped and closed, avoid unnecessary navigation, and wait on meaningful conditions rather than sleeping for a fixed interval. For repeated jobs, make retries limited and deliberate: retry transient navigation or infrastructure failures only when safe, and do not repeat actions that could submit forms or change site state without understanding their effects.
Reliability depends on the site’s DOM, rendering behavior, browser environment, and access policy. Log the URL, failed condition, and exception context without collecting unrelated sensitive page data. Browser automation has resource costs such as browser process time and machine capacity; Selenium itself is an open-source project, and this guide makes no claim about a site’s permission, scraping success rate, or runtime benchmark.
9. Or skip the browser setup
If the job is to capture a page as an image or PDF rather than extract structured records, ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. Its capture can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. AI agents can use its MCP tools, including take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for request options. This cURL example saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
The Node.js snippet uses Bun’s file writer to save the returned bytes. In a Node.js project without Bun, write the response buffer with the Node fs/promises API. ScreenshotNeo offers 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots, and every feature is available on every plan. A screenshot is a visual capture, not structured record extraction.
Sign up for 1,000 free screenshots a month, with no card required.
10. Frequently asked questions
Can Selenium scrape any public page?
No. Public visibility does not itself grant permission. Check the site’s terms and applicable rules, and respect blocks and access controls.
Does Selenium return structured data automatically?
No. WebDriver gives your program access to page elements; your code must locate the relevant fields, extract their text or attributes, validate the values, and store them.
When should I use a screenshot API instead?
Use one when the desired output is a visual image or PDF. For structured records, Selenium extraction or a permitted data API is the relevant approach.


