How to Use Selenium with Java for Browser Automation
Set up Selenium WebDriver in Java, automate a browser, choose reliable locators, wait for dynamic pages, and scale execution when your tests grow.
Selenium WebDriver lets Java programs control a browser through its WebDriver implementation. To get started, add Selenium to a Maven or Gradle project, create a WebDriver, navigate to a page, find elements with locators, interact with them, wait for the page state your next action needs, and call quit() to close the session.
This guide answers “How do I use Selenium with Java for browser automation?” with a runnable first script, locator and wait guidance, test organization, remote execution options, troubleshooting, and practical notes on reliability and cost.
1. Understand Selenium WebDriver
WebDriver is an API and protocol for controlling browsers. Selenium describes it as a W3C Recommendation. The Java binding sends commands to a browser’s WebDriver implementation. The browser can run on the same machine as the Java program or, through Selenium Server, on another machine. Selenium WebDriver documentation
The basic parts are:
- Your Java program: Uses Selenium’s Java bindings to issue browser commands.
- The browser: Chrome, Firefox, or another supported browser must be installed or otherwise available in the environment.
- The browser driver: Implements WebDriver commands for that browser. Selenium Manager, which Selenium bindings use by default for automated browser and driver management, can handle the basic driver setup in many local cases.
Selenium Manager does not install a browser for you. A browser still needs to be present and usable in the environment where the session starts. See Selenium’s Selenium Manager documentation and installation instructions.
2. Add Selenium to a Java project
For Maven, add the selenium-java artifact. Check Selenium’s downloads page for the current release instead of pinning a version copied from an older tutorial. The example below uses a Maven property so you can set the verified version in one place.
<properties>
<maven.compiler.release>17</maven.compiler.release>
<selenium.version>REPLACE_WITH_CURRENT_SELENIUM_VERSION</selenium.version>
</properties>
<dependencies>
<dependency>
<groupId>org.seleniumhq.selenium</groupId>
<artifactId>selenium-java</artifactId>
<version>${selenium.version}</version>
</dependency>
</dependencies>
Replace the version placeholder with a release listed by Selenium. The Java release shown is an example project setting; confirm the currently supported Java baseline in Selenium’s documentation before choosing a project baseline. Selenium’s install-library guide also shows Gradle setup.
3. Run your first Selenium Java script
Save this as FirstScript.java in a Maven project’s src/main/java directory. Run it with your IDE or compile and execute it using your project’s Maven configuration.
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;
public class FirstScript {
public static void main(String[] args) {
WebDriver driver = new ChromeDriver();
try {
driver.get("https://www.selenium.dev/selenium/web/web-form.html");
WebElement textBox = driver.findElement(By.name("my-text"));
WebElement submitButton = driver.findElement(By.cssSelector("button"));
textBox.sendKeys("Selenium");
submitButton.click();
String message = driver.findElement(By.id("message")).getText();
System.out.println(message);
} finally {
driver.quit();
}
}
}
The example follows Selenium’s Java first-script flow. Selenium Manager may arrange the driver automatically. The browser itself must still be available.
| Code | Purpose |
|---|---|
new ChromeDriver() |
Starts a Chrome WebDriver session. |
driver.get(url) |
Navigates the browser to a URL. |
By.name, By.id, By.cssSelector |
Describe how Selenium should locate an element. |
sendKeys, click, getText |
Enter text, activate a control, and read page content. |
driver.quit() |
Ends the session and closes its browser windows. |
Put quit() in a finally block (or your test framework’s teardown hook). If a script exits early after an exception, cleanup still runs and avoids leaving browser processes behind.
4. Choose stable element locators
Selenium supports locator strategies including ID, name, class name, CSS selector, XPath, link text, and partial link text. The locator documentation lists the available strategies.
| Locator | Example | Use when |
|---|---|---|
| ID | By.id("email") |
The page has a stable, unique ID for the target. |
| Name | By.name("my-text") |
A form control has a stable name attribute. |
| CSS selector | By.cssSelector("form button[type='submit']") |
A concise DOM relationship or attribute selector describes the target. |
| Link text | By.linkText("Continue") |
The visible link text is stable and unambiguous. |
| XPath | By.xpath("//button[@type='submit']") |
The required relationship or condition is awkward to express with CSS. |
Prefer a stable ID or name when the application provides one. Use a CSS selector when it clearly expresses a stable relationship. Avoid selectors that depend on a changing class, a long chain of ancestors, or “the third button” when the page structure may change. A clear locator is easier to review when a test fails.
findElement returns one matching element and throws if none is found. findElements returns a list, which can be empty; use it when absence is a valid result you need to check rather than an exceptional condition.
5. Wait for the page state your next action needs
A navigation reaching its load readiness state does not guarantee that JavaScript-driven content has appeared, become visible, or become clickable. Selenium’s waiting-strategies guide describes this timing problem and explains explicit waits. Selenium waiting strategies
Use a condition tied to the action you are about to perform. For example, after triggering an update, wait until its result is visible:
import java.time.Duration;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10));
WebElement result = wait.until(
ExpectedConditions.visibilityOfElementLocated(By.id("result")));
System.out.println(result.getText());
The ten-second duration is an example, not a universal setting. Choose a timeout based on the application’s expected behavior and the cost of a slow failure. Other useful conditions include element presence, element to be clickable, title matching text, and invisibility of a loading indicator. Wait for the state needed by the next command, not an arbitrary pause.
A fixed delay such as Thread.sleep(2000) always waits for the full duration and can still be too short on a slow run. An explicit wait polls for a specific condition and can continue as soon as that condition is true.
Selenium warns that mixing implicit and explicit waits can lead to unpredictable timeout behavior. Keep the implicit wait at its default (zero) when using explicit waits, and make synchronization decisions with explicit conditions. Waiting strategies
6. Turn browser flows into repeatable tests
A one-off script is useful for learning or a small task. For application checks, put repeatable flows in the test framework already used by the project, assert observable outcomes, and keep browser setup and teardown predictable.
- Keep test setup and browser cleanup in framework hooks so failures do not leak sessions.
- Assert a meaningful outcome, such as a confirmation message or changed page state, rather than only checking that a click did not throw.
- Centralize repeated page operations and selectors in page-specific helper classes or methods where that makes tests easier to maintain.
- Keep tests independent where possible so they can be rerun and distributed without relying on another test’s browser state.
- Capture useful failure context in your test runner, such as the failing step and browser logs where available.
These are test-design practices; Selenium supplies the browser-control API, while your project’s test framework supplies test discovery, assertions, and reporting.
7. Choose local or remote execution
Local execution is the simplest starting point: Java, Selenium, and a browser run on one machine. When you need execution across browser and operating-system combinations or across multiple machines, Selenium Grid is Selenium’s documented path for distributed execution. Selenium Grid documentation
| Approach | Setup and control | Coverage and scale |
|---|---|---|
| Local browser | Least infrastructure for a personal development machine or a small CI worker; direct control of that environment. | Limited to browsers and environments available on that machine. |
| Self-hosted Selenium Grid | You operate the Selenium Server/Grid infrastructure and its browser nodes. | Designed to distribute execution across machines, browsers, and operating systems. |
| Managed remote browser service | A provider operates remote browser infrastructure; evaluate its setup, control, and operating terms for your project. | Can reduce infrastructure work, but browser/platform coverage and costs depend on the chosen provider. |
The appropriate point to move beyond local runs depends on required browser coverage, environment reproducibility, and execution capacity. The research sources establish Selenium Grid as the distributed option; they do not substantiate comparisons or performance claims about specific commercial providers.
8. Troubleshoot common Selenium Java errors
| Symptom | Likely cause | What to check or change |
|---|---|---|
SessionNotCreatedException when creating a driver |
Browser startup failed, the browser is unavailable, or browser and driver setup is incompatible. | Confirm the browser is installed and can launch in the current environment. Review the exception details and Selenium Manager output; update Selenium or correct the browser/driver setup. |
NoSuchElementException |
The locator does not match, the element is in a different browsing context, or it has not appeared yet. | Check the locator against the current DOM, wait for the required state, and switch to the relevant frame or window if the element is there. |
TimeoutException from an explicit wait |
The expected condition never became true before the timeout. | Verify the application reached the expected state, the locator is correct, and the condition describes what the page actually does. Increase the timeout only if the behavior legitimately needs more time. |
StaleElementReferenceException |
The page replaced or refreshed the element after it was located. | Locate the element again after the page update. Wait for the new state before interacting; do not retain a reference across a known rerender. |
| Click intercepted or element not interactable | An overlay, animation, hidden state, or layout change prevents the intended interaction. | Wait for the element to be clickable and for blocking overlays to disappear. Check whether the control is actually visible and enabled. |
| Browser window remains after the script | The flow failed before cleanup or only closed one window. | Call driver.quit() in a finally block or teardown hook; quit() ends the session. |
| Tests pass locally but fail in CI | The CI browser, permissions, display environment, or timing differs from the local setup. | Check browser availability and startup logs in CI, make required environment settings explicit, and wait for application conditions rather than relying on timing assumptions. |
9. Performance, reliability, and cost
Performance
Browser startup, navigation, and page behavior contribute to end-to-end runtime. Avoid redundant navigations and unnecessary fixed sleeps. Wait for the specific state needed to proceed so a fast page does not always pay the full fixed delay. There are no benchmark figures in the cited research, so measure your own suite under its real browser and CI environment.
Reliability
Keep locators stable, waits condition-based, and session cleanup unconditional. Reproduce failures with the same browser and environment where possible. If execution must cover multiple browsers or operating systems, plan for that matrix explicitly; Selenium Grid can distribute sessions across environments.
Cost
For local runs, account for the machines and CI capacity your team operates. For Grid, account for the infrastructure and maintenance you choose to run. For a managed service, check its current pricing and terms directly; no provider-specific prices are established by this guide’s research.
10. Or skip the browser setup
If your goal is a screenshot rather than an interactive browser test, ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. It can return a PNG, JPEG, WebP, or PDF from one GET request. See the ScreenshotNeo API documentation for parameters and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie and consent banners are accepted like a visitor would accept them, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response includes X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.
11. Frequently asked questions
Do I need to download ChromeDriver manually?
For a basic Selenium setup, Selenium Manager is used by the bindings for automated browser and driver management. You still need an available browser, and specialized environments may need explicit driver configuration.
Can Selenium run without opening a visible browser window?
Browser visibility depends on the browser and execution environment configuration. Choose and document the mode appropriate to your local or CI environment, and verify it with the browser setup you use.
Should I use Selenium for screenshots?
Use Selenium when the task requires browser interaction or application testing. For a screenshot-only workflow, a screenshot API can avoid maintaining browser automation code; ScreenshotNeo provides a one-request capture API and an MCP server.
Where should I look for Selenium updates?
Check Selenium’s official documentation and downloads page for current installation guidance, release versions, and supported setup details.


