ScreenshotNeo

BlogGuides

Selenium WebDriver: A Beginner’s Guide

Learn what Selenium WebDriver does, install it, write a first browser automation script, wait for dynamic elements, and troubleshoot common setup issues.

By the ScreenshotNeo team4 October 20269 min read

Selenium WebDriver lets a program control a real browser: open pages, find elements, click, type, and inspect results. To write your first script, install a Selenium language binding, have a supported browser available, create a driver session, navigate to a page, interact with or inspect an element, and call quit to close the session. In current Selenium versions, Selenium Manager usually handles browser-driver setup for you.

This guide uses Python for the main runnable example and includes equivalent starting points for JavaScript, Java, and cURL context. WebDriver itself is a browser automation interface, so cURL is not a replacement for it: cURL makes HTTP requests but does not drive a rendered browser.

1. What Selenium WebDriver is

Selenium WebDriver is a language-neutral interface and protocol for controlling browsers from an external program. Selenium provides language bindings, while browser-specific driver implementations connect those commands to browsers. The W3C describes WebDriver as a remote-control interface for user agents; the current WebDriver 2 document cited here is a Working Draft published on 2 July 2026, so its status should not be confused with a finalized Recommendation. Selenium WebDriver documentation · W3C WebDriver draft.

A typical script follows this lifecycle:

  1. Create a browser driver and start a session.
  2. Navigate to a URL.
  3. Locate elements using a locator such as an ID or CSS selector.
  4. Interact with the page or check its state.
  5. End the session with quit, including when an error occurs.

WebDriver can start a browser on the same machine as the script or connect to remote Selenium infrastructure. A local browser is the simplest place to begin.

2. Install Selenium and prepare a browser

Setup has three conceptual pieces: a Selenium binding for your programming language, a browser, and a driver implementation for that browser. Install a current Selenium binding and make sure the browser you intend to automate is installed. For ordinary recent local setups, Selenium Manager, built into Selenium, can resolve and manage the browser driver automatically. You may still need explicit configuration in locked-down environments, with custom browser builds, or when connecting to remote infrastructure. Selenium getting started.

Python installation

python -m pip install selenium

Save the example below as first_selenium.py, then run python first_selenium.py. It uses the official Selenium documentation site as its target and prints the page title.

from selenium import webdriver


def main():
    driver = webdriver.Chrome()
    try:
        driver.get("https://www.selenium.dev/")
        print(driver.title)
    finally:
        driver.quit()


if __name__ == "__main__":
    main()

The finally block closes the browser even if navigation or inspection raises an exception. Selenium’s Python API documentation describes the current binding and its driver usage. Selenium Python API.

JavaScript installation

The Selenium JavaScript API documentation currently specifies Node.js 22 or later. Check that API page for current supported Node release lines before setting up a project. Install the package with npm:

npm init -y
npm install selenium-webdriver

Save as first.js and run node first.js:

const { Builder } = require('selenium-webdriver');

(async () => {
  const driver = await new Builder().forBrowser('chrome').build();
  try {
    await driver.get('https://www.selenium.dev/');
    console.log(await driver.getTitle());
  } finally {
    await driver.quit();
  }
})();

Selenium JavaScript API.

Java setup outline

For Java, add Selenium’s Java binding to your build using the dependency instructions in Selenium’s getting-started guide, then create a ChromeDriver, navigate, and call quit in a finally block. Use the dependency version and build-tool coordinates shown in the current official documentation rather than copying an unmaintained version from an older tutorial. Selenium also provides bindings for C#, Ruby, and Kotlin. Official installation guide.

3. Write a useful first script

Printing the title confirms that a session started and navigation completed. A useful next step is locating a stable element and interacting with it. Prefer IDs or deliberate CSS selectors over fragile positional selectors or text that may change. The target below is illustrative: use a page whose markup and behavior you control when adapting it.

from selenium import webdriver
from selenium.webdriver.common.by import By


driver = webdriver.Chrome()
try:
    driver.get("https://example.com")
    heading = driver.find_element(By.CSS_SELECTOR, "h1")
    print(heading.text)
finally:
    driver.quit()

find_element returns one matching element and raises an exception if none is found. Use find_elements when zero or more matches are valid and you want a list. Keep session cleanup even in small scripts: an unclosed browser can leave processes and sessions running.

4. Wait for dynamic pages correctly

A page’s initial navigation finishing does not guarantee that every asynchronous widget, search result, or client-rendered element is ready. Wait for the condition the next action depends on. Avoid making fixed sleeps the default: they either waste time on fast runs or fail on slow ones.

Python example using an explicit wait for an element to become clickable:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


driver = webdriver.Chrome()
try:
    driver.get("https://example.com")
    button = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.CSS_SELECTOR, "button.submit"))
    )
    button.click()
finally:
    driver.quit()

Choose a condition that reflects what must be true: presence, visibility, or clickability are different states. If the page displays a loading indicator, waiting for it to disappear may be more reliable than waiting an arbitrary duration. Check your language binding’s current API for the exact wait and expected-condition syntax.

5. Browser drivers, local sessions, and remote execution

Selenium Manager automates much of driver acquisition in current bindings and is used by default. This means many beginners do not need to download ChromeDriver manually. ChromeDriver remains a separate executable maintained by the Chromium team with WebDriver contributors, and direct configuration is still useful when an environment requires a pinned or custom executable. Consult the official ChromeDriver guide when configuring it directly.

  • Local WebDriver: runs a browser on the machine executing the script; simplest for learning and small jobs.
  • Remote WebDriver: sends commands to a remote endpoint, useful when browsers run elsewhere or are centrally managed.
  • Selenium Grid: distributes browser sessions across machines, useful as a later scaling step for parallel runs. It is not required for a first local script. See the Selenium project documentation.

When using remote execution, the driver builder and endpoint configuration depend on your infrastructure. Keep credentials out of source code and use the configuration documented for the remote Selenium server you connect to.

6. Selenium WebDriver, Selenium IDE, and WebDriver BiDi

Tool or mode What it is for When to choose it
WebDriver Code-based browser control through language bindings. When you need reusable scripts, assertions, and control over automation logic.
Selenium IDE Record-and-playback, low-code browser automation. When you want an approachable way to record a flow or explore automation before writing code.
Grid / remote WebDriver Runs sessions through remote or distributed infrastructure. When local execution is no longer sufficient for your environment or run volume.
WebDriver BiDi A bidirectional WebSocket protocol for streaming and reacting to browser events. Advanced cases involving events such as network requests, console messages, or JavaScript errors, where the browser and binding support the needed capability.

Selenium’s documentation describes BiDi as a W3C protocol developed with browser vendors. Support can vary by browser and binding, so check the current support information for the exact capability you need. It is not required for a basic WebDriver session. WebDriver documentation.

7. cURL and browser screenshots

cURL is useful for making HTTP requests, but it does not start a browser, execute page JavaScript as a browser does, or interact with rendered controls. If your goal is browser automation, use WebDriver code such as the examples above. If your goal is only a rendered screenshot, a screenshot API is a simpler fit than building and maintaining a browser session.

For a raw HTTP request unrelated to browser rendering, a basic cURL command looks like this:

curl -L "https://example.com"

That returns the server’s HTTP response body; it is not a screenshot and does not provide Selenium’s browser interactions.

8. Troubleshooting common problems

Symptom Likely cause What to do
Driver or browser cannot be found Browser is missing, Selenium Manager cannot obtain a driver in the environment, or a custom setup is required. Install the target browser, check network and filesystem restrictions, and consult Selenium Manager or the browser-driver configuration guide for your binding.
Session creation fails with a version or compatibility message The browser and driver are incompatible, or a manually pinned driver is stale. Use Selenium Manager for a standard setup, or align the explicit driver with the installed browser using the browser vendor’s current instructions.
Element cannot be located The locator is wrong, the page differs from expectations, or the element has not appeared yet. Inspect the page and locator, verify the active frame or window, and wait for the required condition before locating or interacting.
Click intercepted or element not interactable An overlay covers the target, it is hidden, or the page has not reached the needed state. Wait for overlays to clear and for the element to become visible or clickable. Confirm the selector identifies the intended control.
Script exits but browser remains open The session did not reach cleanup, often because quit was omitted or not protected by cleanup logic. Call quit in a finally block and avoid terminating the process abruptly.
Works locally but fails on a remote machine The remote browser environment, permissions, network, or available browser differs from the local setup. Check the remote endpoint configuration and browser availability there; local driver assumptions do not automatically apply remotely.

9. Performance, reliability, and cost

Selenium is open-source browser automation software, and the core first-run example does not require buying a Selenium product. Browser startup and page loading take time, so reuse a session for a sequence of related actions when that is appropriate, and always clean it up. Explicit waits help a script continue when a meaningful condition is met rather than adding unnecessary fixed delays. For parallel execution across machines, Selenium Grid is the project’s scaling path; no speed or reliability benchmark is implied here. See the official project overview.

For screenshots, a browser automation script means you manage the browser session and its setup. ScreenshotNeo is a purpose-built website screenshot API and MCP server by Yorker Media. It returns PNG, JPEG, WebP, or PDF from a GET request and includes features such as full-page capture, selector capture, device presets, custom CSS and JavaScript, waits, and caching. For the full parameter list and API setup, see ScreenshotNeo documentation.

10. Or skip the browser setup

If you only need a website screenshot, ScreenshotNeo makes it one API call. This example saves a WebP response for Stripe; replace the target URL with the page you need. Put your key in an environment or secret manager in a real application.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', bytes);

ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Learn more at ScreenshotNeo, read the API documentation, or sign up for 1,000 free screenshots a month with no card.

11. Frequently asked questions

Do I need Selenium Grid to learn WebDriver?

No. A local browser session is enough for a first script. Grid becomes relevant when you need remote or distributed execution.

Does WebDriver work with only one programming language?

No. Selenium provides language bindings including Python, Java, JavaScript, C#, Ruby, and Kotlin; the workflow is similar while the APIs differ.

Should I start with Selenium IDE or WebDriver?

Use IDE to explore a record-and-playback workflow with little code. Choose WebDriver when you want to own the automation logic in a programming language.

Do I need BiDi for ordinary browser tests?

No. Basic navigation, element interaction, and inspection use the ordinary WebDriver workflow. Consider BiDi only for event-driven capabilities supported by your browser and binding.