Selenium WebDriver Tutorial with Examples
Write your first Selenium WebDriver script in Python: set up a browser session, interact with a page, wait for dynamic content, and clean up reliably.
Selenium WebDriver lets a script control a browser through a language binding and a browser-specific driver. This tutorial uses Python to open a page, find and fill a field, click a button, wait for the result, inspect it, and close the browser session.
The example uses Selenium’s public web form at https://www.selenium.dev/selenium/web/web-form.html. It demonstrates a local Chrome session. Selenium can also drive browsers through Selenium Server, including remote browser sessions. See the Selenium WebDriver overview.
1. How WebDriver works
Your script calls Selenium’s Python binding. The binding sends WebDriver commands to a driver associated with the browser, and that driver controls the browser. Selenium describes WebDriver as driving browsers natively, locally or through Selenium Server. The WebDriver specification is a W3C Recommendation.
Selenium’s WebDriver overview also describes WebDriver BiDi, which uses a WebSocket connection for browser events such as network requests and console messages. The example below uses ordinary WebDriver commands and does not require BiDi.
2. Set up Python and a browser
- Install Python and a browser supported by the Selenium binding you plan to use. This example uses Chrome.
- Install the Selenium Python package in your project environment:
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install selenium
Selenium Manager, used by Selenium bindings by default, helps manage browser drivers and browsers for ordinary setups. It does not guarantee that every network, permissions, browser-version, or compatibility problem will resolve automatically. If session creation fails, see the troubleshooting section.
For a repeatable project setup, record the Selenium package version in your dependency lock or requirements file. Install the browser in each environment where the script runs. Remote execution requires a Selenium Server or compatible remote WebDriver endpoint and different session configuration; it is not required for this local example. Read Selenium’s getting started guide and Selenium Manager documentation for current setup details.
3. Write and run the first script
Save this as first_selenium.py. It creates a session, navigates, reads the title, locates elements using name, CSS selector, and ID, enters text, clicks, waits for the result, checks it, and quits the session even if an exception occurs.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URL = "https://www.selenium.dev/selenium/web/web-form.html"
# Selenium Manager may manage the driver for this local browser session.
driver = webdriver.Chrome()
try:
driver.get(URL)
print("Title:", driver.title)
print("Current URL:", driver.current_url)
wait = WebDriverWait(driver, 10)
# Locate by name, then enter text.
text_input = wait.until(
EC.visibility_of_element_located((By.NAME, "my-text"))
)
text_input.send_keys("Selenium WebDriver")
# Locate the submit button by CSS selector and click it.
submit_button = wait.until(
EC.element_to_be_clickable((By.CSS_SELECTOR, "button"))
)
submit_button.click()
# Wait for the page's result, then inspect it.
message = wait.until(
EC.visibility_of_element_located((By.ID, "message"))
)
print("Result:", message.text)
print("Result URL:", driver.current_url)
finally:
# Close the browser session, including when an earlier step raises.
driver.quit()
Run it from the activated environment with python first_selenium.py. The sample page and locators follow Selenium’s official first-script material. Treat this as a tutorial example: adapt the locators and expected result to your own application.
4. Find elements with locators
A locator describes how to find an element; it does not guarantee that the element is visible, enabled, or ready for the next action. Prefer stable attributes intended for automation, such as an application-owned ID or a dedicated test attribute. Avoid depending on long, layout-specific XPath expressions or CSS chains that change whenever the page structure changes.
| Locator | Python form | Useful when |
|---|---|---|
| ID | (By.ID, "message") |
The page has a unique, stable ID. |
| Name | (By.NAME, "my-text") |
Finding form fields by their name attribute. |
| CSS selector | (By.CSS_SELECTOR, "button[type='submit']") |
Selecting by a stable CSS attribute or relationship. |
| XPath | (By.XPATH, "//button[@type='submit']") |
A relationship or text-based lookup is needed and remains stable. |
Use find_element when one match is expected; it raises an exception if no match is found. Use find_elements when zero or more matches are valid; it returns a list, which may be empty. If the page can render asynchronously, put the lookup inside a suitable explicit wait.
5. Wait for the page state you need
A navigation returning does not mean a dynamic application has finished rendering the particular control your next command needs. Wait for that condition directly. Selenium’s waiting strategies guide describes implicit waits, explicit waits, and the risks of combining them.
Explicit waits (recommended for a specific condition)
WebDriverWait(driver, 10).until(condition) polls until the condition succeeds or the timeout is reached. Common conditions include element presence, visibility, clickability, a URL change, and an alert being present. Choose a condition that matches the next action: presence means the node exists; visibility means it is displayed; clickability checks visibility and enabled state.
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
wait = WebDriverWait(driver, 10)
button = wait.until(
EC.element_to_be_clickable((By.CSS_SELECTOR, "button[type='submit']"))
)
button.click()
The timeout is a maximum, not a fixed sleep: the wait returns as soon as the condition succeeds. If it expires, inspect the locator, page state, and whether the element is in a frame or a different window.
Implicit waits
An implicit wait sets a global delay for element lookups. For example, driver.implicitly_wait(5) asks Selenium to keep retrying lookups for up to five seconds before failing. It can be convenient for simple scripts, but it does not wait for a specific state such as visibility or clickability.
Do not casually mix implicit and explicit waits. Their combined timing can be unpredictable because an explicit wait repeatedly performs lookups that themselves may wait implicitly. For condition-specific synchronization, leave the implicit wait unset and use explicit waits.
Page-load strategy is separate
| Strategy | Navigation waits for | Use with care |
|---|---|---|
normal |
The page load event. | Default behavior for a full navigation. |
eager |
DOMContentLoaded. |
Returns earlier while subresources may still load. |
none |
Only the initial document download. | Returns early; your script must synchronize before interacting. |
These settings control when navigation returns. They do not establish that an application-specific element is ready. Keep an explicit condition wait for that. See Selenium’s driver options documentation.
6. Common browser interactions
Read navigation state
driver.get("https://example.com")
print(driver.title)
print(driver.current_url)
Use the returned title or URL as information for your script, or assert an expected value in a test. A title alone may not prove that a dynamic application has reached the desired state.
Handle an alert
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
alert = WebDriverWait(driver, 10).until(EC.alert_is_present())
print(alert.text)
alert.accept() # Or use alert.dismiss() when the dialog should be canceled.
Wait for the alert before interacting with it. Browser alerts are not ordinary page elements and cannot be found with a CSS selector.
Read and set cookies
driver.get("https://example.com")
# A cookie can be added only after navigating to its domain.
driver.add_cookie({"name": "example", "value": "value"})
print(driver.get_cookies())
driver.delete_cookie("example")
Cookie rules are enforced by the browser, including domain and security restrictions. Navigate to the appropriate site before adding a cookie.
Switch to an iframe
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
wait = WebDriverWait(driver, 10)
wait.until(EC.frame_to_be_available_and_switch_to_it((By.ID, "payment-frame")))
try:
driver.find_element(By.NAME, "card-number").send_keys("example")
finally:
driver.switch_to.default_content()
Elements inside a frame are not found from the top-level document. Switch into the frame first, then return to default content when finished.
Switch tabs or windows
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
original = driver.current_window_handle
# Perform the action that opens a new tab or window here.
WebDriverWait(driver, 10).until(EC.number_of_windows_to_be(2))
new_handle = next(handle for handle in driver.window_handles if handle != original)
driver.switch_to.window(new_handle)
print(driver.current_url)
driver.close() # Closes the current tab.
driver.switch_to.window(original)
close() closes the current window. quit() ends the WebDriver session and closes its associated windows; use it for orderly cleanup.
7. Local and remote sessions
A local session is simplest for learning and runs the browser where the script runs. For a remote session, Selenium communicates with a Selenium Server or Grid endpoint that creates the browser session. Remote execution is useful when browsers run on separate machines or in a managed test environment, but it adds server availability, network, and capability configuration to diagnose.
The Python shape is webdriver.Remote(command_executor=server_url, options=options), where server_url is the endpoint supplied by your Selenium Server or Grid and options is the browser options object. Endpoint URLs and supported capabilities depend on the server deployment; do not copy an arbitrary endpoint from an unrelated environment. Always call quit() so the remote session is released.
8. Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Driver or browser cannot be found | Browser is missing, Selenium Manager cannot download or resolve a driver, or the environment blocks access. | Confirm the browser is installed and launchable. Check network and permissions, then consult Selenium Manager and driver setup documentation. |
| Session creation reports a version or compatibility error | Browser and driver versions do not work together, or the installed browser differs from the expected one. | Confirm which browser is being launched and review the browser and driver compatibility requirements. Allow Selenium Manager to manage the ordinary setup or configure a compatible driver deliberately. |
NoSuchElementException |
The locator is wrong, the element has not appeared, or it is inside a frame. | Inspect the current page and locator. Use an explicit wait for presence or visibility. Switch to the right frame when needed. |
TimeoutException |
The expected condition never became true before the timeout. | Check the locator, network and application state, frame/window context, and whether the expected state actually occurs. Increase the timeout only when the application legitimately needs more time. |
ElementClickInterceptedException |
An overlay or another element covers the target, or it is not yet in a usable state. | Wait for the overlay to disappear and for the target to be clickable. Check whether the page has a consent dialog or animation. |
StaleElementReferenceException |
The page re-rendered and replaced the element after it was located. | Locate the element again after the update, then wait for the new element’s needed condition. |
| Script ends but browser processes remain | Cleanup did not run after an exception or the session was not quit. | Put work in a try/finally block and call driver.quit() in finally. |
| Element is found but input or click has no effect | It may be hidden, disabled, covered, outside the expected frame, or replaced during interaction. | Wait for visibility or clickability, confirm frame and window context, and inspect the page state before retrying. |
9. Reliability, performance, and cost
- Use condition waits: Wait for the exact state needed, not a guessed delay. Fixed sleeps make fast runs slower and slow runs flaky.
- Keep locators stable: Prefer application-owned IDs or test attributes and avoid selectors tied to incidental layout.
- Always clean up: A browser session consumes local or remote resources until it is closed. Use
quit()in afinallyblock, including after test failures. - Expect environment costs: Selenium itself is an automation framework, but browser automation still uses compute, storage, and network resources. Remote browser infrastructure may have its own provider costs; Selenium’s setup pages do not establish a universal price.
- Keep timing deliberate: A shorter page-load strategy can return earlier, but it does not make an application ready. Add the explicit waits that protect the next interaction.
10. Or skip the browser setup
If your goal is a page image or PDF rather than interactive browser automation, ScreenshotNeo provides a website screenshot API and MCP server. It does not replace Selenium for workflows that need to type, click through an application, or inspect interactive state. For a capture, one GET request can return an image or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.
11. FAQ
Is Selenium WebDriver only for testing?
No. It is a browser automation API that can be used by test scripts and other browser automation programs. Choose it when you need to control browser interactions.
Should I use Selenium for a static screenshot?
Selenium can open a browser, but it is designed for browser control. If all you need is a screenshot or PDF from a URL, a screenshot API can avoid setting up and managing a browser session.
Why does driver.get() return before my app is ready?
Navigation completion is based on a page-load strategy, while applications can render or fetch content afterward. Wait for the element or application condition your next action needs.
What should I use if an element appears more than once?
Use a more specific stable locator when possible. If multiple matches are expected, use find_elements and choose the intended match based on a meaningful attribute or relationship.


