Selenium 4 WebDriver Commands: A Practical Guide
A practical Selenium 4 Python guide to starting sessions, navigating, locating elements, waiting for dynamic pages, switching contexts, and capturing evidence.
Selenium 4 WebDriver commands control a browser through a session. A reliable workflow is to configure browser options, start a session, navigate, locate and operate on elements, wait for the application state the next action needs, switch to the right tab or frame when necessary, capture evidence, and always end the session.
The examples below use the Selenium Python binding documented as version 4.50.0. Selenium method names and available features differ by language binding and release. Install the binding with python -m pip install selenium, and use a Chrome installation that your environment can launch. Selenium Manager can obtain a driver in supported situations, but setup behavior depends on the installed browser, network access, and environment.
1. Start a browser session and configure options
A WebDriver session is the browser context in which commands run. In Selenium 4, configure the browser with its Options class. For remote sessions, pass an options instance that specifies the browser; older Desired Capabilities examples are not the Selenium 4 setup pattern.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
options = Options()
# Uncomment to run without a visible browser window:
# options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
print("Session started")
finally:
driver.quit()
Headless mode is an environment choice, not a guarantee that rendering will match every headed browser setup. Browser options can also configure arguments, binary locations, profiles, proxies, and other browser-specific settings; consult the binding and browser documentation for supported values.
Navigation can wait for different document readiness states. Selenium’s default normal strategy waits for readyState complete; eager waits for interactive, while none does not block on page-load readiness. These strategies do not guarantee that a single-page application has rendered the specific element your test needs. If you choose an earlier return, synchronize explicitly with the next required condition. See Selenium’s Browser Options documentation.
2. Navigate and inspect the page
Use get() to open a URL. The Python API documents it as waiting for the page load/onload event in the current tab. Use history navigation and refresh when the scenario calls for them. Inspect the current URL and title to check navigation; page source is a diagnostic snapshot, not a replacement for locating and interacting with live elements.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
options = Options()
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com")
print("Title:", driver.title)
print("URL:", driver.current_url)
# Diagnostic snapshot of the current document:
print(driver.page_source[:500])
driver.back()
driver.forward()
driver.refresh()
finally:
driver.quit()
A successful get() only tells you the configured readiness policy was met. JavaScript can still add or change application content afterward, so wait for the state that matters to your test.
3. Find elements and interact with them
find_element returns one match and raises an exception if none is found. find_elements returns a list, which may be empty. Selenium supports locators such as ID, name, CSS selector, XPath, class name, tag name, and link text. Choose a stable locator tied to application semantics; avoid relying on fragile positional selectors when a meaningful attribute is available.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
options = Options()
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com")
heading = driver.find_element(By.TAG_NAME, "h1")
print("Heading:", heading.text)
print("Visible:", heading.is_displayed())
links = driver.find_elements(By.CSS_SELECTOR, "a")
print("Link count:", len(links))
if links:
print("First link:", links[0].get_attribute("href"))
finally:
driver.quit()
Common element operations include click(), clear(), send_keys(), text, get_attribute(), is_displayed(), and is_enabled(). Check the state needed before acting when the page can update asynchronously.
from selenium.webdriver.common.by import By
email = driver.find_element(By.NAME, "email")
if email.is_displayed() and email.is_enabled():
email.clear()
email.send_keys("developer@example.com")
driver.find_element(By.CSS_SELECTOR, "button[type='submit']").click()
For the complete interaction model and locator guidance, see Selenium’s Web Elements and Browser Interactions documentation.
4. Wait for dynamic application state
Race conditions happen when the test issues a command before the application reaches the state the command requires. Selenium describes this as a primary cause of flaky tests. A document readiness event does not ensure a dynamically rendered button is visible or clickable.
| Approach | Scope and behavior | Good fit |
|---|---|---|
| Explicit wait | Waits for a named condition; returns when it becomes true or times out. | Dynamic pages, near the operation that depends on a state. |
| Implicit wait | Session-wide timeout applied to element-location calls; documented default is zero. | A deliberately chosen global lookup policy. |
| Fixed sleep | Always waits for the specified duration without checking readiness. | A real fixed delay that is itself part of the behavior being tested. |
Explicit waits are usually clearer for dynamic interfaces because they express the condition required by the next step. Avoid casually mixing implicit and explicit waits: Selenium warns that the combination can produce unpredictable wait durations. Selenium’s Waiting Strategies guide explains the tradeoffs.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
options = Options()
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com")
wait = WebDriverWait(driver, 10)
# Wait until the element exists and is visible before reading it.
banner = wait.until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "h1"))
)
print(banner.text)
# For an action, wait for clickability.
# submit = wait.until(EC.element_to_be_clickable(
# (By.CSS_SELECTOR, "button[type='submit']")
# ))
# submit.click()
finally:
driver.quit()
Other useful conditions include presence, invisibility, staleness of an old element, title or URL changes, and frame availability. Pick the condition that represents the state your next command needs. If a lookup still times out, verify the locator and whether the element is inside a frame or shadow root before increasing the timeout.
5. Switch tabs and windows
WebDriver commands operate in the currently selected window. Track window handles, identify the new handle by comparing before and after sets, and switch explicitly. Do not assume a particular handle ordering.
from selenium.webdriver.support.ui import WebDriverWait
before = set(driver.window_handles)
# Trigger the site action that opens a tab or window:
driver.find_element(By.LINK_TEXT, "Open details").click()
WebDriverWait(driver, 10).until(
lambda d: len(set(d.window_handles) - before) == 1
)
new_handle = (set(driver.window_handles) - before).pop()
driver.switch_to.window(new_handle)
print(driver.current_url)
# Return to the original window when needed:
original_handle = next(iter(before))
driver.switch_to.window(original_handle)
The site may open a new window only after a click or another user gesture, and a delayed popup can require a wait. See Selenium’s Windows and Tabs guide.
6. Work with frames and JavaScript dialogs
Locate a frame and switch into it before finding its content. Switch back to the top-level document with default_content(), or to the containing frame with parent_frame().
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
wait = WebDriverWait(driver, 10)
frame = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "iframe.payment")))
driver.switch_to.frame(frame)
try:
driver.find_element(By.NAME, "cardnumber").send_keys("4111111111111111")
finally:
driver.switch_to.default_content()
JavaScript alerts, prompts, and confirmations must be handled as dialogs before continuing with page commands that depend on them being gone.
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
alert = WebDriverWait(driver, 10).until(EC.alert_is_present())
print(alert.text)
# Choose the action appropriate to the test:
alert.accept() # or alert.dismiss()
# For a prompt, set text before accepting:
# alert.send_keys("sample input")
# alert.accept()
Consult Selenium’s documentation for frames and JavaScript alerts.
7. Capture evidence and end the session
Capture a screenshot when a failure occurs, and preserve the test name and error alongside it. A screenshot taken after the page has moved on may not show the state that caused the failure. The Python API supports saving a PNG file and capturing screenshot bytes. Window size and rectangle methods can help diagnose viewport-dependent behavior.
from pathlib import Path
# Save a full viewport screenshot:
driver.save_screenshot("failure.png")
# Or get PNG bytes for a test report or artifact store:
png_bytes = driver.get_screenshot_as_png()
Path("artifacts").mkdir(exist_ok=True)
Path("artifacts/page.png").write_bytes(png_bytes)
print("Window size:", driver.get_window_size())
print("Window rectangle:", driver.get_window_rect())
close() closes the current window. quit() ends the whole WebDriver session and should normally be used during teardown. Put cleanup in finally or your test framework’s teardown hook so failed assertions do not leave a browser or remote session running.
8. Complete workflow example
This runnable example puts session setup, navigation, explicit waits, interaction, evidence capture, and guaranteed teardown together. Replace the example URL and selectors with ones from the application under test.
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
options = Options()
# Uncomment for a headless environment:
# options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
driver.set_window_size(1365, 900)
driver.get("https://example.com")
wait = WebDriverWait(driver, 10)
heading = wait.until(
EC.visibility_of_element_located((By.TAG_NAME, "h1"))
)
print("Page:", driver.title, driver.current_url)
print("Heading:", heading.text)
# Example form flow; adapt selectors to the target page.
# field = wait.until(EC.element_to_be_clickable((By.NAME, "query")))
# field.clear()
# field.send_keys("Selenium")
# driver.find_element(By.CSS_SELECTOR, "button[type='submit']").click()
# wait.until(EC.url_contains("search"))
except Exception:
Path("artifacts").mkdir(exist_ok=True)
driver.save_screenshot("artifacts/failure.png")
raise
finally:
driver.quit()
Or skip the browser setup
For a screenshot asset rather than an interactive browser test, [ScreenshotNeo](https://screenshotneo.com) offers a website screenshot API and MCP server. One GET request returns an image or PDF; the API accepts many parameter names used by other screenshot APIs. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo removes cookie banners, popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free account and get 1,000 screenshots a month with no card.
Troubleshooting Selenium 4 commands
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Driver or browser fails to start | Browser missing, incompatible environment, unavailable driver download, or incorrect browser binary. | Confirm the browser is installed and runnable, check network/proxy restrictions and browser options, and inspect the startup exception. Selenium Manager behavior depends on the environment. |
NoSuchElementException |
Wrong locator, element not yet rendered, or element is in another frame. | Check the live page and selector, wait for the required condition, and switch into the relevant frame. |
TimeoutException |
The condition did not become true within the timeout, or the locator/state is wrong. | Check the expected state and selector, confirm the page URL and browsing context, and set a timeout appropriate to the application. Do not mask a broken condition with a large global wait. |
| Element is not interactable or click is intercepted | Element is hidden, disabled, covered, or not yet ready for interaction. | Wait for visibility or clickability, inspect overlays and element state, and ensure the test is in the correct frame. |
| Stale element reference | The page replaced or refreshed the element after it was located. | Wait for the new state and locate the element again instead of reusing the old WebElement. |
| Commands target the wrong page | The active window or frame is not the intended one. | Check window handles, switch to the target handle, or switch into/out of the frame explicitly. |
| Dialog blocks later commands | A JavaScript alert, prompt, or confirmation remains open. | Wait for the alert and accept or dismiss it before sending further page commands. |
| Test passes locally but fails in CI | Different browser setup, viewport, timing, permissions, or display environment. | Record browser and binding versions, set a deliberate viewport, use condition-based waits, and preserve screenshots and logs at failure. |
Performance, reliability, and cost considerations
- Wait for conditions: explicit waits can return as soon as the needed state is true. A fixed sleep always consumes its full duration and still may be too short.
- Choose page-load policy deliberately:
eagerornonecan reduce navigation blocking, while transferring more synchronization responsibility to the test. Verify application state explicitly afterward. - Keep sessions bounded: one session per test or a controlled test fixture makes cleanup and failure diagnosis easier. Always quit even on exceptions.
- Capture useful evidence: save artifacts at the point of failure and include test context. Screenshot capture helps diagnosis but does not itself establish correctness.
- Account for infrastructure: local and remote browser sessions consume machine or grid capacity and run time. Selenium’s cited documentation does not establish a universal cost or performance figure; actual cost depends on your infrastructure and test design.
Advanced note: WebDriver BiDi
The Selenium 4.50.0 Python API includes BiDi-related interfaces for areas such as browsing contexts, input, network, and scripts. These APIs extend beyond the classic command flow, and availability and syntax differ across bindings and releases. Verify the exact API in the documentation for your chosen binding before adopting version-specific examples. The Python API reference is at Selenium WebDriver Python API.
FAQ
How do I wait for an element in Selenium?
Use WebDriverWait with a condition such as visibility or clickability when the next step depends on that state. Avoid combining implicit and explicit waits casually.
How do I switch tabs or frames in Selenium WebDriver?
Switch to the intended window handle with driver.switch_to.window(handle). For frames, switch with driver.switch_to.frame(...) and return with default_content() or parent_frame().
Does driver.get() mean the whole app is ready?
No. It waits according to the page-load strategy and document readiness; JavaScript-driven content may still be loading or changing afterward.
Should I use close() or quit()?
Use close() to close the current window when needed. Use quit() to end the WebDriver session during teardown.


