ScreenshotNeo

BlogGuides

Selenium 4 WebDriver Commands: A Practical Guide

A practical Selenium 4 Python guide to starting sessions, navigating, locating elements, waiting for dynamic pages, switching contexts, and capturing evidence.

By the ScreenshotNeo team4 October 202611 min read

Selenium 4 WebDriver commands control a browser through a session. A reliable workflow is to configure browser options, start a session, navigate, locate and operate on elements, wait for the application state the next action needs, switch to the right tab or frame when necessary, capture evidence, and always end the session.

The examples below use the Selenium Python binding documented as version 4.50.0. Selenium method names and available features differ by language binding and release. Install the binding with python -m pip install selenium, and use a Chrome installation that your environment can launch. Selenium Manager can obtain a driver in supported situations, but setup behavior depends on the installed browser, network access, and environment.

1. Start a browser session and configure options

A WebDriver session is the browser context in which commands run. In Selenium 4, configure the browser with its Options class. For remote sessions, pass an options instance that specifies the browser; older Desired Capabilities examples are not the Selenium 4 setup pattern.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
# Uncomment to run without a visible browser window:
# options.add_argument("--headless")

driver = webdriver.Chrome(options=options)
try:
    print("Session started")
finally:
    driver.quit()

Headless mode is an environment choice, not a guarantee that rendering will match every headed browser setup. Browser options can also configure arguments, binary locations, profiles, proxies, and other browser-specific settings; consult the binding and browser documentation for supported values.

Navigation can wait for different document readiness states. Selenium’s default normal strategy waits for readyState complete; eager waits for interactive, while none does not block on page-load readiness. These strategies do not guarantee that a single-page application has rendered the specific element your test needs. If you choose an earlier return, synchronize explicitly with the next required condition. See Selenium’s Browser Options documentation.

2. Navigate and inspect the page

Use get() to open a URL. The Python API documents it as waiting for the page load/onload event in the current tab. Use history navigation and refresh when the scenario calls for them. Inspect the current URL and title to check navigation; page source is a diagnostic snapshot, not a replacement for locating and interacting with live elements.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")
    print("Title:", driver.title)
    print("URL:", driver.current_url)
    # Diagnostic snapshot of the current document:
    print(driver.page_source[:500])

    driver.back()
    driver.forward()
    driver.refresh()
finally:
    driver.quit()

A successful get() only tells you the configured readiness policy was met. JavaScript can still add or change application content afterward, so wait for the state that matters to your test.

3. Find elements and interact with them

find_element returns one match and raises an exception if none is found. find_elements returns a list, which may be empty. Selenium supports locators such as ID, name, CSS selector, XPath, class name, tag name, and link text. Choose a stable locator tied to application semantics; avoid relying on fragile positional selectors when a meaningful attribute is available.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By

options = Options()
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")

    heading = driver.find_element(By.TAG_NAME, "h1")
    print("Heading:", heading.text)
    print("Visible:", heading.is_displayed())

    links = driver.find_elements(By.CSS_SELECTOR, "a")
    print("Link count:", len(links))
    if links:
        print("First link:", links[0].get_attribute("href"))
finally:
    driver.quit()

Common element operations include click(), clear(), send_keys(), text, get_attribute(), is_displayed(), and is_enabled(). Check the state needed before acting when the page can update asynchronously.

from selenium.webdriver.common.by import By

email = driver.find_element(By.NAME, "email")
if email.is_displayed() and email.is_enabled():
    email.clear()
    email.send_keys("developer@example.com")

driver.find_element(By.CSS_SELECTOR, "button[type='submit']").click()

For the complete interaction model and locator guidance, see Selenium’s Web Elements and Browser Interactions documentation.

4. Wait for dynamic application state

Race conditions happen when the test issues a command before the application reaches the state the command requires. Selenium describes this as a primary cause of flaky tests. A document readiness event does not ensure a dynamically rendered button is visible or clickable.

Approach Scope and behavior Good fit
Explicit wait Waits for a named condition; returns when it becomes true or times out. Dynamic pages, near the operation that depends on a state.
Implicit wait Session-wide timeout applied to element-location calls; documented default is zero. A deliberately chosen global lookup policy.
Fixed sleep Always waits for the specified duration without checking readiness. A real fixed delay that is itself part of the behavior being tested.

Explicit waits are usually clearer for dynamic interfaces because they express the condition required by the next step. Avoid casually mixing implicit and explicit waits: Selenium warns that the combination can produce unpredictable wait durations. Selenium’s Waiting Strategies guide explains the tradeoffs.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = Options()
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")
    wait = WebDriverWait(driver, 10)

    # Wait until the element exists and is visible before reading it.
    banner = wait.until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "h1"))
    )
    print(banner.text)

    # For an action, wait for clickability.
    # submit = wait.until(EC.element_to_be_clickable(
    #     (By.CSS_SELECTOR, "button[type='submit']")
    # ))
    # submit.click()
finally:
    driver.quit()

Other useful conditions include presence, invisibility, staleness of an old element, title or URL changes, and frame availability. Pick the condition that represents the state your next command needs. If a lookup still times out, verify the locator and whether the element is inside a frame or shadow root before increasing the timeout.

5. Switch tabs and windows

WebDriver commands operate in the currently selected window. Track window handles, identify the new handle by comparing before and after sets, and switch explicitly. Do not assume a particular handle ordering.

from selenium.webdriver.support.ui import WebDriverWait

before = set(driver.window_handles)
# Trigger the site action that opens a tab or window:
driver.find_element(By.LINK_TEXT, "Open details").click()

WebDriverWait(driver, 10).until(
    lambda d: len(set(d.window_handles) - before) == 1
)
new_handle = (set(driver.window_handles) - before).pop()
driver.switch_to.window(new_handle)
print(driver.current_url)

# Return to the original window when needed:
original_handle = next(iter(before))
driver.switch_to.window(original_handle)

The site may open a new window only after a click or another user gesture, and a delayed popup can require a wait. See Selenium’s Windows and Tabs guide.

6. Work with frames and JavaScript dialogs

Locate a frame and switch into it before finding its content. Switch back to the top-level document with default_content(), or to the containing frame with parent_frame().

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

wait = WebDriverWait(driver, 10)
frame = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "iframe.payment")))
driver.switch_to.frame(frame)
try:
    driver.find_element(By.NAME, "cardnumber").send_keys("4111111111111111")
finally:
    driver.switch_to.default_content()

JavaScript alerts, prompts, and confirmations must be handled as dialogs before continuing with page commands that depend on them being gone.

from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

alert = WebDriverWait(driver, 10).until(EC.alert_is_present())
print(alert.text)
# Choose the action appropriate to the test:
alert.accept()  # or alert.dismiss()

# For a prompt, set text before accepting:
# alert.send_keys("sample input")
# alert.accept()

Consult Selenium’s documentation for frames and JavaScript alerts.

7. Capture evidence and end the session

Capture a screenshot when a failure occurs, and preserve the test name and error alongside it. A screenshot taken after the page has moved on may not show the state that caused the failure. The Python API supports saving a PNG file and capturing screenshot bytes. Window size and rectangle methods can help diagnose viewport-dependent behavior.

from pathlib import Path

# Save a full viewport screenshot:
driver.save_screenshot("failure.png")

# Or get PNG bytes for a test report or artifact store:
png_bytes = driver.get_screenshot_as_png()
Path("artifacts").mkdir(exist_ok=True)
Path("artifacts/page.png").write_bytes(png_bytes)

print("Window size:", driver.get_window_size())
print("Window rectangle:", driver.get_window_rect())

close() closes the current window. quit() ends the whole WebDriver session and should normally be used during teardown. Put cleanup in finally or your test framework’s teardown hook so failed assertions do not leave a browser or remote session running.

8. Complete workflow example

This runnable example puts session setup, navigation, explicit waits, interaction, evidence capture, and guaranteed teardown together. Replace the example URL and selectors with ones from the application under test.

from pathlib import Path

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = Options()
# Uncomment for a headless environment:
# options.add_argument("--headless")

driver = webdriver.Chrome(options=options)
try:
    driver.set_window_size(1365, 900)
    driver.get("https://example.com")
    wait = WebDriverWait(driver, 10)

    heading = wait.until(
        EC.visibility_of_element_located((By.TAG_NAME, "h1"))
    )
    print("Page:", driver.title, driver.current_url)
    print("Heading:", heading.text)

    # Example form flow; adapt selectors to the target page.
    # field = wait.until(EC.element_to_be_clickable((By.NAME, "query")))
    # field.clear()
    # field.send_keys("Selenium")
    # driver.find_element(By.CSS_SELECTOR, "button[type='submit']").click()
    # wait.until(EC.url_contains("search"))

except Exception:
    Path("artifacts").mkdir(exist_ok=True)
    driver.save_screenshot("artifacts/failure.png")
    raise
finally:
    driver.quit()

Or skip the browser setup

For a screenshot asset rather than an interactive browser test, [ScreenshotNeo](https://screenshotneo.com) offers a website screenshot API and MCP server. One GET request returns an image or PDF; the API accepts many parameter names used by other screenshot APIs. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo removes cookie banners, popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free account and get 1,000 screenshots a month with no card.

Troubleshooting Selenium 4 commands

Symptom Likely cause What to check or change
Driver or browser fails to start Browser missing, incompatible environment, unavailable driver download, or incorrect browser binary. Confirm the browser is installed and runnable, check network/proxy restrictions and browser options, and inspect the startup exception. Selenium Manager behavior depends on the environment.
NoSuchElementException Wrong locator, element not yet rendered, or element is in another frame. Check the live page and selector, wait for the required condition, and switch into the relevant frame.
TimeoutException The condition did not become true within the timeout, or the locator/state is wrong. Check the expected state and selector, confirm the page URL and browsing context, and set a timeout appropriate to the application. Do not mask a broken condition with a large global wait.
Element is not interactable or click is intercepted Element is hidden, disabled, covered, or not yet ready for interaction. Wait for visibility or clickability, inspect overlays and element state, and ensure the test is in the correct frame.
Stale element reference The page replaced or refreshed the element after it was located. Wait for the new state and locate the element again instead of reusing the old WebElement.
Commands target the wrong page The active window or frame is not the intended one. Check window handles, switch to the target handle, or switch into/out of the frame explicitly.
Dialog blocks later commands A JavaScript alert, prompt, or confirmation remains open. Wait for the alert and accept or dismiss it before sending further page commands.
Test passes locally but fails in CI Different browser setup, viewport, timing, permissions, or display environment. Record browser and binding versions, set a deliberate viewport, use condition-based waits, and preserve screenshots and logs at failure.

Performance, reliability, and cost considerations

  • Wait for conditions: explicit waits can return as soon as the needed state is true. A fixed sleep always consumes its full duration and still may be too short.
  • Choose page-load policy deliberately: eager or none can reduce navigation blocking, while transferring more synchronization responsibility to the test. Verify application state explicitly afterward.
  • Keep sessions bounded: one session per test or a controlled test fixture makes cleanup and failure diagnosis easier. Always quit even on exceptions.
  • Capture useful evidence: save artifacts at the point of failure and include test context. Screenshot capture helps diagnosis but does not itself establish correctness.
  • Account for infrastructure: local and remote browser sessions consume machine or grid capacity and run time. Selenium’s cited documentation does not establish a universal cost or performance figure; actual cost depends on your infrastructure and test design.

Advanced note: WebDriver BiDi

The Selenium 4.50.0 Python API includes BiDi-related interfaces for areas such as browsing contexts, input, network, and scripts. These APIs extend beyond the classic command flow, and availability and syntax differ across bindings and releases. Verify the exact API in the documentation for your chosen binding before adopting version-specific examples. The Python API reference is at Selenium WebDriver Python API.

FAQ

How do I wait for an element in Selenium?

Use WebDriverWait with a condition such as visibility or clickability when the next step depends on that state. Avoid combining implicit and explicit waits casually.

How do I switch tabs or frames in Selenium WebDriver?

Switch to the intended window handle with driver.switch_to.window(handle). For frames, switch with driver.switch_to.frame(...) and return with default_content() or parent_frame().

Does driver.get() mean the whole app is ready?

No. It waits according to the page-load strategy and document readiness; JavaScript-driven content may still be loading or changing afterward.

Should I use close() or quit()?

Use close() to close the current window when needed. Use quit() to end the WebDriver session during teardown.