ScreenshotNeo

BlogGuides

Selenium with Python: A Practical Guide to Browser Automation

Install Selenium, automate a browser, find elements, wait for dynamic pages, and choose between local and remote execution with practical Python examples.

By the ScreenshotNeo team4 October 202610 min read

Selenium with Python lets you control a real browser from a Python script. Install the Selenium package, create a WebDriver session such as webdriver.Chrome(), navigate with driver.get(), locate elements, and use explicit waits when pages render asynchronously. Selenium Manager usually handles a missing browser driver for common setups. A local script does not need a Java server; use Selenium Grid and Remote WebDriver when you need a browser on another machine.

1. Install Selenium and prepare Python

The current Selenium Python bindings require Python 3.10 or later. Use a virtual environment so the project’s dependencies stay separate from other Python projects.

python --version
python -m venv .venv

# macOS or Linux
source .venv/bin/activate

# Windows PowerShell
.venv\Scripts\Activate.ps1

python -m pip install -U selenium

Check the Selenium Python API documentation for current requirements and supported APIs. If your machine has multiple Python installations, use the same interpreter to create the environment, install the package, and run the script.

2. Open and close a local browser

Save this as quickstart.py and run python quickstart.py. Chrome must be available in an environment Selenium can use. On common setups, Selenium Manager discovers the browser version and resolves, downloads, and caches a matching driver when one is missing.

from selenium import webdriver


driver = webdriver.Chrome()
try:
    driver.get("https://example.com")
    print(driver.title)
finally:
    driver.quit()

driver.quit() closes the whole browser session. Put it in a finally block so cleanup happens even if navigation or later work raises an exception. driver.close() closes the current window; it is not a substitute for quitting the session when the script is finished.

For local scripts, the Selenium project says, “For local Selenium scripts, the Java server is not needed.” Java is not a prerequisite for this quickstart.

3. Find elements and interact with a page

A locator tells Selenium which page element to operate on. Prefer a locator that identifies the intended element clearly and remains meaningful as the page changes. Use find_element when you expect one match and find_elements when you want a list; the latter returns an empty list when nothing matches.

from selenium import webdriver
from selenium.webdriver.common.by import By


driver = webdriver.Chrome()
try:
    driver.get("https://example.com")

    heading = driver.find_element(By.TAG_NAME, "h1")
    print(heading.text)

    links = driver.find_elements(By.CSS_SELECTOR, "a")
    for link in links:
        print(link.text, link.get_attribute("href"))
finally:
    driver.quit()

Common locator strategies include:

Strategy Example Use when
ID By.ID, "submit" The element has a stable, unique ID.
Name By.NAME, "email" A form element has a suitable name attribute.
CSS selector By.CSS_SELECTOR, "form button[type='submit']" The page structure offers a clear CSS selector.
XPath By.XPATH, "//button[normalize-space()='Continue']" You need to locate by text or a relationship in the document tree.
Tag name By.TAG_NAME, "h1" The element type is enough to identify the target.

These are examples, not a universal ranking: inspect the page and choose a locator that points to the intended element without accidentally matching unrelated elements. Selenium’s locator strategies guide covers the supported approaches.

4. Wait for dynamic content

Navigation does not guarantee that an application has finished rendering the element your next action needs. Use an explicit wait for that condition instead of guessing with a fixed sleep. The timeout is an upper bound for waiting for the condition; it does not promise that the page will load within that duration.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


driver = webdriver.Chrome()
try:
    driver.get("https://example.com")

    wait = WebDriverWait(driver, 10)
    heading = wait.until(
        EC.visibility_of_element_located((By.TAG_NAME, "h1"))
    )
    print(heading.text)
finally:
    driver.quit()

Choose an expected condition that matches the next operation. For example, wait for an element to be present before reading its attributes, visible before relying on its displayed text, or clickable before clicking. See Selenium’s waiting strategies for the available conditions.

Avoid mixing implicit and explicit waits. Selenium warns that combining them can produce unpredictable timeout behavior. For most scripts, keep the implicit wait at its default and use explicit waits at the points where the page state matters.

5. A practical form automation example

This example shows a typical sequence: open a page, locate fields, enter values, wait for a result, and clean up. Replace the example URL and selectors with ones from a page you are authorized to use. The page must contain elements matching these selectors for the example to work as written.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


driver = webdriver.Chrome()
try:
    driver.get("https://example.com/login")
    wait = WebDriverWait(driver, 10)

    email = wait.until(EC.visibility_of_element_located((By.NAME, "email")))
    password = driver.find_element(By.NAME, "password")
    email.send_keys("person@example.com")
    password.send_keys("replace-with-a-test-password")

    submit = wait.until(
        EC.element_to_be_clickable((By.CSS_SELECTOR, "button[type='submit']"))
    )
    submit.click()

    result = wait.until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "[role='status']"))
    )
    print(result.text)
finally:
    driver.quit()

Do not commit real passwords or tokens to source control. For a test suite, load secrets from an appropriate environment or secret store and use test accounts and data.

6. Choose a browser and execution location

The Python API currently lists Chrome, Edge, Firefox, Safari, WebKitGTK, WPEWebKit, and Remote protocol support. Selenium Manager’s automatic browser downloads are a narrower capability: its documentation describes managed downloads for Chrome, Firefox, and Edge, subject to platform and configuration caveats. The list of API-supported browser implementations should not be read as a promise that Manager can install each one.

Execution choice What it means Considerations
Local browser The browser runs on the machine running the script. Best starting point for learning and local scripts; no Java server required.
Installed browser with Selenium Manager Use a browser already on the machine; Manager can resolve a compatible driver. Manager can discover the browser and cache a driver. Network or platform restrictions can affect setup.
Manager-managed browser Manager can download certain browser builds as well as drivers. Its browser and platform support has specific caveats; Linux system libraries may be missing, and Windows Edge installation through Manager requires administrator permissions.
Remote WebDriver / Grid The Python client controls a browser session on a Grid or remote endpoint. Useful when browser execution belongs on another machine or in a managed test environment; the endpoint must be configured and reachable.

For remote execution, the shape is a Remote WebDriver session pointed at your Grid endpoint. Configure the endpoint and browser capabilities for your Grid deployment; those values are environment-specific.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options


options = Options()
driver = webdriver.Remote(
    command_executor="http://grid-host:4444",
    options=options,
)
try:
    driver.get("https://example.com")
    print(driver.title)
finally:
    driver.quit()

Replace grid-host with the actual Grid address. A Grid is not needed for a first local exercise. Consult the Selenium documentation for Grid setup and configuration.

7. Use Selenium in a test suite

Selenium controls the browser; a test framework organizes test cases, assertions, and setup/cleanup. The Selenium Python API includes examples for both unittest and pytest. Choose the framework that fits your project and keep browser lifecycle management explicit so a failed assertion does not leave a session behind.

import unittest
from selenium import webdriver
from selenium.webdriver.common.by import By


class ExamplePageTest(unittest.TestCase):
    def setUp(self):
        self.driver = webdriver.Chrome()

    def tearDown(self):
        self.driver.quit()

    def test_homepage_has_heading(self):
        self.driver.get("https://example.com")
        heading = self.driver.find_element(By.TAG_NAME, "h1")
        self.assertTrue(heading.text)


if __name__ == "__main__":
    unittest.main()

This is a basic structure. Add explicit waits for conditions that depend on asynchronous rendering, and select assertions that describe behavior your application is expected to provide.

8. Screenshots: browser automation or a screenshot API?

Selenium is appropriate when you need to interact with a browser: fill forms, click controls, validate flows, or inspect behavior after a sequence of actions. If the task is simply to capture a URL as an image or PDF, a screenshot API can avoid managing a browser session in your script.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. The API accepts screenshot options for tasks such as full-page capture, element selection, device and viewport settings, waits, and custom CSS or JavaScript. See the ScreenshotNeo API documentation for parameters and response details.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Equivalent requests are available with cURL and Node.js:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up for ScreenshotNeo and get 1,000 screenshots a month free, with no card.

9. Troubleshooting common Selenium problems

Symptom Likely cause What to do
ModuleNotFoundError: No module named 'selenium' Selenium was installed into a different Python environment. Activate the project virtual environment, then run python -m pip install -U selenium using the same python that runs the script.
Driver or browser setup fails The browser is unavailable, Selenium Manager cannot reach required metadata or downloads, or the environment has restricted network access. Confirm the browser is installed and available to the account running the script. Check network and proxy access. In controlled environments, configure the browser and driver explicitly as described in Selenium Manager documentation.
Browser starts locally but not in a container or Linux host Required system libraries or other platform dependencies may be absent. Review the browser’s system requirements and the Selenium Manager platform caveats; install dependencies in the execution environment or use a suitable remote browser.
Element lookup raises NoSuchElementException The locator is wrong, the element has not appeared yet, or the element is in a different browsing context such as an iframe. Inspect the current page structure and locator. Wait for the required condition. If the element is in an iframe, switch into the correct frame before locating it.
Wait raises TimeoutException The expected condition did not become true before the timeout. Check whether navigation reached the expected page, whether the locator matches, and whether the application is showing an error or different state. Increase the timeout only when the page can reasonably take longer; do not use a longer timeout to conceal a wrong locator.
Click is intercepted or does nothing An overlay, animation, or layout change blocks the target, or the target is not yet clickable. Wait for the overlay to disappear or the target to become clickable, then inspect the page state. Avoid repeated blind clicks.
Waits take much longer than expected Implicit and explicit waits may have been combined, or the chosen condition never becomes true. Remove the implicit wait and use an explicit wait for the condition needed by the next step. Selenium warns that mixing wait types can cause unpredictable timing.
Browser process remains after a failure The script exited before closing its session. Put browser work in try/finally and call driver.quit() in the cleanup path.
Remote session cannot be created The Grid URL is wrong or unreachable, the Grid has no matching capacity, or requested capabilities are unsupported. Check the endpoint from the client machine, inspect Grid availability, and request capabilities supported by that deployment.

10. Performance, reliability, and cost

  • Use the right wait. Condition-based waits let the script continue as soon as the required state appears, while fixed sleeps always spend the chosen delay and may still be too short.
  • Keep sessions short and clean. Reuse a session for a coherent workflow, and quit it when the work is done. A failed cleanup can leave browser processes consuming resources.
  • Account for environment setup. Selenium Manager can resolve and cache drivers, but the initial setup can depend on browser availability, platform support, network access, and permissions. Linux dependencies and Windows Edge permissions can matter.
  • Local and remote costs differ by setup. Local execution uses the machine’s browser and resources. Remote execution depends on the Grid or hosted environment you operate; Selenium itself does not make that environment free or provision it automatically.
  • Use a screenshot API for capture-only tasks. If you only need an image or PDF of a page, ScreenshotNeo offers a free monthly tier and paid plans with stated quotas. Browser interaction and application testing remain a different job from requesting a page capture.

11. FAQ

Do I need Java to use Selenium with Python?

No Java server is needed for local Python scripts. Remote sessions use a Selenium Grid or another Remote WebDriver endpoint.

Do I have to download ChromeDriver manually?

Usually not for common setups. Selenium Manager can manage missing drivers. Manual or explicit browser and driver configuration can still be appropriate in controlled environments.

Can Selenium automate browsers other than Chrome?

Yes. The Python API lists Chrome, Edge, Firefox, Safari, WebKitGTK, WPEWebKit, and Remote protocol support. Setup and automation details vary by browser and environment.

Should I use Selenium or a screenshot API?

Use Selenium when the task requires browser interaction or UI testing. Use a screenshot API when the required output is a capture of a URL and you do not need to drive a local browser through a workflow.

Official references