ScreenshotNeo

BlogHow-to

Getting Started with Selenium and Python for Browser Automation

Install Selenium, automate a browser with Python, and learn reliable element selection, waits, cleanup, and troubleshooting.

By the ScreenshotNeo team4 October 20268 min read

Selenium lets a Python program control a real browser: open a page, find elements, type or click, and read the result. For a first script, install the Selenium package, install a supported browser, start a WebDriver session, and close it with driver.quit(). Current Selenium versions usually use Selenium Manager to handle the browser driver automatically.

This guide uses Chrome and Selenium’s documented sample form. The same workflow applies to other supported browsers with their corresponding WebDriver. See the Selenium WebDriver documentation for browser-specific details.

1. Install Python, a browser, and Selenium

Use a virtual environment so this project’s Python packages stay separate from other projects. Install Python and a supported browser first, then run:

python -m venv .venv

# macOS or Linux
source .venv/bin/activate

# Windows PowerShell
# .venv\Scripts\Activate.ps1

python -m pip install selenium

For repeatable installs, record a pinned version in requirements.txt after choosing and validating the version for your environment:

selenium==4.49.0

4.49.0 is the version shown in Selenium’s installation documentation as an example; it is not a claim that this is the latest version when you read this. Install a browser such as Chrome as well. Selenium Manager can generally resolve the matching driver for current Selenium installations. See the official Selenium Manager documentation and Python installation guide.

2. Run a complete first script

Save this as first_selenium.py. It starts Chrome, opens Selenium’s sample form, types into the field, submits it, checks the displayed message, and always closes the browser session.

from selenium import webdriver
from selenium.webdriver.common.by import By


def main():
    driver = webdriver.Chrome()
    try:
        driver.get("https://www.selenium.dev/selenium/web/web-form.html")
        print("Page title:", driver.title)

        text_box = driver.find_element(By.NAME, "my-text")
        submit_button = driver.find_element(By.CSS_SELECTOR, "button")

        text_box.send_keys("Selenium")
        submit_button.click()

        message = driver.find_element(By.ID, "message")
        print("Result:", message.text)
    finally:
        driver.quit()


if __name__ == "__main__":
    main()

Run it from the activated environment with python first_selenium.py. The browser opens, performs the interaction, prints the page title and result, then exits. The calls shown follow the Selenium Project’s first Python script example; try/finally ensures cleanup if a lookup or action raises an exception.

3. Understand the WebDriver lifecycle

  1. Create a session: webdriver.Chrome() starts Chrome and creates a WebDriver session. The equivalent entry point varies by browser.
  2. Navigate: driver.get(url) asks the browser to open a page. Navigation completion does not guarantee every asynchronous widget or API-driven element is ready.
  3. Locate: find_element returns one matching element and raises an exception if none is found. find_elements returns a list, which is empty when there are no matches.
  4. Interact and inspect: common actions include send_keys, click, and reading .text or an attribute.
  5. Clean up: call driver.quit() when the work is done. It ends the WebDriver session and closes the browser by default. Use finally so this also happens on errors.

4. Find elements with the right locator

Selenium’s Python locators are available through selenium.webdriver.common.by.By. Prefer stable attributes that identify the intended control. An application-owned ID or a deliberate test attribute is often less fragile than a long CSS path or an element’s position in the page.

Locator Example Good fit
ID By.ID, "message" A unique, stable ID.
Name By.NAME, "my-text" Form fields with a name attribute.
CSS selector By.CSS_SELECTOR, "button[type='submit']" Combining tag names, classes, attributes, and relationships.
Class name By.CLASS_NAME, "notice" A single class token; do not pass multiple classes separated by spaces.
Tag name By.TAG_NAME, "button" Finding elements by HTML tag, often followed by filtering or scoping.
Link text By.LINK_TEXT, "Documentation" A link whose visible text is known and stable.
Partial link text By.PARTIAL_LINK_TEXT, "Document" A link when only a stable portion of its visible text is known.
XPath By.XPATH, "//button[@type='submit']" Relationships or conditions that are awkward in CSS; keep expressions readable.

Use find_elements to handle optional or repeated matches without an exception:

from selenium.webdriver.common.by import By

buttons = driver.find_elements(By.CSS_SELECTOR, "button.save")
if not buttons:
    print("No save button is present")
else:
    buttons[0].click()

When a page contains repeated controls, narrow the search to a parent element first, then locate within that element. This avoids clicking the wrong matching button.

5. Wait for pages and elements correctly

Web pages update asynchronously. A page can finish its initial navigation while JavaScript is still rendering controls, loading data, or changing the DOM. A timing bug often looks like an intermittent “element not found” or a click against an element that is not ready.

For most interaction scripts, use an explicit wait for the specific condition you need. This example waits until the result is visible:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

wait = WebDriverWait(driver, 10)
result = wait.until(
    EC.visibility_of_element_located((By.ID, "message"))
)
print(result.text)

Other useful expected conditions include presence_of_element_located, element_to_be_clickable, title_contains, and alert_is_present. Choose the condition that describes what the next step needs: presence does not necessarily mean visible, and visible does not necessarily mean enabled or clickable.

Selenium’s first-script example demonstrates an implicit wait as a simple introduction, but says it is rarely the best solution. Avoid mixing implicit and explicit waits because their combined timing can be difficult to reason about. Read the official waiting strategies guide for the details.

6. Make the browser setup reproducible

  • Keep dependencies local: activate the same virtual environment when installing and running the script.
  • Pin dependencies when needed: record the Selenium version used by a project and update it deliberately.
  • Use a clean browser session: for isolated runs, create a new WebDriver session rather than relying on browser state left by an earlier run.
  • Keep selectors maintainable: prefer stable IDs, names, or test-specific attributes over layout-dependent selectors.
  • Put cleanup in finally: this prevents a failed assertion or lookup from leaving a browser process running.
  • Use a test runner for suites: Selenium’s usage documentation shows an example organized with pytest. That section of the docs is marked incomplete, so treat it as a starting pointer rather than a complete testing reference.

For larger test runs across multiple machines or browsers, Selenium Grid is the project’s scaling pathway. Selenium IDE is a record-and-playback option for people seeking a lower-code introduction. A local Python WebDriver script remains the most direct way to learn browser control.

7. Troubleshoot common failures

Symptom Likely cause What to do
ModuleNotFoundError: No module named 'selenium' Selenium was installed into a different Python environment. Activate the project virtual environment, then run python -m pip install selenium using the same python that runs the script.
Driver or browser startup error The browser is missing, Selenium is old, or the environment prevents Selenium Manager from resolving a driver. Install a supported browser and use a current Selenium release. In restricted or older environments, consult Selenium’s driver setup guidance and configure a compatible driver explicitly.
NoSuchElementException The selector is wrong, the element has not appeared yet, or it is inside a frame. Inspect the page and selector, wait for the appropriate condition, and switch into the correct frame before searching when needed.
TimeoutException from a wait The condition never became true within the timeout, perhaps because the page state or selector differs from expectation. Check the locator and expected state, confirm navigation succeeded, and choose a realistic timeout. Do not simply add a long sleep without understanding the missing condition.
ElementClickInterceptedException An overlay or another element covers the target, or the page has not settled. Wait for the overlay to disappear or for the button to become clickable. Verify that the intended control is visible and unobstructed.
StaleElementReferenceException The page replaced or rerendered the element after it was located. Wait for the new state and locate the element again instead of reusing the old reference.
Browser stays open after an error Cleanup was skipped on an exception path. Wrap browser work in try/finally and call driver.quit() in the finally block.
Site shows a CAPTCHA or blocks automation The site is detecting or disallowing automated access. Respect the site’s terms and access rules. Do not try to bypass access controls; use an authorized API or request permission.

8. Performance, reliability, and cost

A WebDriver session launches and controls a real browser, so it uses more time and resources than a direct HTTP request. Keep sessions scoped to the work, avoid unnecessary browser restarts inside a small batch, and wait for the exact state needed rather than inserting arbitrary delays. For parallel work, account for the CPU and memory used by each browser session and use Grid when execution needs to span machines or browser instances.

Reliability comes from stable selectors, explicit waits, fresh element lookups after page updates, and guaranteed cleanup. Browser automation can still vary with network conditions, page changes, browser versions, and site behavior; make failures observable by recording which step failed and the browser/page context.

The Selenium package is open source and does not itself charge per browser action. Operating cost comes from running the browser and infrastructure: developer machine resources, hosted runners, or any grid environment. Validate current browser support and package versions against the official docs when setting up a maintained project.

9. Or skip the browser setup

If your goal is a clean screenshot rather than interacting with controls, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. Its API accepts common screenshot API parameter names, which can make switching easier. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say which page verdict occurred and whether it was billed. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card.

10. Frequently asked questions

Do I need to download ChromeDriver separately?

Usually not with a current Selenium installation on a supported setup: Selenium Manager handles driver management in most cases. You still need the browser, and older or restricted environments may require explicit configuration.

Can Selenium automate a site without an API?

It can control a browser for many automation tasks, but check the site’s terms and access rules. Some sites prohibit scraping or block Selenium.

Should I use Selenium IDE or write Python?

Selenium IDE offers record and playback for a lower-code start. A Python WebDriver script gives direct control over program logic, assertions, and integration with Python tools.

What is the key cleanup call?

Use driver.quit() to end the WebDriver session and close the browser by default.

Sources