ScreenshotNeo

BlogGuides

7 Selenium Tutorials for Beginners

Write your first Selenium script in Python: install Selenium, open a browser, interact with a form, verify the result, and plan what to learn next.

By the ScreenshotNeo team4 October 202611 min read

Selenium WebDriver lets code control a real browser: open a page, find elements, interact with them, and check what happened. This beginner guide uses Python for the main path and builds one working local script in seven steps. You need Python, the Selenium package, and a supported browser. With current Selenium releases, Selenium Manager can usually manage the browser driver automatically, so a manual driver download is not the default first step.

The [Selenium Project describes WebDriver](https://www.selenium.dev/documentation/webdriver/) as a browser automation interface and a W3C Recommendation. The official project documentation also says: “WebDriver drives a browser natively; learn more about it.”

Before you start: choose your learning path

Path Best for What you learn
Selenium IDE A low-code introduction Record and play back browser actions with an extension.
Selenium WebDriver Code-based automation Write scripts and tests in a language binding. This is the path used below.
Local WebDriver Your first script Run a browser on your own machine.
Selenium Grid or a hosted browser service Remote or distributed execution Run tests across machines or browser environments. This is optional after local basics.

For this walkthrough, use Python 3, pip, and a supported browser such as Chrome, Firefox, or Edge. Selenium Manager is included with Selenium releases and can manage a missing driver in supported setups. If your environment supplies a driver explicitly, Selenium can use that instead. See the official [Selenium installation guide](https://www.selenium.dev/documentation/webdriver/getting_started/install_library/) and [Selenium Manager documentation](https://www.selenium.dev/documentation/selenium_manager/).

1. Install Selenium and prepare a browser

Create a project directory and virtual environment, then install the Python binding:

mkdir selenium-beginner
cd selenium-beginner
python -m venv .venv

Activate the environment. On macOS or Linux:

source .venv/bin/activate

On Windows PowerShell:

.venv\Scripts\Activate.ps1

Install Selenium:

python -m pip install --upgrade selenium

Install or update a browser supported by your Selenium version. The package and browser are separate setup components. In supported setups Selenium Manager resolves and manages the driver when you create a session; first startup may need network access to obtain a driver. You generally do not need to download a driver manually for this tutorial.

Confirm the installation

python -c "import selenium; print(selenium.__version__)"

If this prints a version, Python can import the package. If it says the module is missing, check that the virtual environment is active and that pip and Python refer to the same environment.

2. Start and close a WebDriver session

A WebDriver session controls a browser process. Start a session by creating a browser-specific driver object, and always call quit() to close the session even if a later step fails. Selenium’s [first-script guide](https://www.selenium.dev/documentation/webdriver/getting_started/first_script/) follows this lifecycle.

from selenium import webdriver

browser = webdriver.Chrome()
try:
    print("Browser session started")
finally:
    browser.quit()

For Firefox, replace webdriver.Chrome() with webdriver.Firefox(). For Edge, use webdriver.Edge(). Browser availability and management depend on your installed browser, Selenium version, and environment.

3. Navigate to a page and inspect it

Use the stable Selenium sample form for practice. Reading the page title confirms that navigation reached the expected page.

from selenium import webdriver

browser = webdriver.Chrome()
try:
    browser.get("https://www.selenium.dev/selenium/web/web-form.html")
    print(browser.title)
finally:
    browser.quit()

get() navigates to the URL and waits for the page load according to the browser’s page-load behavior. It does not guarantee that every asynchronous application update or lazy resource has completed. Use an explicit wait for the specific state your next step needs.

4. Find page elements with locators

A locator tells Selenium how to identify an element. The sample form provides a text field, a submit button, and a result message. Prefer stable attributes such as an ID, name, or purpose-built test attribute when the page provides them.

from selenium import webdriver
from selenium.webdriver.common.by import By

browser = webdriver.Chrome()
try:
    browser.get("https://www.selenium.dev/selenium/web/web-form.html")
    text_field = browser.find_element(By.NAME, "my-text")
    submit_button = browser.find_element(By.CSS_SELECTOR, "button")
    print(text_field.tag_name, submit_button.text)
finally:
    browser.quit()

The official sample uses locators such as name, CSS selector, and ID. Common choices are:

Locator Example Use
ID By.ID, "submit" A unique, stable element ID.
Name By.NAME, "my-text" A form control name.
CSS selector By.CSS_SELECTOR, "button[type='submit']" Precise selection using HTML attributes and relationships.
XPath By.XPATH, "//button[@type='submit']" Cases that need XPath relationships or text matching.
Class name By.CLASS_NAME, "result" A single class token; not a space-separated list of classes.

find_element returns one match and raises NoSuchElementException if none exists. find_elements returns a list, which is empty when there are no matches. Avoid brittle selectors tied to generated class names or deep page structure where a stable locator is available.

5. Enter text and trigger an action

Use send_keys() to type into the field, then click the button. Clear a field first if a test must replace existing text.

from selenium import webdriver
from selenium.webdriver.common.by import By

browser = webdriver.Chrome()
try:
    browser.get("https://www.selenium.dev/selenium/web/web-form.html")
    field = browser.find_element(By.NAME, "my-text")
    field.clear()
    field.send_keys("Selenium beginner")
    browser.find_element(By.CSS_SELECTOR, "button").click()
finally:
    browser.quit()

These commands interact with the rendered page. If an element is covered by a dialog, disabled, or not yet interactable, the browser may reject the action. Wait for the relevant state and handle the page’s actual interaction flow rather than forcing a click blindly.

6. Wait for and verify an outcome

A script that clicks a button but never checks the result can silently fail. Use an explicit wait for a meaningful page condition, then assert it. Explicit waits poll for a condition until it succeeds or the timeout expires.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

browser = webdriver.Chrome()
try:
    browser.get("https://www.selenium.dev/selenium/web/web-form.html")
    browser.find_element(By.NAME, "my-text").send_keys("Selenium beginner")
    browser.find_element(By.CSS_SELECTOR, "button").click()

    result = WebDriverWait(browser, 10).until(
        EC.visibility_of_element_located((By.ID, "message"))
    )
    assert result.text == "Received!", f"Unexpected result: {result.text!r}"
    print("Form submission verified")
finally:
    browser.quit()

Save the full example as first_selenium.py and run python first_selenium.py. A successful run prints Form submission verified. If the sample page changes, update the expected result to match the page’s current behavior.

Wait choices

  • Explicit wait: use WebDriverWait with a condition for the element or state needed next. This is the usual choice for dynamic pages.
  • Implicit wait: configures a default search wait for element lookups with browser.implicitly_wait(seconds). Avoid mixing implicit and explicit waits casually; their combined timing can be difficult to reason about.
  • Fixed sleep: time.sleep(seconds) pauses for a fixed duration regardless of whether the page is ready. Reserve it for cases where a fixed delay is specifically required, not as a general synchronization strategy.

7. Organize the script and choose the next step

Keep browser setup and cleanup in one place, and keep assertions close to the behavior they verify. The following standard-library test makes the example repeatable without adding a test-runner dependency:

import unittest
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

class FormTest(unittest.TestCase):
    def setUp(self):
        self.browser = webdriver.Chrome()

    def tearDown(self):
        self.browser.quit()

    def test_form_submission(self):
        self.browser.get("https://www.selenium.dev/selenium/web/web-form.html")
        self.browser.find_element(By.NAME, "my-text").send_keys("Selenium beginner")
        self.browser.find_element(By.CSS_SELECTOR, "button").click()
        result = WebDriverWait(self.browser, 10).until(
            EC.visibility_of_element_located((By.ID, "message"))
        )
        self.assertEqual(result.text, "Received!")

if __name__ == "__main__":
    unittest.main()

Save as test_form.py and run python -m unittest. As a project grows, choose a test runner that fits your language and team; no single runner is right for every Selenium binding. Selenium IDE remains an option for record-and-playback learning, while WebDriver is the code-based approach used here.

Run locally first. [Selenium Grid](https://www.selenium.dev/documentation/grid/) is the next option when you need to distribute execution across machines and browsers. A hosted service is another optional route to remote execution; [Sauce Labs documents a Selenium cloud quickstart](https://docs.saucelabs.com/web-apps/automated-testing/selenium/). Neither Grid nor a cloud service is needed to complete the local tutorial.

Or skip the browser setup

If your task is to capture a page as an image or PDF rather than interact with it as a test, [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. The API has options for full-page or selected-element captures, viewport and device presets, waits, custom CSS and JavaScript, cookies, headers, and more. See the [API documentation](https://screenshotneo.com/docs/).

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace YOUR_API_KEY with your key. For Node.js versions without Bun.write, write the response bytes with Node’s file system API:

import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Selenium remains the right tool when you need browser interaction and assertions; ScreenshotNeo is the alternative when the deliverable is a screenshot or PDF.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

Troubleshooting

Symptom Likely cause What to do
ModuleNotFoundError: No module named 'selenium' Selenium was installed into a different Python environment. Activate the project environment and run python -m pip install selenium with the same Python executable used to run the script.
Driver or browser session fails to start Browser is missing or incompatible, Selenium Manager cannot resolve a driver, or the environment blocks required downloads. Install/update a supported browser, check network and proxy access, and inspect the Selenium Manager error details. In restricted environments, provide a compatible driver using the setup supported by your Selenium version.
NoSuchElementException Wrong locator, page not at the expected state, or element inside a different frame. Confirm the current URL and locator against the rendered page; wait for the element and switch to the relevant frame when needed.
TimeoutException The expected condition did not occur before the wait expired. Check that the page action succeeded, the locator is correct, and the page has no validation or network error. Increase the timeout only when the page legitimately needs more time.
ElementClickInterceptedException An overlay or another element covers the target. Wait for the overlay to disappear or handle the dialog, then click when the target is interactable.
Script exits but browser remains open Cleanup did not run or the process was terminated abruptly. Put quit() in a finally block or a test teardown method. close() closes a window; quit() ends the session.
Works locally but fails in automation Different browser version, headless environment, permissions, network, or timing. Record browser and Selenium versions, use explicit waits, and configure the execution environment deliberately before adding Grid or hosted execution.

Performance, reliability, and cost

  • Startup: creating a browser session is heavier than a simple HTTP request. Reuse a session for related steps within a test, then close it reliably. Avoid sharing one session across independent parallel tests.
  • Waits: wait for specific conditions rather than sleeping for the slowest expected page. This helps avoid both premature actions and unnecessary delay.
  • Reliability: use stable locators, assert observable outcomes, and keep the browser and driver environment consistent. A passing sequence of commands is not proof of the intended page state.
  • Execution scale: a local browser is sufficient for learning. Grid distributes execution; a hosted provider can run remote browsers. These add setup and service costs that depend on the chosen environment, so check the provider’s current terms for your use case.
  • Screenshot cost: local Selenium’s software setup does not require a screenshot API. If the job is image/PDF capture, ScreenshotNeo has 1,000 free shots per month; its listed paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. Every feature is on every plan.

FAQ

How do I write my first Selenium script?

Install a language binding, create a WebDriver session, navigate, locate and interact with elements, wait for an outcome, assert it, and call quit(). The complete Python script above follows that sequence.

Do I need to download a browser driver?

Usually not for a supported current Selenium setup: Selenium Manager can manage a missing driver automatically. Manual driver setup remains useful when the environment requires a pinned or preinstalled driver.

Is Selenium IDE the same as WebDriver?

No. Selenium IDE is a record-and-playback extension; WebDriver is the code-based browser automation interface.

Do I need Selenium Grid to learn Selenium?

No. A local browser session is enough for the first script. Grid is relevant when you want distributed execution across machines and browsers.

Can Selenium replace a screenshot API?

Selenium can capture browser screenshots as part of automation, but if you only need a website image or PDF, a screenshot API such as ScreenshotNeo avoids setting up a local browser session.

Primary references