ScreenshotNeo

BlogHow-to

How to Build a Selenium Automation Framework

Build a maintainable Selenium framework from a first browser test to reliable waits, Page Objects, CI execution, and Selenium Grid.

By the ScreenshotNeo team4 October 20269 min read

Selenium automation frameworks combine browser setup, readable tests, synchronization, and execution infrastructure. Start with the language and test runner your team already supports, get one local WebDriver test passing, then add reusable page abstractions and explicit waits. Introduce Selenium Grid when remote browsers, parallel capacity, or cross-platform coverage justify the added infrastructure.

Selenium is the umbrella project. WebDriver is the central API for controlling browsers; Selenium also includes Selenium IDE, Grid, and Selenium Manager. WebDriver is a W3C Recommendation. Selenium does not require one language, runner, or framework architecture. Selenium project overview · WebDriver documentation

1. Choose a language and test runner

Use a language the team can maintain and a runner that already fits the project’s build and CI setup. Selenium provides language bindings and a language-neutral WebDriver interface; the documentation does not prescribe one runner for every project.

Before committing to a design, answer these questions:

  • Which language is already used for application code or tests?
  • Which runner integrates with your package manager, reporting, and CI?
  • Which browsers and operating systems must be covered?
  • Will tests run serially at first, or is parallel capacity already needed?
  • Who will maintain selectors and any remote browser infrastructure?

The examples below use Python with pytest for the complete starter framework. Installation and setup details vary by binding and version; consult the current Selenium documentation for your chosen language.

2. Install Selenium and verify a local browser session

Install Python, a browser such as Chrome or Firefox, and the Selenium binding. Selenium Manager is used by Selenium bindings by default to help manage browsers and drivers, which can reduce manual driver configuration. Check the current binding documentation for version-specific behavior and any environment requirements.

python -m pip install selenium pytest

Create tests/test_homepage.py with a small end-to-end test. This example uses Selenium’s documentation page as its target and verifies a visible page heading. The test owns the browser lifecycle so the session is closed even when an assertion fails.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


def test_selenium_documentation_opens():
    driver = webdriver.Chrome()
    try:
        driver.get("https://www.selenium.dev/documentation/")
        heading = WebDriverWait(driver, 10).until(
            EC.visibility_of_element_located((By.TAG_NAME, "h1"))
        )
        assert "Selenium" in heading.text
    finally:
        driver.quit()

Run it from the project root:

python -m pytest -q

Use a target page and assertion owned by your project for production tests. A minimal test should prove the whole path works: start a browser, navigate, wait for a meaningful state, assert an outcome, and close the session.

3. Organize tests around behavior

Keep each test focused on a user-visible behavior and its outcome. Avoid putting all browser logic in one large test file, but do not add layers until they make repeated work easier to understand or maintain.

A practical starter layout

project/
  tests/
    conftest.py
    pages/
      home_page.py
    test_homepage.py
  requirements.txt

Once more than one test needs the same browser setup, put lifecycle management in a pytest fixture:

# tests/conftest.py
import pytest
from selenium import webdriver


@pytest.fixture
def driver():
    browser = webdriver.Chrome()
    yield browser
    browser.quit()

Then tests request the fixture rather than constructing their own session:

# tests/test_homepage.py
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


def test_documentation_heading_is_visible(driver):
    driver.get("https://www.selenium.dev/documentation/")
    heading = WebDriverWait(driver, 10).until(
        EC.visibility_of_element_located((By.TAG_NAME, "h1"))
    )
    assert "Selenium" in heading.text

4. Add Page Objects when they reduce duplication

A Page Object keeps knowledge of a page’s structure and operations in one place. Use one when multiple tests share selectors or actions, or when page changes would otherwise require edits across many test cases. For a tiny suite, direct interactions can be clearer than an abstraction.

For example, create tests/pages/home_page.py:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


class DocumentationHome:
    URL = "https://www.selenium.dev/documentation/"
    HEADING = (By.TAG_NAME, "h1")

    def __init__(self, driver):
        self.driver = driver

    def open(self):
        self.driver.get(self.URL)
        return self

    def heading_text(self):
        heading = WebDriverWait(self.driver, 10).until(
            EC.visibility_of_element_located(self.HEADING)
        )
        return heading.text

The test describes the behavior and keeps the ordinary assertion in the test:

from tests.pages.home_page import DocumentationHome


def test_documentation_home_has_heading(driver):
    page = DocumentationHome(driver).open()
    assert "Selenium" in page.heading_text()

Keep ordinary test assertions in the test rather than embedding them in page methods. A page object may verify that the expected page loaded when that check establishes the object represents the correct page. The Selenium guide describes this separation in its Page Object Models guidance.

5. Make tests wait for the state they need

JavaScript-driven applications can still be changing after navigation returns. Selenium describes the resulting race between application readiness and the next command as a common source of flaky tests. Wait for the condition that matters at the point it matters: an element visible, clickable, or present.

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

save_button = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.CSS_SELECTOR, "button.save"))
save_button.click()

Common explicit-wait conditions include presence, visibility, clickability, and invisibility of an element. Choose the condition that matches the next action or assertion. A fixed sleep can be too short on a slow run and needlessly long on a fast one.

Avoid combining implicit and explicit waits. Selenium warns that mixing them can produce unpredictable total wait times. Prefer explicit waits for dynamic states, and keep wait behavior easy to see in the test or page operation. Selenium waits documentation

6. Run tests in CI and add Grid when needed

Keep the first version local and serial. Add remote execution when the browser and operating-system matrix, CI runtime, or available local capacity calls for it. Selenium Grid routes WebDriver commands to remote browser instances; it supports parallel execution across machines and coverage across browser versions and platforms. A standalone Grid server is one available deployment path; hub/node deployment is another. The documentation does not define a universal point at which every team should adopt Grid.

Grid introduces infrastructure to configure, monitor, and maintain. Compare the coverage and capacity it enables with its setup and operational cost. A remote session uses Selenium’s remote WebDriver client; the exact endpoint and capabilities depend on the Grid deployment and browser configuration.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("--headless")

driver = webdriver.Remote(
    command_executor="http://localhost:4444",
    options=options,
)
try:
    driver.get("https://www.selenium.dev/documentation/")
    print(driver.title)
finally:
    driver.quit()

Start a Grid server using the current Grid getting-started instructions for your chosen deployment. Confirm that the remote endpoint is reachable and that the requested browser is available before pointing CI tests at it.

7. Configuration choices that affect maintainability

Choice Start with Change when
Browser setup Selenium Manager and a local installed browser CI or managed infrastructure needs explicit browser and driver provisioning
Wait strategy Explicit waits around the needed state Review individual timeouts when the application has measured slower transitions
Test structure Behavior-focused tests; direct driver use for a small suite Add Page Objects or component objects as selectors and actions are reused
Execution Local, serial browser sessions Add Grid for remote, parallel, cross-browser, or cross-platform needs
Assertions Assert outcomes in the test Use page-level checks only to establish that the expected page is represented

Keep browser and runner configuration visible to the team. Avoid hiding navigation, assertions, and synchronization inside generic helpers whose behavior is difficult to inspect when a test fails.

8. Troubleshooting common failures

Symptom Likely cause Fix
Driver or browser cannot start Browser missing, incompatible setup, or environment cannot obtain the required driver Confirm the browser is installed and review Selenium Manager and binding setup guidance for the installed Selenium version. Check network and permissions if automated management cannot complete.
Element not found immediately after navigation Dynamic content has not reached the needed state, or the locator is wrong Verify the locator against the current page and wait explicitly for presence or visibility before using the element.
Click intercepted or element not clickable Overlay, animation, or page state prevents interaction Wait for the element to be clickable and for any known blocking overlay to disappear; confirm the application reached the intended state.
Intermittent timeout Timing race, unstable environment, or an incorrect readiness condition Wait for the actual application state needed by the test. Inspect the failed page and environment; do not simply increase every timeout.
Wait takes much longer than expected Implicit and explicit waits are mixed Remove the mixed wait configuration and use explicit waits for state-dependent operations.
Remote session cannot connect Grid endpoint is unavailable, wrong, or inaccessible from the test process Check the configured command executor URL, server availability, network path, and browser capabilities against the Grid setup.
Tests pass locally but fail in CI Different browser availability, timing, permissions, or environment configuration Compare browser and Selenium versions, use explicit conditions, and make the CI browser setup and endpoint configuration visible.
Browser remains open after a failure Session cleanup did not run Use a fixture teardown or a finally block that calls quit().

9. Performance, reliability, and cost

There is no universal speed or flakiness figure for a Selenium framework. Runtime depends on the application, browser, environment, waits, and how many sessions run concurrently. Use parallel execution only when the suite and infrastructure can safely support independent sessions; Grid can distribute sessions, but also adds infrastructure to operate.

Reliability comes from synchronizing on application state, keeping tests focused, making selectors maintainable, and reliably closing sessions. More layers or longer timeouts do not automatically make a suite robust. Start with the smallest architecture that keeps behavior and failure causes understandable.

Cost is primarily engineering and infrastructure effort: maintaining browsers and CI environments, operating Grid if used, and spending time diagnosing failures. The Selenium documentation describes product capabilities, not a single cost model or a benchmark for choosing a deployment size.

Or skip the browser setup

If your task is to capture a website image or PDF rather than drive an interactive browser test, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, consent prompts, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free for 1,000 screenshots a month, no card required.

Frequently asked questions

Is Selenium WebDriver a test framework?

WebDriver is Selenium’s browser-control API. A test runner and your project’s organization provide the surrounding test framework.

Do I need Page Objects from the first test?

No. Add them when centralizing page selectors and operations makes repeated tests easier to maintain.

When should a team move from local runs to Grid?

When remote browsers, cross-platform or cross-browser coverage, or execution capacity justify the additional infrastructure. There is no documentation-backed universal threshold.

Can Selenium test a site that renders content with JavaScript?

Yes. Synchronize on the state the test needs, since navigation completion alone may not mean the application is ready.