ScreenshotNeo

BlogHow-to

How to Build a Hybrid Framework in Selenium

Build a maintainable Selenium framework by combining data-driven tests with Page Objects, explicit waits, and a reusable WebDriver lifecycle.

By the ScreenshotNeo team4 October 202610 min read

A maintainable Selenium hybrid framework combines a test runner, reusable Page Objects, test data, and a small layer for browser configuration and lifecycle. In this guide, “hybrid” means data-driven testing plus Page Objects: test cases supply different inputs to the same test intent, while page objects hold page-specific locators and actions. Selenium does not prescribe one canonical hybrid framework recipe.

This example uses Python, pytest, and Selenium WebDriver. WebDriver controls the browser; pytest runs tests and provides assertions. Page Objects keep UI operations reusable, while explicit waits synchronize actions with the application. Start with a local browser and use Selenium Grid when remote sessions or broader parallel browser coverage become useful. Selenium’s documentation describes these responsibilities and components separately: Selenium overview, Page Objects, and waits.

1. Decide what “hybrid” means for your project

Teams use “hybrid” to mean different combinations, such as data-driven tests, keyword-driven steps, or behavior-driven scenarios. State the combination explicitly so the architecture does not become a collection of patterns without clear responsibilities. Here the combination is:

  • Data-driven: parameterize a test with multiple input and expected-result cases.
  • Page Object Model: encapsulate page-specific locators and actions behind focused methods.
  • Test runner: use pytest for test discovery, execution, fixtures, and assertions.

Add a keyword or behavior layer only when it helps the team express and maintain test scenarios. Selenium’s WebDriver does not provide test assertions or test execution; those belong to a test framework. Selenium lists framework options by language, including pytest and unittest for Python, JUnit and TestNG for Java, NUnit and MSTest for .NET, and Jest and Mocha for JavaScript: organizing and executing Selenium code.

2. Create the project and install dependencies

Use a supported Python runtime and a browser installed in the local environment. Selenium Manager, included with Selenium releases, can manage a browser driver when one has not been supplied. Pin dependency versions in your project’s dependency file and confirm the runtime, browser, binding, runner, and CI environment work together.

mkdir selenium-hybrid
cd selenium-hybrid
python -m venv .venv

Activate the environment, then install the dependencies:

# macOS or Linux
source .venv/bin/activate

# Windows PowerShell
# .venv\Scripts\Activate.ps1

python -m pip install selenium pytest

For repeatable builds, record the resolved versions after installation and install from a pinned requirements file in CI. For example, create one with python -m pip freeze > requirements.txt; review and maintain it as part of normal dependency updates. Selenium’s installation guidance shows pip installation and pinned dependencies: installing Selenium libraries.

3. Organize tests, pages, and support code

This layout is an example, not a Selenium requirement. Keep test intent and assertions in tests, page operations in page objects, input cases in test data, and browser setup in a fixture.

selenium-hybrid/
├── pages/
│   ├── __init__.py
│   └── example_page.py
├── tests/
│   └── test_example_page.py
├── conftest.py
└── requirements.txt

Keep page objects focused on the services a page offers. Tests should assert the observed outcome; a page object generally should not make test assertions or expose its locator internals. That keeps a UI change localized and avoids duplicated page-specific code, consistent with Selenium’s Page Object guidance.

4. Add browser configuration and lifecycle

Create a pytest fixture that opens a session and always quits it, including when a test fails. The optional HEADLESS environment variable selects headless mode. Selenium Manager handles the driver when no driver path is supplied.

# conftest.py
import os

import pytest
from selenium import webdriver
from selenium.webdriver.chrome.options import Options


@pytest.fixture
def driver():
    options = Options()
    if os.getenv("HEADLESS", "0").lower() in {"1", "true", "yes"}:
        options.add_argument("--headless=new")

    browser = webdriver.Chrome(options=options)
    browser.set_window_size(1440, 1000)
    browser.set_page_load_timeout(30)
    try:
        yield browser
    finally:
        browser.quit()

For another browser, use its Selenium driver and browser-specific options in this fixture. Keep browser selection in configuration or an environment variable if the suite needs multiple browsers. Do not silently swallow setup errors: a missing browser, unavailable driver download, or invalid browser option should be visible as an infrastructure failure.

5. Build a Page Object with condition-based waits

A completed navigation does not mean a JavaScript-driven page is ready for the next interaction. Wait for the specific condition the next step needs, such as a visible element or a clickable control. Avoid fixed sleeps for ordinary synchronization because they either waste time or proceed too early when the page is slower.

# pages/example_page.py
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


class ExamplePage:
    URL = "https://www.selenium.dev/"
    HEADING = (By.CSS_SELECTOR, "main h1")

    def __init__(self, driver, timeout=10):
        self.driver = driver
        self.wait = WebDriverWait(driver, timeout)

    def open(self):
        self.driver.get(self.URL)
        self.wait.until(EC.visibility_of_element_located(self.HEADING))
        return self

    def heading_text(self):
        heading = self.wait.until(
            EC.visibility_of_element_located(self.HEADING)
        )
        return heading.text

The page object exposes open() and heading_text(), operations meaningful to a test. Its locator remains an implementation detail. Choose wait conditions based on the application: presence may be enough for a hidden element that will be read from the DOM, while visibility or clickability is usually more appropriate before interacting. Selenium documents explicit waits and condition support in its waiting strategies.

6. Write a data-driven test

Pytest parameterization runs the same test intent with each input and expected result. Keep assertions in the test, where the pass/fail decision belongs.

# tests/test_example_page.py
import pytest

from pages.example_page import ExamplePage


@pytest.mark.parametrize(
    "expected_text",
    ["Selenium automates browsers"],
)
def test_homepage_heading(driver, expected_text):
    page = ExamplePage(driver).open()
    assert expected_text in page.heading_text()

Run the suite from the project root:

python -m pytest -q

To add cases, supply additional input and expected-result rows. If cases become large or need to be maintained outside Python, load structured data in a dedicated test-data layer, validate its schema, and report which case failed. Do not put arbitrary test control flow into a page object simply to make a data table drive every action.

7. Choose a runner and execution topology

Choose the runner that fits the language and team, then compare its parameterization, parallel execution, plugins, and CI/reporting integration. Selenium’s examples name several runners; for Java, its guidance calls out TestNG’s parameterized and parallel features. The runner owns test orchestration; WebDriver owns browser communication.

Need Starting choice When to consider a change
Python tests and data cases pytest Use another runner if team conventions or integrations make it a better fit.
Local browser session WebDriver with Selenium Manager Move when browser installation, network access, or machine capacity makes local setup impractical.
Remote sessions and distributed coverage Selenium Grid Use when parallel runs, browser/OS combinations, or remote execution are requirements.

Grid routes remote browser sessions and can distribute execution across machines. Selenium’s Grid getting started guide demonstrates a standalone server and the endpoint clients use: Grid getting started. Grid adds infrastructure and network configuration to operate, so begin locally while the suite is small and move when the coverage or capacity need is clear.

8. Run against Selenium Grid

Start a Grid server according to the official setup guide, then point the fixture at its remote endpoint. A common pattern is to make the endpoint configurable so local and CI runs share the test and page code.

# conftest.py: remote-capable fixture fragment
import os
import pytest
from selenium import webdriver
from selenium.webdriver.chrome.options import Options


@pytest.fixture
def driver():
    options = Options()
    if os.getenv("HEADLESS", "0").lower() in {"1", "true", "yes"}:
        options.add_argument("--headless=new")

    grid_url = os.getenv("SELENIUM_REMOTE_URL")
    if grid_url:
        browser = webdriver.Remote(
            command_executor=grid_url,
            options=options,
        )
    else:
        browser = webdriver.Chrome(options=options)

    try:
        yield browser
    finally:
        browser.quit()

Set SELENIUM_REMOTE_URL to the Grid URL accessible from the test runner. Configure Grid nodes and browser capabilities for the browsers you intend to run; do not assume a local browser option or installed browser exists on every remote node. See the Grid overview for its remote and distributed session model.

9. Make the framework reliable and efficient

  • Wait for state: use explicit waits for the next required condition. Selenium identifies races between application state and test commands as a common source of flakiness. Avoid mixing implicit and explicit waits without understanding the resulting wait behavior.
  • Keep sessions isolated: create and quit sessions predictably. Shared browser state can make tests order-dependent; use fresh sessions when isolation matters.
  • Use stable locators: prefer selectors tied to stable application attributes or accessible names over brittle layout-dependent paths. Keep locator changes inside the page object.
  • Control parallelism: parallel workers consume browser and Grid capacity. Increase concurrency only while the environment can serve sessions reliably, and ensure test data does not collide across workers.
  • Separate application failures from infrastructure failures: capture enough context in CI logs to tell whether a page assertion failed, a browser session could not start, or a remote node was unavailable.
  • Use Selenium Manager with network constraints in mind: it may need to reach driver or browser version endpoints. Corporate proxies and restricted networks can prevent downloads; plan an approved driver or browser provisioning path. The documentation also describes platform support limitations, including Linux ARM/aarch64 limitations. See Selenium Manager.

There is no universal cost figure for a Selenium framework. The main operational costs are engineering time to maintain tests and page objects, browser and CI capacity, and any infrastructure needed to run Grid. Local execution is simpler to start; remote distribution can provide more execution capacity and platform coverage but adds setup and operations.

10. Troubleshooting common failures

Symptom Likely cause Fix
Driver or browser session fails to start Browser unavailable, incompatible environment, or Selenium Manager cannot reach required endpoints. Check browser installation, network/proxy access, Selenium and browser versions, and platform support. Use the documented driver provisioning approach for restricted environments.
Element not found immediately after navigation The application has not rendered the target element yet, or the locator is stale or incorrect. Wait for the relevant condition and verify the locator against the current page state. Keep locator definitions in the Page Object.
Element is present but interaction fails It may be hidden, disabled, covered by an overlay, or not yet clickable. Wait for visibility or clickability; handle the application overlay or state explicitly instead of adding a blind delay.
Test passes alone but fails in the suite Tests may share browser state, test data, or depend on execution order. Use isolated sessions and unique test data; remove order dependencies.
Remote session cannot connect Incorrect Grid endpoint, network path, or unavailable node/capability. Check the configured remote URL from the test runner, Grid status, network access, and requested browser capability.
Suite is slow despite short assertions Fixed sleeps, repeated browser setup, excessive navigation, or constrained parallel capacity. Replace unnecessary sleeps with condition waits, remove redundant navigation, and tune parallelism to available local or Grid capacity.
Page object changes keep breaking many tests Page-specific implementation details leak into tests or page methods are too generic. Keep locators private to the page object and expose focused page services; keep test assertions in tests.

11. Capture a browser result for debugging

When investigating a visual state in a test failure, a screenshot can help show what the browser rendered. Selenium can save the current viewport from a WebDriver session:

# Add inside a failure-handling path when the driver is still available
from pathlib import Path

Path("artifacts").mkdir(exist_ok=True)
driver.save_screenshot("artifacts/failure.png")

For a rendered page capture without managing the browser setup yourself, ScreenshotNeo provides a website screenshot API and MCP server. The Selenium-generated screenshot above is useful inside a test workflow; the API below is a separate capture path. Full ScreenshotNeo API documentation covers its request options.

Or skip the browser setup

Send one GET request to capture a URL as an image or PDF. This Python example writes the returned image bytes to a file:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Equivalent cURL and Node.js calls:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with the result identified in response headers. Its MCP server lets AI agents using Claude, Cursor, or another MCP client take screenshots. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo and the API docs. Create a free account and get 1,000 screenshots a month with no card.

FAQ

Does Selenium provide a hybrid framework generator?

No. Selenium provides browser automation components and documentation, not one prescribed hybrid framework structure. Choose and describe the patterns your project combines.

Should every test use a Page Object?

Use one where it improves reuse or keeps UI details contained. For a tiny one-off script, the extra abstraction may not help; for a suite with repeated page interactions, it can reduce duplication.

Do I need Grid to run Selenium tests?

No. A local WebDriver session is a direct starting point. Grid is for remote sessions and distributed browser execution when your coverage or capacity needs justify it.

Can I use a different language?

Yes. Selenium provides language bindings; pair the binding with a test runner used by that language and keep the same boundaries between tests, page operations, data, and browser lifecycle.