ScreenshotNeo

BlogHow-to

How to Use the Page Object Model in Selenium with Python

Learn to organize Selenium tests with Python page objects, explicit waits, reusable components, and clear assertions—with a runnable login example.

By the ScreenshotNeo team4 October 20269 min read

The Page Object Model (POM) organizes Selenium tests by giving each meaningful page or reusable UI region a Python class. The class owns the locators and browser interactions for that area; tests call its user-level methods and assert the expected outcome. This keeps selectors and interaction details in one place, so a UI change is often easier to fix. Use explicit waits for the specific UI condition a test needs.

What a page object should do

A page object is an interface to part of your application, not another test case. It should know how to find and operate the elements it owns, and expose actions that make sense to a user, such as login_as() or search_for().

  • Page object: owns page-specific locators and useful interactions. It may make a narrow check that the expected page loaded.
  • Test: arranges the scenario and asserts business or acceptance outcomes.
  • Component object: represents a substantial or repeated region, such as a navigation menu, when it has meaningful behavior of its own.

This division avoids duplicated UI knowledge while keeping behavioral assertions visible in the test. Selenium’s Page Object Models guidance describes this separation and the use of page components.

Set up Selenium with Python

Use a virtual environment and install the Selenium Python package:

python -m venv .venv
# macOS or Linux
source .venv/bin/activate
# Windows PowerShell
.venv\Scripts\Activate.ps1
python -m pip install selenium

Recent Selenium releases can manage supported browser drivers through Selenium Manager. Install a compatible browser, then run the example below. In a managed CI environment, install the browser and driver using that environment’s documented setup and make the driver available on PATH.

Build a login page object

The example assumes a test application at https://example.test/login with fields identified by username and password, a submit button with ID submit, and a successful-login heading with ID welcome. Replace those selectors and the URL with your application’s stable markup.

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


class LoginPage:
    URL = "https://example.test/login"
    USERNAME = (By.ID, "username")
    PASSWORD = (By.ID, "password")
    SUBMIT = (By.ID, "submit")

    def __init__(self, driver):
        self.driver = driver
        self.wait = WebDriverWait(driver, 10)
        self.driver.get(self.URL)
        self.wait.until(EC.visibility_of_element_located(self.USERNAME))

    def login_as(self, username, password):
        self.wait.until(EC.visibility_of_element_located(self.USERNAME)).send_keys(username)
        self.driver.find_element(*self.PASSWORD).send_keys(password)
        self.wait.until(EC.element_to_be_clickable(self.SUBMIT)).click()
        return HomePage(self.driver)


class HomePage:
    WELCOME = (By.ID, "welcome")

    def __init__(self, driver):
        self.driver = driver
        self.wait = WebDriverWait(driver, 10)

    def welcome_message(self):
        element = self.wait.until(EC.visibility_of_element_located(self.WELCOME))
        return element.text

The constructor navigates to the login URL and waits for a page-specific element, a narrow readiness check. The login_as() method describes a user action and returns the page expected after successful login. It does not assert that login succeeded; the test owns that assertion.

Write a test that uses the page object

This standalone script creates a browser, exercises the page object, checks the result, and quits the browser even if the assertion fails. Set credentials through environment variables in real projects rather than committing secrets.

import os

from selenium import webdriver

from pages.login_page import LoginPage


def test_user_can_log_in():
    driver = webdriver.Chrome()
    try:
        home = LoginPage(driver).login_as(
            os.environ["TEST_USERNAME"],
            os.environ["TEST_PASSWORD"],
        )
        assert home.welcome_message() == "Welcome"
    finally:
        driver.quit()


if __name__ == "__main__":
    test_user_can_log_in()
    print("Login flow passed")

Save the page classes in pages/login_page.py and the script as test_login.py. Run it after setting TEST_USERNAME and TEST_PASSWORD. To use pytest, install it and let pytest discover the test function; keep the same setup and cleanup or move browser creation into a fixture.

Choose locators that survive UI changes

Keep each locator beside the page or component that owns the element. Prefer stable IDs, names, or test-specific attributes supplied by the application. CSS selectors can be concise and robust when based on stable attributes. XPath can express relationships when needed, but long paths tied to DOM layout tend to break when markup changes.

Selenium supports ID, name, CSS selector, link text, partial link text, class name, tag name, and XPath strategies. Choose for clarity and stability in the actual application; no strategy is best in every case. See the official locator strategies reference.

# Stable test attribute, if the application provides it
SAVE_BUTTON = (By.CSS_SELECTOR, '[data-testid="save-profile"]')

# Accessible link text, when the visible label is stable
HELP_LINK = (By.LINK_TEXT, "Help")

# Relationship-based XPath, useful when no stable attribute exists
DELETE_FOR_ROW = (By.XPATH, ".//tr[.//td[normalize-space()='Sample']]//button[@aria-label='Delete']")

Prefer a locator that communicates what element the test needs, rather than one that encodes incidental layout. If your team can change the application, adding stable test attributes can make browser tests less sensitive to visual refactoring.

Wait for the condition the next action needs

Modern pages often update asynchronously. A navigation call returning does not guarantee that JavaScript has rendered the next control. Selenium identifies these races as a common source of flaky tests in its waiting strategies documentation.

Use WebDriverWait with a condition tied to the next step:

wait = WebDriverWait(driver, 10)

# Element exists in the DOM
field = wait.until(EC.presence_of_element_located((By.ID, "search")))

# Element is visible to the user
message = wait.until(EC.visibility_of_element_located((By.ID, "results")))

# Element can be clicked
wait.until(EC.element_to_be_clickable((By.ID, "continue"))).click()

# URL changed after a navigation action
wait.until(EC.url_contains("/dashboard"))

Presence does not imply visibility, and visibility does not necessarily mean an element is enabled. Select the condition that matches the action. Avoid fixed sleeps as the normal synchronization method: they can waste time when the page is fast and still fail when it is slow. Keep a consistent wait policy and avoid casually mixing implicit and explicit waits because their timing can interact.

Organize pages and reusable components

A small project can keep classes in one module. As the suite grows, one module per page makes ownership easier to see:

project/
  pages/
    __init__.py
    login_page.py
    home_page.py
    components/
      __init__.py
      main_navigation.py
  tests/
    test_login.py

Do not create a class for every fragment. Extract a component when it is repeated or has a coherent set of interactions. For example, if multiple pages use a navigation menu with its own links and actions, give it a component object and compose it into the relevant page.

from selenium.webdriver.common.by import By


class MainNavigation:
    ACCOUNT_LINK = (By.CSS_SELECTOR, "nav a.account")

    def __init__(self, driver):
        self.driver = driver

    def open_account(self):
        self.driver.find_element(*self.ACCOUNT_LINK).click()

Locator classes or modules are optional. Selenium’s Python tutorial demonstrates separating locators, but that is one possible organization. Keep UI knowledge close enough that someone changing a page can find its selectors and operations without searching through unrelated tests.

Common design choices

Choice Use it when Watch for
One class per page The page has a distinct role and interactions. A single site-wide class becomes hard to navigate.
Component object A region is reused or has meaningful behavior of its own. Tiny one-off wrappers add indirection without reuse.
Locators on the page class The page is small or selectors are easiest to read beside methods. Do not scatter selectors across tests.
Separate locator class/module A larger page benefits from a clear locator inventory. Do not separate so far that ownership becomes unclear.
Return next page object An action has a clear expected destination, such as successful login. Do not make the page object verify the business outcome.
Return observable state The test needs text, a value, or another result to assert. Keep the assertion itself in the test.

Troubleshooting

Symptom Likely cause Fix
NoSuchElementException The locator is wrong, the element has not rendered, or the test is on a different page. Verify the current URL and markup; use a wait for presence or visibility when rendering is asynchronous; choose a stable locator.
TimeoutException The waited-for condition never became true, the selector is incorrect, or the application is slower than the timeout. Inspect the page state and locator first. Increase the timeout only when the observed application behavior justifies it.
ElementClickInterceptedException An overlay, animation, or sticky element covers the target. Wait for the overlay to disappear or for the target to become clickable; handle the actual UI state rather than forcing a JavaScript click by default.
Flaky test after navigation The test assumes navigation completion means the dynamic content is ready. Wait for the next page’s meaningful element or URL condition.
Driver or browser startup error The browser is missing, incompatible, or unavailable in the runtime environment. Install a supported browser and ensure Selenium Manager can resolve the driver, or configure the environment’s browser driver explicitly.
Wait takes unexpectedly long Implicit and explicit waits are both active, or a broad condition is being polled. Use one deliberate wait policy and wait on the specific state needed by the next operation.
Credentials missing from the test Environment variables were not set. Set TEST_USERNAME and TEST_PASSWORD in the local shell or CI secret configuration; do not place real credentials in source control.

Performance, reliability, and maintenance

  • Wait precisely: a condition wait proceeds as soon as the condition is satisfied, while a fixed delay always consumes its full duration.
  • Keep methods task-level: methods such as submit_order() communicate intent better than wrappers that only rename find_element().
  • Keep assertions observable: return text, state, or the next page object so the test can make the expected outcome explicit.
  • Clean up reliably: quit the driver in a finally block or a test fixture teardown so failures do not leave browser processes behind.
  • Limit coupling: avoid methods that silently depend on many unrelated page details. A page object should model its page or component, not encode the whole test workflow.
  • Control test data: use known accounts and resettable data so a passing result does not depend on stale state from a prior run.

POM does not make browser automation inherently fast or eliminate failures caused by application state, network delays, or unstable markup. It makes UI knowledge easier to maintain and test behavior easier to read when the boundaries stay focused.

Or skip the browser setup

If your goal is to capture a website image or PDF rather than interactively test its behavior, ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with the page verdict and billing status returned in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for MCP clients such as Claude and Cursor. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, no card required.

FAQ

Should every page object verify that it loaded?

A narrow readiness check, such as waiting for a page-specific element in the constructor, is reasonable. Keep assertions about ordinary test outcomes in the test.

Can a page-object method return another page object?

Yes. Returning the expected destination page after an action can make a workflow readable. Returning the current page or a component is also reasonable when that better reflects the interaction.

Do I need a separate locator class?

No. Use one when it improves clarity for your project; keeping locators directly on the page class is also valid.

Is Page Object Model a Selenium requirement?

No. It is an organization pattern for reducing duplicated UI knowledge and making tests easier to maintain, especially when multiple tests interact with the same pages.