ScreenshotNeo

BlogHow-to

How to Use the Page Object Model with Selenium and JavaScript

Organize Selenium tests with JavaScript page objects: install the WebDriver package, model pages and components, handle waits and assertions, and troubleshoot common failures.

By the ScreenshotNeo team4 October 20268 min read

The Page Object Model (POM) organizes Selenium tests by putting each page’s locators and user-facing operations in a page object. Tests call those operations and assert the results. When the UI changes, you can often update the page object instead of changing every test that used its selectors.

This guide uses Selenium’s JavaScript binding, selenium-webdriver. The official Selenium JavaScript API reference specifies Node.js 22 or newer. Its Page Object Model guidance uses mainly Java examples, so the JavaScript code below applies the same design principles using JavaScript APIs. [Selenium Page Object Models; Selenium JavaScript API]

1. Install Selenium and choose where the browser runs

Start with Node.js 22 or newer and install the WebDriver package:

mkdir selenium-pom-example
cd selenium-pom-example
npm init -y
npm install selenium-webdriver

Use ES modules by adding "type": "module" to package.json, or use CommonJS consistently. The runnable example below uses ES modules. Selenium Manager can handle browser-driver installation for local runs. For remote runs, point Selenium at a Grid or standalone server with SELENIUM_REMOTE_URL or Builder.usingServer(). See the JavaScript API reference for current runtime and setup details.

2. Create page objects around user-facing operations

Here is a complete example with a login page, a destination page, and a test. Save it as login.test.js. It uses Node’s built-in test runner and assertion library, so it does not require another test framework. Replace the example URL, selectors, and expected heading with values from your application.

import assert from 'node:assert/strict'
import test from 'node:test'
import { Builder, By, until } from 'selenium-webdriver'

class HomePage {
  constructor(driver) {
    this.driver = driver
    this.heading = By.css('h1')
  }

  async waitUntilLoaded() {
    await this.driver.wait(until.elementLocated(this.heading), 10_000)
    return this
  }

  async headingText() {
    return this.driver.findElement(this.heading).getText()
  }
}

class LoginPage {
  constructor(driver) {
    this.driver = driver
    this.username = By.name('username')
    this.password = By.name('password')
    this.submit = By.css('button[type="submit"]')
  }

  async open() {
    await this.driver.get('https://example.test/login')
    await this.driver.wait(until.elementLocated(this.username), 10_000)
    return this
  }

  async signIn(username, password) {
    await this.driver.findElement(this.username).sendKeys(username)
    await this.driver.findElement(this.password).sendKeys(password)
    await this.driver.findElement(this.submit).click()

    const home = new HomePage(this.driver)
    await home.waitUntilLoaded()
    return home
  }
}

test('a user can sign in', async () => {
  const driver = await new Builder().forBrowser('chrome').build()

  try {
    const login = await new LoginPage(driver).open()
    const home = await login.signIn('reader', 'example-password')
    assert.equal(await home.headingText(), 'Welcome')
  } finally {
    await driver.quit()
  }
})

Run it with:

node --test login.test.js

The example keeps selectors inside their page objects, names methods after actions, returns the page reached by a successful transition, waits for an element instead of sleeping a fixed duration, and closes the browser in finally. The assertions remain in the test.

3. Decide what belongs in a page object

Keep locators and page operations together

Store selectors as By values on the page object and use methods such as signIn(), searchFor(), or addItemToCart() to express user actions. This gives tests a stable interface even if the page’s CSS or markup changes. Prefer stable application attributes, accessible names, or meaningful form labels where available; selectors tied to incidental layout or generated class names tend to require more maintenance.

Keep scenario assertions in the test

A test should state what outcome matters, such as the heading text or a confirmation message. Selenium’s guidance says, “Page objects themselves should never make verifications or assertions,” while allowing a narrow check that the page and critical elements loaded correctly during construction. [Selenium Page Object Models]

In practice, a page object may offer observations such as headingText() or isErrorVisible(). The test compares those observations with the expected outcome. Avoid embedding test-specific expectations such as “this user must see this exact message” in a reusable page object.

Represent navigation and multiple outcomes clearly

When an action leads to another page, return that page object after waiting for a useful readiness signal. If the same action can lead to different outcomes, make the test inspect the resulting state or provide clearly named operations for distinct expected paths. For example, a rejected login may remain on the login page and display an error, while a successful login reaches a home page.

Use component objects for repeated regions

A navigation bar, product card, or repeated table row can be a component object when it has reusable behavior of its own. Locate the component’s root element, then scope child lookups to that element so selectors cannot accidentally match a different component elsewhere on the page.

import { By } from 'selenium-webdriver'

class ProductCard {
  constructor(root) {
    this.root = root
    this.name = By.css('.product-name')
    this.addButton = By.css('button.add-to-cart')
  }

  async getName() {
    return this.root.findElement(this.name).getText()
  }

  async addToCart() {
    await this.root.findElement(this.addButton).click()
  }
}

class CatalogPage {
  constructor(driver) {
    this.driver = driver
    this.cards = By.css('[data-testid="product-card"]')
  }

  async productCards() {
    const roots = await this.driver.findElements(this.cards)
    return roots.map(root => new ProductCard(root))
  }
}

Selenium’s JavaScript API supports finding descendants from a WebElement, which makes this scoped component lookup possible. [Selenium WebElement API]

4. Handle waits and changing page state

Modern pages often render or update asynchronously. Prefer waiting for the condition the test needs over pausing for a fixed number of seconds. A fixed delay may waste time when the page is fast and still fail when it is slow.

import { By, until } from 'selenium-webdriver'

const confirmation = By.css('[role="status"]')
await driver.wait(until.elementLocated(confirmation), 10_000)
const message = await driver.findElement(confirmation).getText()

Choose a signal tied to the next operation: an element appearing, becoming visible, or a URL changing. A page object can provide a method such as waitUntilLoaded() to keep that page’s readiness knowledge beside its locators. Use a timeout that matches the application and environment, and include enough context in failures to tell which transition did not complete.

A located element can later become stale if a framework replaces it during rendering. In that case, wait for the replacement condition and locate the element again rather than reusing a stale WebElement. Keep page objects focused: they should help tests interact with the UI, not become a general-purpose framework that hides the browser’s behavior.

5. Run against a remote Selenium server

Local execution is useful for a small setup; remote execution moves the browser session to a Selenium Grid or standalone server. Selenium’s JavaScript API documents both usingServer() and the SELENIUM_REMOTE_URL environment variable. [Selenium JavaScript API]

For an explicit server URL, build the driver like this:

const remoteUrl = process.env.SELENIUM_REMOTE_URL
if (!remoteUrl) throw new Error('Set SELENIUM_REMOTE_URL to your Selenium server URL')

const driver = await new Builder()
  .forBrowser('chrome')
  .usingServer(remoteUrl)
  .build()

Keep the page objects unchanged: they receive a WebDriver instance, whether the browser is local or remote. Configure the remote endpoint and any browser capabilities according to the Selenium server you operate. A remote session adds network and server availability to the test’s dependencies, so preserve cleanup in finally even when navigation or an assertion fails.

6. Common failures and fixes

Symptom Likely cause What to check
Node rejects the package or syntax The runtime is older than the binding’s requirement, or module formats are mixed. Use Node.js 22 or newer per the current API reference. Set "type": "module" for the example above, or convert imports to CommonJS consistently.
Browser or driver cannot start The browser is missing, unsupported, or unavailable to the local setup. Check the installed browser and consult Selenium’s current setup documentation. Selenium Manager handles driver installation in the documented quick start; confirm the environment permits the required setup.
Remote connection fails The remote URL is unset, unreachable, or points to a server that is not accepting sessions. Check SELENIUM_REMOTE_URL, network access, server status, and browser configuration.
Element lookup times out The selector is wrong, the page is not ready, or the element is inside a frame or shadow root. Inspect the current DOM and selector, wait for the right condition, and switch to the relevant browsing context where needed.
Click has no effect or hits an overlay An animation, modal, consent banner, or other element is covering the target. Wait for the overlay to disappear or interact with the intended visible control. Avoid clicking by coordinates when a normal element interaction is possible.
Stale element reference The page replaced a node after the element was found. Wait for the updated state and find the element again; do not retain element handles across rerenders.
Browser stays open after a failure Cleanup is skipped when a test throws. Put await driver.quit() in a finally block around the test’s browser work.

7. Performance, reliability, and maintenance

  • Wait on conditions. This avoids unnecessary fixed delays while keeping tests bounded by an explicit timeout.
  • Keep browser sessions isolated. Build and quit a driver for each test or test group according to the runner’s lifecycle, and avoid sharing mutable browser state across unrelated scenarios.
  • Keep objects small. A page object should expose a page’s useful actions and observations; use a component object when a repeated region has its own behavior.
  • Reduce selector churn. Centralizing selectors means UI changes generally touch the relevant page or component object instead of many tests.
  • Account for remote costs. Remote browsers add network latency and infrastructure use. Batch related interactions into meaningful page operations, but keep failures diagnosable and avoid hiding unexpected states.
  • Always release sessions. Cleanup in finally prevents failed assertions from leaving browser sessions behind.

8. Or skip the browser setup

If your task is to capture a page image or PDF rather than interact with a browser session, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. See the API documentation for parameters and response details.

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The API also has cURL and Python examples in the documentation. Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. An MCP server lets AI agents use screenshot, page information, and PDF capture tools. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.

FAQ

Should every page have its own class?

Create page objects where they clarify a real flow or centralize meaningful behavior. Avoid classes that only wrap one selector without improving reuse or readability.

Can a page object represent a modal or widget?

Yes. A distinct reusable region can be modeled as a component object, especially when it has its own actions and can appear on more than one page.

Where should expected text and business rules live?

Keep scenario-specific expected results in tests. Page objects expose actions and observations so tests can express those expectations clearly.

Can the same page objects run locally and on Grid?

Yes. They operate on the WebDriver passed to them. The driver builder and server configuration determine where the browser session runs.