WebdriverIO Browser Commands: A Practical Tutorial
Learn what WebdriverIO’s browser object controls, how to navigate, inspect pages, manage windows, send input, and wait for the state your test needs.
WebdriverIO browser commands operate on the active browser or mobile session. In a WebdriverIO test-runner project, use the runner-managed global browser object; in a standalone script, get a session object from remote. Use browser-level commands for navigation, page state, windows, timeouts, and scripts, and element-level commands for interacting with a particular element. The exact commands available can depend on the driver backend and environment.
This tutorial covers current WebdriverIO documentation for 8.x and later. Check the documentation for your installed version and selected backend when a command or input type behaves differently.
1. What is the WebdriverIO browser object?
browser represents the active automation session. It is not a browser installation or a physical device. In the test runner, WebdriverIO starts and ends the session as part of the runner lifecycle. In standalone usage, your code creates a session and should close it when finished.
WebdriverIO commands have two broad layers:
- Protocol commands bind more directly to the underlying WebDriver protocol, such as navigation and window operations.
- Convenience commands provide a higher-level API on objects such as
browserandelement.
That distinction helps explain command scope. Use browser for session-wide work; use the element object returned by a selector for work on a page element. Backend-specific commands may also be available, so do not assume every browser session exposes an identical surface. See the WebdriverIO API introduction and the browser object reference.
2. Set up a test-runner example
In a WebdriverIO test-runner test, use the runner-managed browser object. Do not create a second session for every test unless your project specifically requires that lifecycle. Here is a small asynchronous test body; place it in a spec file configured for your project’s runner and browser capabilities:
describe('page navigation', () => {
it('opens a page and checks its location and title', async () => {
await browser.url('https://webdriver.io/');
const currentUrl = await browser.getUrl();
const title = await browser.getTitle();
console.log({ currentUrl, title });
});
});
The runner provides the session and handles its lifecycle. The example logs state so it can run without assuming a particular assertion library. Add the assertion syntax supplied by your project’s test framework when you want the test to fail on a mismatch.
3. Use browser commands in a standalone script
For standalone automation, create the session with remote, perform commands, and close the session in a finally block so it is also closed after an error. This example uses a local WebDriver endpoint; configure its address and capabilities for the driver you run.
import { remote } from 'webdriverio';
const browser = await remote({
hostname: 'localhost',
port: 4444,
capabilities: {
browserName: 'chrome'
}
});
try {
await browser.url('https://webdriver.io/');
console.log('URL:', await browser.getUrl());
console.log('Title:', await browser.getTitle());
} finally {
await browser.deleteSession();
}
The driver must be reachable at the configured endpoint, and it must support the requested capabilities. The session’s available commands and behavior can vary with the backend.
4. Navigate and inspect the current page
Use browser.url(url) as the convenient navigation method. The protocol reference also documents navigateTo(url), which navigates to a URL. Read the current location with getUrl() and the document title with getTitle().
await browser.url('https://example.com/account');
const url = await browser.getUrl();
const title = await browser.getTitle();
console.log({ url, title });
A returned URL or title confirms that the command completed and provides a useful assertion point. It does not prove that a client-rendered page has finished loading data or that a specific control is ready. Wait for the condition your next step actually needs, such as the presence or visibility of a target element.
5. Move through browser history and refresh
Browser-level history commands let a test move backward or forward in the active browsing context. Use back(), forward(), and refresh(), then inspect or wait for the resulting page state.
await browser.url('https://example.com/first');
await browser.url('https://example.com/second');
await browser.back();
console.log('After back:', await browser.getUrl());
await browser.forward();
console.log('After forward:', await browser.getUrl());
await browser.refresh();
console.log('After refresh:', await browser.getUrl());
History behavior can depend on redirects and how the application changes location. If the URL is not a stable indicator for a single-page application, wait for a page-specific element or state instead.
6. Work with windows and tabs
Window handles identify open browsing contexts. Retrieve them, switch to the handle you intend to inspect, and only then make assertions. A safe pattern is to remember the current handle before an action that may open another context, then identify the new handle afterward.
const originalHandle = await browser.getWindowHandle();
const handlesBefore = await browser.getWindowHandles();
// Perform the page action that opens another window or tab here.
const handlesAfter = await browser.getWindowHandles();
const newHandle = handlesAfter.find(
(handle) => !handlesBefore.includes(handle)
);
if (!newHandle) {
throw new Error('No new browsing context appeared');
}
await browser.switchToWindow(newHandle);
console.log('New context URL:', await browser.getUrl());
await browser.switchToWindow(originalHandle);
console.log('Original context URL:', await browser.getUrl());
The snippet deliberately leaves the opening action to your test because the mechanism is application-specific. Do not assume the current context is the one that just opened; identify and switch to it explicitly. The WebDriver protocol reference documents window handles and switching operations. Confirm support and exact behavior with the API reference for your installed version and backend.
7. Send keyboard, pointer, or wheel input
For ordinary test interactions, prefer the higher-level interaction commands provided by WebdriverIO. When you need to compose a low-level keyboard, pointer, or wheel sequence, use browser.action() and finish the chain with perform(). Support can vary by environment, driver, and input type.
await browser.action('key')
.down('CTRL')
.up('CTRL')
.perform();
This illustrates the action-chain shape, but a key sequence is useful only when it matches the application and environment. Consult the browser action reference for supported action types and parameters. If an action is unsupported, use the higher-level interaction API appropriate to the element and verify behavior with your selected backend.
8. Wait for page state and set timeouts deliberately
A navigation command completing is not the same as the application being ready for your next action. Prefer a condition-based wait that expresses the page state the test depends on. For example, wait for a known element using the element API:
await browser.url('https://example.com/account');
const heading = await $('h1');
await heading.waitForDisplayed();
console.log('Heading text:', await heading.getText());
Use a condition that is meaningful for the page: an element displayed, a loading indicator gone, or an application-specific state reached. Avoid arbitrary delays when a state check is available; delays add time when the page is fast and can still be too short when it is slow.
The protocol supports session timeout configuration. Be cautious with implicit timeouts: the current WebdriverIO timeout documentation does not recommend them because they can affect other WebdriverIO commands. Set timeouts to address a known operation or test condition, and consult the current protocol reference and your installed version’s timeout API for exact signatures.
9. Execute JavaScript in the page
Use the browser script-execution command for a small page-side read or operation that the normal element API does not express. Keep browser interaction in WebdriverIO where possible so the test remains clear about what it observes.
const pageInfo = await browser.execute(() => ({
title: document.title,
url: window.location.href
}));
console.log(pageInfo);
The function runs in the page context. Only return values that can be serialized across the WebDriver boundary, and remember that page-side script execution does not establish that asynchronous application work is complete. Use an explicit wait for that.
10. Choose the command scope that matches the task
| Task | Typical scope | Examples |
|---|---|---|
| Navigate or inspect session state | browser |
url, getUrl, getTitle |
| History, windows, or session timeout | browser |
back, forward, refresh, window handles |
| Interact with a particular control | Element object | Find an element, wait for it, then use its interaction commands |
| Compose low-level input | browser action builder |
action(...) followed by perform() |
| Extend a session API | browser custom command |
addCommand or, for advanced cases, overwriteCommand |
Custom commands can make repeated project-specific behavior easier to reuse, but they are an extension point rather than a requirement for basic automation. See the browser object reference before adding or replacing commands.
11. Common errors and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
browser is undefined |
The code is running outside a runner-managed test, or the runner globals are not configured. | Use the runner’s configured global or import the supported globals for your setup. For standalone use, create the object with remote. |
| Could not connect to the WebDriver server | The endpoint is wrong, the driver is not running, or the host/port is unreachable. | Start or locate the driver service and match hostname and port to its endpoint. |
| Session creation fails | The requested capabilities do not match the installed browser or driver backend. | Check the backend’s supported capabilities and the browser/driver configuration. |
| Command is not available | The command is unsupported by the selected backend, environment, or installed WebdriverIO version. | Check the current API reference for that version and use a supported command path. |
| Action chain does not work | The input type may not be supported in that environment, or the chain was not dispatched. | Call perform(), verify the action reference, and try a higher-level interaction where suitable. |
| Page assertion runs too early | The navigation finished before the application’s asynchronous state was ready. | Wait for a page-specific element or condition before reading or asserting state. |
| Window assertion checks the wrong tab | The test did not switch to the intended window handle. | Compare handles before and after the opening action, then call switchToWindow with the intended handle. |
| Standalone browser remains open after failure | Session cleanup was skipped when an earlier command threw. | Put deleteSession() in a finally block. |
12. Performance, reliability, and cost considerations
Browser commands wait on a remote session and page behavior, so avoid doing redundant navigations or fixed sleeps. Wait for the specific condition needed, keep the session lifecycle clear, and always close standalone sessions. These practices reduce avoidable waiting and make failures easier to diagnose; actual execution time depends on the page, driver, and environment.
For reliability, keep checks tied to stable page state rather than assuming that a URL or title means all client-side work has finished. Keep backend-specific assumptions visible in the test, and consult the command reference when changing drivers. The research sources provide no benchmark or fixed timing guarantee, so performance should be measured in your own environment.
WebdriverIO itself is an open-source automation framework; session infrastructure and browser execution costs depend on where and how you run the driver. This tutorial makes no claim about a particular provider’s price.
Or skip the browser setup
If your task is to save a page image or PDF rather than automate an interactive browser session, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its API documentation covers the available parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers say which verdict applied and whether the shot was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Try it with the free ScreenshotNeo sign-up.
Frequently asked questions
Does browser mean the browser installed on my computer?
No. It is the WebdriverIO object representing the active automation session.
Should I use browser commands for every page interaction?
No. Browser commands manage session-wide work; use element-scoped APIs to interact with particular page controls.
Why does a command work with one driver and fail with another?
The command surface can depend on the backend and environment. Verify support in the current API reference for your setup.
Does a successful navigation mean my page is ready?
Not necessarily. Wait for the element or application state required by the next step.


