Selenium 4 Capabilities: What They Are and How to Use Them
Learn what Selenium 4 capabilities do, configure browser Options for local and remote sessions, and avoid common Grid and page-load pitfalls.
Selenium 4 capabilities are name/value settings sent when WebDriver creates a browser session. They select the browser and platform, set standard session behavior such as page-load strategy and timeouts, and can include browser-specific or Grid/provider settings. In Selenium 4, configure these through the browser’s Options class and pass that object to the local or remote driver.
For example, Python uses ChromeOptions; Java uses ChromeOptions as well. A remote session also needs the WebDriver or Grid endpoint. Use standard W3C capability names for portable settings, and the documented namespace for custom settings.
1. What capabilities mean in Selenium 4
A capability is a requested property of a WebDriver session. During session creation, the client sends its requested configuration to the local driver or remote end. The remote end either creates a matching session or reports that it cannot satisfy the request. An Options object is Selenium’s language-binding API for assembling that browser’s session configuration.
Selenium’s upgrade guide describes Selenium 4 as using the W3C WebDriver standard and removing support for the legacy protocol. Its documentation says that browser Options classes must be used in Selenium 4. See Browser Options and Upgrade to Selenium 4.
| Setting family | What it controls | Example |
|---|---|---|
| Standard W3C capabilities | Portable session behavior and browser/platform requests | browserName, pageLoadStrategy |
| Browser-specific options | Features or startup arguments for one browser | Chrome command-line arguments |
| Selenium Grid settings | Grid routing, metadata, or Grid features | se:name, se:downloadsEnabled |
| Provider-specific settings | Cloud vendor configuration | A vendor namespace such as cloud:options |
These families are not interchangeable. A setting accepted by one cloud provider or browser may not be recognized elsewhere. Check the relevant provider’s current documentation for its exact key names and nesting.
2. Standard capabilities to know
| Capability | Purpose | Practical note |
|---|---|---|
browserName |
Selects the browser. | An Options instance normally sets this for its browser. |
browserVersion |
Requests a browser version, particularly useful for remote execution. | The remote environment must be able to provide it. Selenium Manager may download a missing browser version in supported environments. |
platformName |
Identifies the requested operating system/platform. | Commonly used by Grid or a cloud service for matching. |
acceptInsecureCerts |
Allows the session to accept insecure certificates. | Use only when the test needs to exercise an environment with such certificates. |
pageLoadStrategy |
Sets which document readiness state blocks navigation. | Choices are normal, eager, and none. |
timeouts |
Sets script, page-load, and implicit element-location timeouts. | Choose each value based on the test’s needs; these affect session-wide behavior. |
unhandledPromptBehavior |
Defines how an unhandled browser prompt is processed. | The documented default is dismiss and notify. |
proxy |
Configures proxy settings for browser traffic. | A proxy routes traffic; it does not by itself provide a complete capture or mocking system. |
The Selenium Options documentation lists default timeouts of 30,000 ms for scripts, 300,000 ms for page loads, and 0 for implicit element-location waits. These defaults are not a recommendation that every application use those values. Set timeouts deliberately and prefer explicit waits for application-specific readiness.
3. Page-load strategies and readiness
| Strategy | Navigation waits for | Use when |
|---|---|---|
normal |
The document reaches complete. This is the default. |
The test can proceed after ordinary document loading. |
eager |
The document reaches interactive. |
The test will explicitly wait for the particular content it needs. |
none |
No page-readiness condition. | The test intentionally controls all waiting after navigation. |
complete does not guarantee that a JavaScript-heavy single-page application has finished fetching data, rendering components, or enabling controls. If a test uses eager or none, it must wait for an application-specific condition before interacting. Changing the strategy without adding suitable waits often makes tests flaky rather than usefully faster.
4. Configure a local Selenium 4 session in Python
This example creates a local Chrome session, sets standard session behavior and a browser startup argument, opens a page, waits for a known element, and quits even if an operation fails. Install Selenium with python -m pip install selenium. Selenium Manager can configure drivers when enabled and supported by the environment.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
options = Options()
options.page_load_strategy = "eager"
options.accept_insecure_certs = True
options.add_argument("--headless")
options.add_argument("--window-size=1440,1000")
options.timeouts = {
"implicit": 0,
"pageLoad": 60_000,
"script": 30_000,
}
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com")
heading = WebDriverWait(driver, 15).until(
EC.visibility_of_element_located((By.TAG_NAME, "h1"))
)
print(heading.text)
finally:
driver.quit()
The argument --headless is a Chrome startup option, not a portable W3C capability. Browser arguments vary by browser and environment. For a headed local run, remove the headless argument. Do not copy arguments between browsers without checking that browser’s documentation.
Python remote session
For Grid or another remote WebDriver endpoint, keep the Options object and provide the endpoint URL. The requested browser and any requested version or platform must be available remotely.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
options = Options()
options.page_load_strategy = "normal"
options.set_capability("browserVersion", "stable")
# Replace with the URL of your Selenium Grid or remote WebDriver endpoint.
driver = webdriver.Remote(
command_executor="http://localhost:4444",
options=options,
)
try:
driver.get("https://example.com")
print(driver.title)
finally:
driver.quit()
Use a concrete version and platform value supported by your remote infrastructure when you need deterministic matching. A request for an unavailable combination cannot be satisfied merely by putting it in capabilities.
5. Configure capabilities in Java
With Selenium’s Java binding, pass the browser Options instance to the driver constructor. This local Chrome example uses the same concepts as the Python example. Add Selenium to the project using the dependency setup appropriate to your build system.
import java.time.Duration;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
public class CapabilitiesExample {
public static void main(String[] args) {
ChromeOptions options = new ChromeOptions();
options.setPageLoadStrategy(org.openqa.selenium.PageLoadStrategy.EAGER);
options.setAcceptInsecureCerts(true);
options.addArguments("--headless");
options.addArguments("--window-size=1440,1000");
options.setImplicitWaitTimeout(Duration.ZERO);
options.setPageLoadTimeout(Duration.ofSeconds(60));
options.setScriptTimeout(Duration.ofSeconds(30));
WebDriver driver = new ChromeDriver(options);
try {
driver.get("https://example.com");
String text = new WebDriverWait(driver, Duration.ofSeconds(15))
.until(ExpectedConditions.visibilityOfElementLocated(By.tagName("h1")))
.getText();
System.out.println(text);
} finally {
driver.quit();
}
}
}
For a remote session, use new RemoteWebDriver(remoteUrl, options) with your Grid endpoint and a compatible ChromeOptions object. Java binding APIs evolve; use the current Selenium API for your installed version.
6. Custom capabilities, browser options, and namespaces
Use the Options API’s capability-setting method for a standard setting that is not exposed as a dedicated convenience property in your binding. Keep the spelling and value types aligned with the WebDriver standard. In Python, for example:
from selenium.webdriver.chrome.options import Options
options = Options()
options.set_capability("browserVersion", "stable")
options.set_capability("platformName", "linux")
options.set_capability("acceptInsecureCerts", True)
For browser-specific configuration, use that browser’s Options API. For example, Chrome arguments belong in ChromeOptions. Such settings are not guaranteed to work in Firefox, Safari, or a remote provider’s environment.
For a cloud provider, place custom keys in the vendor-prefixed namespace and structure prescribed by that provider. Selenium’s upgrade guide illustrates provider values such as build and name inside a namespace like cloud:options. That illustrates the namespace rule; it is not a universal cloud configuration. Confirm the actual namespace, keys, and accepted values with your provider.
Legacy names such as version and platform should be migrated to standard names browserVersion and platformName. In C#, the Selenium upgrade guide calls out AddAdditionalOption in place of the old AddAdditionalCapability. Avoid combining legacy Desired Capabilities patterns with Selenium 4 Options objects.
7. Selenium Grid matching and metadata
Grid routes a requested session to an available node whose advertised configuration matches the request. A requested browser, version, platform, or custom capability must be supported by the Grid’s nodes. Selenium Grid’s documentation covers starting a standalone server and connecting RemoteWebDriver to its URL: Getting started with Selenium Grid.
Grid supports custom capability matching when the custom values are configured on relevant nodes and included in each matching session request. A custom key sent only by the client does not make an unconfigured node eligible. Selenium’s Grid metadata uses the se: prefix; se:name can label a test in the Grid UI.
Managed downloads are an example of a feature requiring both sides of the configuration: the node must be configured for managed downloads, and the session request must set se:downloadsEnabled. Consult the current Grid CLI options documentation for node flags and matching details.
8. Or skip the browser setup
If the goal is a website screenshot rather than browser automation, ScreenshotNeo provides a screenshot API and MCP server. It accepts a URL in one GET request and returns an image or PDF. The API docs are at ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners, popups, and chat widgets are removed before the shot, and each step can be turned off. Bot checks, blank pages, and failed loads are never billed; response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo for the service and the API documentation for options.
Sign up for 1,000 free screenshots a month, with no card required.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Session creation fails with an unsupported capability error | A legacy key, misspelled standard name, or unprefixed custom key was sent. | Use Selenium 4 Options, standard W3C names, and the namespace required for custom settings. |
| Remote session cannot find a matching node | The requested browser, version, platform, or custom capability does not match any available node. | Check node stereotypes and advertised values; adjust the request or configure the relevant nodes. |
| Navigation returns before app content appears | The page-load strategy only waits for document readiness, not application-specific data. | Wait explicitly for the element, state, or data the test needs. |
| Navigation hangs or times out | The page is slow, never reaches the chosen readiness state, or a timeout is too short. | Inspect the navigation and network behavior, choose an appropriate page-load strategy, and set a deliberate page-load timeout. |
| Element lookup fails immediately | Implicit wait is zero, or the element is not present/visible yet. | Use an explicit wait for the required condition; verify selector and frame context. |
| Browser ignores a startup argument | The argument is browser-specific, unsupported, or passed through the wrong Options class. | Use the matching browser Options API and verify the argument against current browser documentation. |
| Old code fails after Selenium 4 upgrade | It constructs a driver with legacy Desired Capabilities or an outdated driver-path API. | Migrate to browser Options. In Python, use a Service object instead of deprecated executable_path construction; see the upgrade guide. |
10. Performance, reliability, and cost considerations
- Page-load strategy:
eagerornonecan reduce time spent waiting for document readiness, but only if the test waits for the application state it uses. Measure your own suite; there is no universal best setting. - Explicit waits: Wait for meaningful conditions instead of sleeping for a fixed interval. This avoids waiting longer than needed and reduces timing assumptions.
- Remote execution: Grid or cloud execution adds a network hop and depends on capacity and node matching. Keep requested versions and custom capabilities aligned with available infrastructure.
- Reliability: Always close sessions with
quit()in a cleanup path. Keep browser-specific options isolated and avoid relying on undocumented provider keys. - Cost: Selenium itself does not define a cloud provider’s session price. Remote browser infrastructure may have its own pricing and limits; check the provider’s current terms. Local execution uses your own compute resources.
11. FAQ
Are Desired Capabilities still the right Selenium 4 API?
Use the browser Options class. “Desired Capabilities” is a legacy term and old construction patterns can fail under Selenium 4’s W3C-based session model.
Does every capability work with every browser?
No. Standard W3C settings are intended to be shared, but browsers add their own options, and remote providers define their own namespaces and values.
Can Selenium Manager select a browser version?
Recent Selenium Manager versions can download a browser version not found locally in supported conditions. Remote Grid matching still depends on what the remote infrastructure offers.
Should I set an implicit wait and explicit waits together?
Prefer a consistent wait strategy. This guide sets implicit wait to zero and uses explicit waits so the condition being awaited is visible in the test.


