Selenium 4 Observability: Monitoring Grid and Test Sessions
Monitor Selenium Grid’s live capacity with GraphQL and trace failed requests across Grid components with OpenTelemetry.
Selenium Grid observability combines two views: use the Grid UI, /status, or GraphQL to inspect current capacity, nodes, slots, and sessions; use OpenTelemetry traces to follow a request across Grid components and find where it slowed down or failed. Live state tells you what is happening now. A trace explains the path and timing of a particular request.
This guide covers Selenium 4’s built-in observability, practical commands and queries, session labeling, trace collection, version checks, and troubleshooting. Exact options can vary by Selenium Server version and deployment topology, so check the running version’s configuration help before applying settings.
1. Choose the right observability view
| Question | Use | What it shows |
|---|---|---|
| Is the Grid responding? | Grid UI or /status |
Quick health and status information. |
| How much capacity is in use? | GraphQL | Counts, nodes, slots, and session details in structured form. |
| Where did this request fail or spend time? | OpenTelemetry trace | The request journey, timed spans, and timestamped events. |
| How do I inspect traces locally? | FINE-level console logging | Trace and event details in server logs. |
| How do I query traces centrally? | A trace backend such as Jaeger | Collected traces that can be searched, filtered, and visualized. |
Selenium Server is instrumented with OpenTelemetry tracing. A trace represents a request journey, spans represent timed operations, and events add timestamped context within spans. The built-in tracing documentation describes tracing as enabled by default; console trace and event output is visible at FINE log level. See the Selenium Project’s Observability in Selenium Grid guide.
2. Check Grid health and live sessions
Use the Grid UI or status endpoint for a quick check
Open the Grid UI in a browser, or request the server’s /status endpoint. These are useful first checks when the immediate question is whether the Grid is up and reporting status. The exact host and port depend on how you launched or deployed the server.
curl -sS http://localhost:4444/status
For a distributed Grid, query the Router address exposed to clients. A successful HTTP response alone does not establish that the Grid has free capacity; inspect the returned status and use GraphQL for structured capacity and session data.
Query capacity and sessions with GraphQL
GraphQL is useful when you need machine-readable details, such as the maximum session capacity, current session count, node status, slots, or session metadata. The following examples query commonly documented fields; verify the schema against the Selenium version in use.
curl -sS http://localhost:4444/graphql \
-H 'Content-Type: application/json' \
--data-raw '{"query":"{ grid { maxSession sessionCount } }"}'
To inspect nodes and slots:
curl -sS http://localhost:4444/graphql \
-H 'Content-Type: application/json' \
--data-raw '{"query":"{ nodes { id uri availability maxSession sessionCount slots { id stereotype session { sessionId, capabilities, startTime, uri } } } }"}'
GraphQL fields and nesting can differ across server versions. If a query is rejected, use the official GraphQL query support examples for the version you run. The documented session information can include capabilities, start time, node information, and duration.
Give sessions useful names
Add the se:name capability when creating a session so that it can be recognized in the Grid UI and queried through Grid APIs. Add other appropriate se: metadata where useful for your setup. Avoid putting secrets or personal data in capability values because session metadata can be visible to Grid operators and monitoring tools.
// Java example: label a remote session with se:name
ChromeOptions options = new ChromeOptions();
options.setCapability("se:name", "checkout-chrome-linux");
WebDriver driver = new RemoteWebDriver(gridUrl, options);
Use a label that makes the run easy to identify, such as a suite, browser, or job identifier. Selenium documents Grid UI, /status, GraphQL, and se: metadata in its Grid getting started guide.
3. Trace a request across Grid components
Grid may include a Router, New Session Queue, Distributor, Node, Session Map, and Event Bus. In a standalone deployment, fewer boundaries are involved; in a fully distributed deployment, a request can cross several separately running roles. Read spans in the context of the roles actually deployed.
- Identify the failed or unusually slow operation and its approximate time. Preserve the client-side error and session identifier if available.
- Find the corresponding trace in your configured output or trace backend. Match using available trace context and timing; do not assume every client error automatically contains a trace ID.
- Follow spans in order to see which Grid component handled each part of the request.
- Inspect the span timing and events around the delay or error. Use the component boundary to narrow whether the issue occurred while routing, queuing, distributing, mapping a session, or executing on a node.
- Correlate with GraphQL or UI state at the same time. For example, a saturated session count is useful live context, while a trace explains what happened to an individual request.
- Check the relevant node and session details, then compare with server logs from the implicated component.
Traces show request history and timing; GraphQL, UI, and status endpoints show current state. Neither replaces the other when diagnosing intermittent failures.
4. Enable useful trace output or a trace backend
Inspect the server’s actual configuration options
Selenium’s configuration help is generated from the running implementation. Check it before copying flags from a different release:
java -jar selenium-server-4.x.jar info config
java -jar selenium-server-4.x.jar info tracing
Replace 4.x with the exact server JAR filename. The commands provide configuration and tracing details for that version, which is especially useful when server flags or defaults have changed. See the official configuration help.
Use FINE logging for direct inspection
When you want to inspect traces and events in a local or diagnostic environment, run the server at FINE log level using the option supported by the deployed release. For releases that accept the following command-line form:
java -jar selenium-server-4.x.jar standalone --log-level FINE
For Hub/Node or fully distributed deployments, apply the logging configuration to the relevant process or processes according to that release’s configuration help. FINE output can be verbose; use it deliberately and ensure logs are collected and retained appropriately.
Use Jaeger when you need to query traces
Selenium’s observability guide presents Jaeger as a backend for collecting, querying, filtering, and visualizing traces. Follow info tracing for the deployed Selenium version to configure the exporter and endpoint. The precise properties and startup flags are version-sensitive, so don’t assume a Jaeger configuration copied from another release will work unchanged.
5. Monitor Selenium Grid with Prometheus: check the version first
The Selenium Grid 4.41.0 release article describes a Session Event API and a native Prometheus metrics endpoint. Treat these as version-bound features: verify that the deployed server includes them before designing dashboards, alerts, or integrations around them. This research does not establish a full compatibility matrix across releases.
For a deployment that does not have the documented endpoint, use the supported status, GraphQL, logs, and tracing mechanisms available in that exact version. Do not assume that a Prometheus endpoint exists merely because the deployment uses Selenium 4.
6. A practical monitoring routine
- Check reachability: inspect the Grid UI or request
/status. - Check capacity: query GraphQL for session counts, nodes, and slots; compare current use with the configured maximum.
- Label sessions: set
se:nameand useful metadata so active work is identifiable. - Investigate a failure: locate the request trace, follow spans through the deployed topology, and correlate with node/session state and component logs.
- Centralize when needed: configure a trace backend such as Jaeger using the deployed server’s
info tracinghelp. - Validate release-specific telemetry: confirm the Selenium version before depending on newer event or metrics features.
7. Performance, reliability, and cost considerations
- Logging volume: FINE-level diagnostic output can be verbose. Collect only what is useful, and account for log storage, retention, and access controls.
- Trace backend availability: a backend adds an operational dependency for centralized querying. Keep the Grid’s own health checks and logs available so a trace backend outage does not remove all diagnostic options.
- Sampling and completeness: if you configure sampling or exporter behavior, understand the deployed configuration before relying on traces to contain every request. Check the version’s tracing help.
- Topology matters: a trace is more valuable when you know which roles run where. Record deployment role boundaries and correlate each span with the right process.
- Capacity is not a performance benchmark:
maxSessionandsessionCountdescribe configured and current capacity; they do not alone establish browser startup speed or test throughput. - Cost: Selenium’s documented observability mechanisms do not imply a specific hosting, storage, or licensing cost. Actual infrastructure cost depends on your Grid, log retention, trace backend, and deployment choices.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
/status cannot connect |
Wrong host/port, Grid process stopped, or endpoint not reachable from the caller. | Check the Router or standalone address, process state, network path, and deployment port mapping. |
| GraphQL returns an error or unknown field | Query schema differs from the deployed release, or the query shape is invalid. | Use the GraphQL examples for that release and inspect the current schema/field names. |
| No useful session names appear | Session creation omitted se:name, or metadata is not being inspected in the right UI/API view. |
Set the capability when creating new sessions and query the session fields supported by the server version. |
| Trace output is missing from console | Log level is below FINE, logging is configured on a different process, or version-specific tracing behavior differs. | Check info tracing and the logging configuration for each relevant Grid role. |
| Spans appear incomplete | Only some components are configured to export, the backend cannot receive data, or sampling/export settings omit spans. | Verify tracing configuration and connectivity for the relevant processes, then reproduce a request and inspect timestamps and available events. |
| Cannot find a trace by session ID | The trace backend may not index that capability or identifier, or the request happened outside the selected time window. | Search by available trace context and time first; correlate with server logs and GraphQL session details. |
| Prometheus endpoint is absent | The deployed release may not include the version-bound feature. | Confirm the exact server version and release documentation before configuring a scrape target. |
| FINE logs overwhelm storage | Verbose diagnostic output is enabled for too many processes or retained too long. | Limit diagnostic logging to the needed components and period; configure suitable collection and retention. |
9. Or skip the browser setup
For the separate task of capturing a website as an image or PDF, ScreenshotNeo provides a screenshot API and MCP server. It does not monitor Selenium Grid sessions or traces. One GET request can return a screenshot; see the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners are accepted and removed before capture; known newsletter popups and chat widgets are removed too, and each step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card.
10. FAQ
Does Selenium Grid observability require a third-party tracing product?
No. Selenium documents built-in OpenTelemetry instrumentation and console trace output. A backend such as Jaeger is useful when you need centralized collection and querying.
Can GraphQL show historical request timing?
GraphQL is primarily the structured view of Grid state and sessions. Use traces to investigate the path and timing of a request.
Should I alert on session count alone?
Session count is a useful capacity signal, but it does not identify the cause of slow or failed requests. Pair capacity checks with health status and trace-based investigation.
Can I use these instructions unchanged on every Selenium 4 release?
No. Check the deployed server’s info config and info tracing output, especially for flags and newer telemetry features.


