How to Scale Selenium Grid with KEDA
Use KEDA’s Selenium Grid scaler to match browser-node capacity to queued sessions, with configuration examples, scale-down guidance, and troubleshooting.
Direct answer: Install KEDA, run Selenium Grid with browser-node workloads, and attach a KEDA ScaledObject to each browser pool. The built-in Selenium Grid scaler reads pending session demand from Grid’s GraphQL endpoint and scales the matching node workload. Match the scaler’s nodeMaxSessions to the node’s actual session capacity, and bound replicas to what your cluster can schedule.
KEDA’s Selenium Grid scaler is available in KEDA v2.4 and later. The examples below use a persistent Chrome node Deployment. They are configuration templates: adapt names, versions, resource settings, namespaces, authentication, and capabilities to your Grid and cluster. Check the documentation for the exact KEDA release you deploy because scaler metadata and behavior can evolve. KEDA Selenium Grid scaler documentation
How queue-aware scaling works
Selenium Grid routes WebDriver scripts to remote browser instances for parallel, cross-browser, or cross-platform testing. When no compatible browser slot is free, a new session request waits in Grid’s session queue. KEDA polls the Grid GraphQL endpoint, matches queued demand to a browser pool’s capability metadata, and supplies a metric that drives Kubernetes scaling.
For a persistent node pool, KEDA scales a Deployment. The scaler’s capacity model must agree with the number of simultaneous sessions each node can really serve. For example, if each Chrome node supports two sessions, set nodeMaxSessions to 2 and configure the node for two sessions. A mismatch can lead to over-provisioning or a queue that remains despite apparently available capacity.
Queue-aware scaling is a closer signal of unmet test demand than CPU utilization alone. Browser nodes can be busy without crossing a CPU threshold, and scaling down a node with an active session can interrupt a test. Queue scaling still requires graceful termination and enough Kubernetes capacity.
Deploy a persistent browser pool
This example assumes that a Grid Hub is reachable inside the cluster as selenium-hub, with its GraphQL endpoint at http://selenium-hub:4444/graphql. The Chrome node Deployment and ScaledObject should normally be in the same namespace. Install KEDA according to its installation guide before applying the resource below.
- Deploy and verify the Grid Hub and a Chrome node workload.
- Expose the Hub’s GraphQL endpoint to KEDA using an in-cluster service URL. Confirm network policies allow KEDA to reach it.
- Apply a
ScaledObjectfor each capability pool. Start with a conservative maximum replica count based on cluster capacity. - Submit concurrent test sessions and observe the queue, KEDA metrics, Deployment replicas, pod scheduling, and session completion.
apiVersion: apps/v1
kind: Deployment
metadata:
name: selenium-chrome-node
spec:
replicas: 0
selector:
matchLabels:
app: selenium-chrome-node
template:
metadata:
labels:
app: selenium-chrome-node
spec:
containers:
- name: chrome
image: selenium/node-chrome:4.41.0
ports:
- containerPort: 5555
env:
- name: SE_EVENT_BUS_HOST
value: selenium-hub
- name: SE_NODE_PLATFORM_NAME
value: Linux
- name: SE_NODE_MAX_SESSIONS
value: "1"
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
memory: 3Gi
---
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: selenium-chrome
spec:
scaleTargetRef:
name: selenium-chrome-node
minReplicaCount: 0
maxReplicaCount: 8
triggers:
- type: selenium-grid
metadata:
url: http://selenium-hub:4444/graphql
browserName: chrome
platformName: Linux
nodeMaxSessions: "1"
# Optional scaler controls are version-specific; consult your KEDA docs.
The image tag, resources, event-bus wiring, and session setting are examples, not universal values. Use compatible, pinned Grid images for your deployment. Ensure the node’s event-bus settings, registration, ports, and Hub service agree with your Grid topology. The sample sets one session per node, so nodeMaxSessions is also one. If you intentionally enable multiple sessions, configure SE_NODE_MAX_SESSIONS (and any required override setting for your Selenium image/version) and the KEDA metadata to the same value.
The trigger’s matching fields can include browserName, browserVersion, and platformName. Use the same capability values that your node stereotypes advertise. Set one trigger per browser pool you intend to scale. If Grid authentication is enabled, keep credentials in a Kubernetes Secret and configure KEDA TriggerAuthentication as documented; do not commit credentials in a public manifest.
Configure capacity and scale behavior
Replica bounds and node slots
minReplicaCount controls the warm floor; setting it to zero permits scale-to-zero. A nonzero minimum can reduce cold-start waits at the cost of keeping idle browser pods available. maxReplicaCount caps the workload. Choose it from available node-pool capacity, per-pod resource requests, and competing workloads. Kubernetes may leave replicas pending if the cluster cannot schedule them; KEDA cannot create underlying compute capacity unless your cluster autoscaler is configured to do so.
nodeMaxSessions represents slots per browser node for the scaler’s capacity calculations. Keep it synchronized with --max-sessions or SE_NODE_MAX_SESSIONS. Revisit both settings after changing node images, environment variables, stereotypes, or browser concurrency. Higher sessions per pod may reduce pod count, but can increase contention for CPU and memory; size and validate using your own tests.
Capability pools
Separate Chrome, Firefox, Edge, browser versions, or platform stereotypes into workloads when they need independent scaling. Each trigger must describe the capability pool served by its target. If capability values do not match, queued demand may not activate that pool. Avoid overlapping or ambiguous stereotypes unless you have deliberately designed how Grid routes those requests.
Scale to zero and cold starts
With minReplicaCount: 0, KEDA can activate a pool from zero when matching requests wait. A test still incurs the time to schedule a pod, pull its image if uncached, start the browser, register the node, and create the session. Keep a warm minimum if that delay is unacceptable. Scale-to-zero also depends on KEDA reaching the Grid endpoint and on the configured trigger matching the capabilities submitted by clients.
Scale-down and safe termination
Do not treat an idle CPU reading as proof that a browser node is safe to remove. A node can be serving a long-running session. Configure and validate the Grid and Kubernetes termination path so active tests can finish or be drained before pod termination. Kubernetes termination grace periods, Grid session timeouts, and test duration need to be considered together. The sources do not prescribe one universal drain configuration or cooldown value.
ScaledJob alternative
A ScaledJob can create browser-node Jobs instead of maintaining a persistent Deployment. It suits a lifecycle where a browser node handles work and then terminates. The scaler’s includeOngoingSessions setting must fit the selected ScaledJob scaling strategy:
| Target and strategy | Recommended ongoing-session setting | Reason |
|---|---|---|
| ScaledObject | Default true |
Running replicas are accounted for by the HPA calculation. |
ScaledJob: default or common custom |
Default true |
The strategy accounts for running Jobs when calculating additional demand. |
ScaledJob: accurate or eager |
false |
These strategies do not subtract ongoing sessions in a way that permits counting them again; reporting them can repeatedly create unnecessary Jobs. |
apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
name: selenium-chrome-job
spec:
maxReplicaCount: 8
scalingStrategy:
strategy: accurate
jobTargetRef:
template:
spec:
containers:
- name: chrome
image: selenium/node-chrome:4.41.0
env:
- name: SE_EVENT_BUS_HOST
value: selenium-hub
- name: SE_NODE_PLATFORM_NAME
value: Linux
resources:
requests:
cpu: "1"
memory: 2Gi
restartPolicy: Never
triggers:
- type: selenium-grid
metadata:
url: http://selenium-hub:4444/graphql
browserName: chrome
platformName: Linux
includeOngoingSessions: "false"
For a Job-per-node pattern, ensure the node’s own session capacity and Job lifecycle match the one-session-per-Job assumption. Validate the precise strategy behavior against your deployed KEDA version. KEDA documents the strategy and ongoing-session interaction.
Security and Kubernetes configuration
Keep the GraphQL endpoint on a network path accessible to KEDA, but do not expose an unauthenticated management endpoint publicly. When authentication is enabled, store secrets in Kubernetes and reference them through KEDA authentication resources. Restrict who can read those secrets and edit ScaledObjects.
Browser pods need sufficient resources and access to the Grid event bus and Hub. Set resource requests so the scheduler can place pods predictably, and limits appropriate to your workload. Review image pull permissions, network policies, service accounts, and pod security requirements. Selenium Grid’s Kubernetes session factory has additional settings for API endpoint, namespace, service account, image-to-capability mappings or Job templates, and image pull policy; consult its CLI reference when using that design.
When to use Grid’s Kubernetes session factory
Selenium Grid 4.41.0 introduced a Kubernetes session factory that provisions an ephemeral browser Pod per session and removes it when the session closes. This is a different operating model from KEDA scaling a persistent browser-node workload. It can reduce separate scaler configuration if Grid-managed per-session pods fit your lifecycle, but requires Kubernetes configuration and permissions for the Grid node. Check the feature’s availability and exact setup for your Selenium release before adopting it. The cited sources do not establish that this approach is faster or cheaper for every workload. Selenium Grid 4.41.0 release discussion
| Consideration | KEDA with persistent node pool | Grid Kubernetes session factory |
|---|---|---|
| Provisioning unit | Browser-node replicas driven by queued demand | One ephemeral browser Pod per session |
| Configuration | KEDA trigger plus Grid node and session settings | Grid Kubernetes session configuration and permissions |
| Useful fit | Existing node-pool operations and shared slots | Per-session pod lifecycle managed through Grid |
| Evidence limit | No universal latency, throughput, or cost winner is established | Measure with representative images and concurrency |
Monitor and validate the scaling loop
- Watch the Grid session queue and confirm requests advertise the expected browser capabilities.
- Inspect KEDA operator and metrics-adapter logs for scaler polling, GraphQL, and authentication errors.
- Track ScaledObject or ScaledJob status, desired replicas/jobs, actual pods, and Kubernetes pending reasons.
- Track node registration, active sessions, session creation failures, and test completion times.
- Exercise a burst, a sustained queue, and an idle period. Confirm scale-up occurs, the queue drains, and scale-down does not interrupt active tests.
There is no universal replica count, concurrency setting, or cost estimate in the cited documentation. Measure startup delay, queue wait, resource use, and cluster spend with your browser images and workload before changing capacity limits.
Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Queued sessions do not scale the pool | Bad GraphQL URL, network block, scaler auth failure, or mismatched capability filters | Check KEDA logs, endpoint reachability, credentials, and exact browser/platform/version values. |
| Pool scales too far | nodeMaxSessions differs from actual node capacity, or ongoing work is counted incorrectly for a ScaledJob strategy |
Align scaler and node session settings. For accurate/eager ScaledJobs, set includeOngoingSessions: "false". |
| Queue remains while replicas exist | Pods are Pending, nodes have not registered, wrong stereotypes, or each node has fewer usable slots than assumed | Inspect pod events/resources, Grid node status, capability stereotypes, and session capacity. |
| Pods are Pending | Insufficient schedulable CPU/memory, taints, quota, pull errors, or unavailable image | Read pod events; adjust bounded capacity, node selectors/tolerations, quota, or image pull setup as appropriate. |
| Scale-to-zero requests wait too long | Cold pod scheduling, image pull, browser startup, or KEDA polling delay | Use a warm minimum if latency matters, pre-pull images where suitable, and inspect the full startup path. |
| Active tests fail during scale-down | Termination interrupts a live session or drain grace is too short | Review Grid draining behavior, Kubernetes termination grace period, and session/test duration. |
| Repeated Jobs are created unnecessarily | Ongoing sessions are included under a ScaledJob strategy that does not subtract them | Set includeOngoingSessions: "false" with accurate or eager; confirm version-specific behavior. |
| Manifest rejected or trigger not recognized | KEDA CRDs/operator version mismatch or invalid metadata type | Check installed CRDs and KEDA version; use the corresponding scaler documentation and string-valued examples where required. |
Or skip the browser setup
If your task is to capture website screenshots rather than execute interactive browser tests, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. The API accepts parameters used by other screenshot APIs, with options including full-page capture, CSS selectors, device presets, custom CSS and JavaScript, and caching. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners are accepted as a visitor and removed before capture, along with supported consent banners, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server lets AI agents use screenshot, page-info, and PDF tools. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots.
Start with 1,000 free screenshots a month; no card required.
FAQ
Does KEDA replace Selenium Grid?
No. Grid routes WebDriver sessions and manages browser nodes; KEDA adjusts Kubernetes workload capacity from queue demand.
Can this scale more than one browser?
Yes. Configure a pool and matching trigger for each browser capability set you need to serve.
Does queue-aware scaling guarantee a test starts immediately?
No. Scheduling, image pulls, browser startup, node registration, cluster capacity, and KEDA polling still affect start time.
Should I choose KEDA or the Grid-native Kubernetes session factory?
Choose based on whether persistent shared node pools or ephemeral per-session Pods fit your operations. Validate compatibility and measure your own workload; published source material does not show a universal winner.


