ScreenshotNeo

BlogHow-to

How to Scale Selenium Grid with KEDA

Use KEDA’s Selenium Grid scaler to match browser-node capacity to queued sessions, with configuration examples, scale-down guidance, and troubleshooting.

By the ScreenshotNeo team4 October 20269 min read

Direct answer: Install KEDA, run Selenium Grid with browser-node workloads, and attach a KEDA ScaledObject to each browser pool. The built-in Selenium Grid scaler reads pending session demand from Grid’s GraphQL endpoint and scales the matching node workload. Match the scaler’s nodeMaxSessions to the node’s actual session capacity, and bound replicas to what your cluster can schedule.

KEDA’s Selenium Grid scaler is available in KEDA v2.4 and later. The examples below use a persistent Chrome node Deployment. They are configuration templates: adapt names, versions, resource settings, namespaces, authentication, and capabilities to your Grid and cluster. Check the documentation for the exact KEDA release you deploy because scaler metadata and behavior can evolve. KEDA Selenium Grid scaler documentation

How queue-aware scaling works

Selenium Grid routes WebDriver scripts to remote browser instances for parallel, cross-browser, or cross-platform testing. When no compatible browser slot is free, a new session request waits in Grid’s session queue. KEDA polls the Grid GraphQL endpoint, matches queued demand to a browser pool’s capability metadata, and supplies a metric that drives Kubernetes scaling.

For a persistent node pool, KEDA scales a Deployment. The scaler’s capacity model must agree with the number of simultaneous sessions each node can really serve. For example, if each Chrome node supports two sessions, set nodeMaxSessions to 2 and configure the node for two sessions. A mismatch can lead to over-provisioning or a queue that remains despite apparently available capacity.

Queue-aware scaling is a closer signal of unmet test demand than CPU utilization alone. Browser nodes can be busy without crossing a CPU threshold, and scaling down a node with an active session can interrupt a test. Queue scaling still requires graceful termination and enough Kubernetes capacity.

Deploy a persistent browser pool

This example assumes that a Grid Hub is reachable inside the cluster as selenium-hub, with its GraphQL endpoint at http://selenium-hub:4444/graphql. The Chrome node Deployment and ScaledObject should normally be in the same namespace. Install KEDA according to its installation guide before applying the resource below.

  1. Deploy and verify the Grid Hub and a Chrome node workload.
  2. Expose the Hub’s GraphQL endpoint to KEDA using an in-cluster service URL. Confirm network policies allow KEDA to reach it.
  3. Apply a ScaledObject for each capability pool. Start with a conservative maximum replica count based on cluster capacity.
  4. Submit concurrent test sessions and observe the queue, KEDA metrics, Deployment replicas, pod scheduling, and session completion.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: selenium-chrome-node
spec:
  replicas: 0
  selector:
    matchLabels:
      app: selenium-chrome-node
  template:
    metadata:
      labels:
        app: selenium-chrome-node
    spec:
      containers:
        - name: chrome
          image: selenium/node-chrome:4.41.0
          ports:
            - containerPort: 5555
          env:
            - name: SE_EVENT_BUS_HOST
              value: selenium-hub
            - name: SE_NODE_PLATFORM_NAME
              value: Linux
            - name: SE_NODE_MAX_SESSIONS
              value: "1"
          resources:
            requests:
              cpu: "1"
              memory: 2Gi
            limits:
              memory: 3Gi
---
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: selenium-chrome
spec:
  scaleTargetRef:
    name: selenium-chrome-node
  minReplicaCount: 0
  maxReplicaCount: 8
  triggers:
    - type: selenium-grid
      metadata:
        url: http://selenium-hub:4444/graphql
        browserName: chrome
        platformName: Linux
        nodeMaxSessions: "1"
        # Optional scaler controls are version-specific; consult your KEDA docs.

The image tag, resources, event-bus wiring, and session setting are examples, not universal values. Use compatible, pinned Grid images for your deployment. Ensure the node’s event-bus settings, registration, ports, and Hub service agree with your Grid topology. The sample sets one session per node, so nodeMaxSessions is also one. If you intentionally enable multiple sessions, configure SE_NODE_MAX_SESSIONS (and any required override setting for your Selenium image/version) and the KEDA metadata to the same value.

The trigger’s matching fields can include browserName, browserVersion, and platformName. Use the same capability values that your node stereotypes advertise. Set one trigger per browser pool you intend to scale. If Grid authentication is enabled, keep credentials in a Kubernetes Secret and configure KEDA TriggerAuthentication as documented; do not commit credentials in a public manifest.

Configure capacity and scale behavior

Replica bounds and node slots

minReplicaCount controls the warm floor; setting it to zero permits scale-to-zero. A nonzero minimum can reduce cold-start waits at the cost of keeping idle browser pods available. maxReplicaCount caps the workload. Choose it from available node-pool capacity, per-pod resource requests, and competing workloads. Kubernetes may leave replicas pending if the cluster cannot schedule them; KEDA cannot create underlying compute capacity unless your cluster autoscaler is configured to do so.

nodeMaxSessions represents slots per browser node for the scaler’s capacity calculations. Keep it synchronized with --max-sessions or SE_NODE_MAX_SESSIONS. Revisit both settings after changing node images, environment variables, stereotypes, or browser concurrency. Higher sessions per pod may reduce pod count, but can increase contention for CPU and memory; size and validate using your own tests.

Capability pools

Separate Chrome, Firefox, Edge, browser versions, or platform stereotypes into workloads when they need independent scaling. Each trigger must describe the capability pool served by its target. If capability values do not match, queued demand may not activate that pool. Avoid overlapping or ambiguous stereotypes unless you have deliberately designed how Grid routes those requests.

Scale to zero and cold starts

With minReplicaCount: 0, KEDA can activate a pool from zero when matching requests wait. A test still incurs the time to schedule a pod, pull its image if uncached, start the browser, register the node, and create the session. Keep a warm minimum if that delay is unacceptable. Scale-to-zero also depends on KEDA reaching the Grid endpoint and on the configured trigger matching the capabilities submitted by clients.

Scale-down and safe termination

Do not treat an idle CPU reading as proof that a browser node is safe to remove. A node can be serving a long-running session. Configure and validate the Grid and Kubernetes termination path so active tests can finish or be drained before pod termination. Kubernetes termination grace periods, Grid session timeouts, and test duration need to be considered together. The sources do not prescribe one universal drain configuration or cooldown value.

ScaledJob alternative

A ScaledJob can create browser-node Jobs instead of maintaining a persistent Deployment. It suits a lifecycle where a browser node handles work and then terminates. The scaler’s includeOngoingSessions setting must fit the selected ScaledJob scaling strategy:

Target and strategy Recommended ongoing-session setting Reason
ScaledObject Default true Running replicas are accounted for by the HPA calculation.
ScaledJob: default or common custom Default true The strategy accounts for running Jobs when calculating additional demand.
ScaledJob: accurate or eager false These strategies do not subtract ongoing sessions in a way that permits counting them again; reporting them can repeatedly create unnecessary Jobs.
apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
  name: selenium-chrome-job
spec:
  maxReplicaCount: 8
  scalingStrategy:
    strategy: accurate
  jobTargetRef:
    template:
      spec:
        containers:
          - name: chrome
            image: selenium/node-chrome:4.41.0
            env:
              - name: SE_EVENT_BUS_HOST
                value: selenium-hub
              - name: SE_NODE_PLATFORM_NAME
                value: Linux
            resources:
              requests:
                cpu: "1"
                memory: 2Gi
        restartPolicy: Never
  triggers:
    - type: selenium-grid
      metadata:
        url: http://selenium-hub:4444/graphql
        browserName: chrome
        platformName: Linux
        includeOngoingSessions: "false"

For a Job-per-node pattern, ensure the node’s own session capacity and Job lifecycle match the one-session-per-Job assumption. Validate the precise strategy behavior against your deployed KEDA version. KEDA documents the strategy and ongoing-session interaction.

Security and Kubernetes configuration

Keep the GraphQL endpoint on a network path accessible to KEDA, but do not expose an unauthenticated management endpoint publicly. When authentication is enabled, store secrets in Kubernetes and reference them through KEDA authentication resources. Restrict who can read those secrets and edit ScaledObjects.

Browser pods need sufficient resources and access to the Grid event bus and Hub. Set resource requests so the scheduler can place pods predictably, and limits appropriate to your workload. Review image pull permissions, network policies, service accounts, and pod security requirements. Selenium Grid’s Kubernetes session factory has additional settings for API endpoint, namespace, service account, image-to-capability mappings or Job templates, and image pull policy; consult its CLI reference when using that design.

When to use Grid’s Kubernetes session factory

Selenium Grid 4.41.0 introduced a Kubernetes session factory that provisions an ephemeral browser Pod per session and removes it when the session closes. This is a different operating model from KEDA scaling a persistent browser-node workload. It can reduce separate scaler configuration if Grid-managed per-session pods fit your lifecycle, but requires Kubernetes configuration and permissions for the Grid node. Check the feature’s availability and exact setup for your Selenium release before adopting it. The cited sources do not establish that this approach is faster or cheaper for every workload. Selenium Grid 4.41.0 release discussion

Consideration KEDA with persistent node pool Grid Kubernetes session factory
Provisioning unit Browser-node replicas driven by queued demand One ephemeral browser Pod per session
Configuration KEDA trigger plus Grid node and session settings Grid Kubernetes session configuration and permissions
Useful fit Existing node-pool operations and shared slots Per-session pod lifecycle managed through Grid
Evidence limit No universal latency, throughput, or cost winner is established Measure with representative images and concurrency

Monitor and validate the scaling loop

  • Watch the Grid session queue and confirm requests advertise the expected browser capabilities.
  • Inspect KEDA operator and metrics-adapter logs for scaler polling, GraphQL, and authentication errors.
  • Track ScaledObject or ScaledJob status, desired replicas/jobs, actual pods, and Kubernetes pending reasons.
  • Track node registration, active sessions, session creation failures, and test completion times.
  • Exercise a burst, a sustained queue, and an idle period. Confirm scale-up occurs, the queue drains, and scale-down does not interrupt active tests.

There is no universal replica count, concurrency setting, or cost estimate in the cited documentation. Measure startup delay, queue wait, resource use, and cluster spend with your browser images and workload before changing capacity limits.

Troubleshooting

Symptom Likely cause What to check or change
Queued sessions do not scale the pool Bad GraphQL URL, network block, scaler auth failure, or mismatched capability filters Check KEDA logs, endpoint reachability, credentials, and exact browser/platform/version values.
Pool scales too far nodeMaxSessions differs from actual node capacity, or ongoing work is counted incorrectly for a ScaledJob strategy Align scaler and node session settings. For accurate/eager ScaledJobs, set includeOngoingSessions: "false".
Queue remains while replicas exist Pods are Pending, nodes have not registered, wrong stereotypes, or each node has fewer usable slots than assumed Inspect pod events/resources, Grid node status, capability stereotypes, and session capacity.
Pods are Pending Insufficient schedulable CPU/memory, taints, quota, pull errors, or unavailable image Read pod events; adjust bounded capacity, node selectors/tolerations, quota, or image pull setup as appropriate.
Scale-to-zero requests wait too long Cold pod scheduling, image pull, browser startup, or KEDA polling delay Use a warm minimum if latency matters, pre-pull images where suitable, and inspect the full startup path.
Active tests fail during scale-down Termination interrupts a live session or drain grace is too short Review Grid draining behavior, Kubernetes termination grace period, and session/test duration.
Repeated Jobs are created unnecessarily Ongoing sessions are included under a ScaledJob strategy that does not subtract them Set includeOngoingSessions: "false" with accurate or eager; confirm version-specific behavior.
Manifest rejected or trigger not recognized KEDA CRDs/operator version mismatch or invalid metadata type Check installed CRDs and KEDA version; use the corresponding scaler documentation and string-valued examples where required.

Or skip the browser setup

If your task is to capture website screenshots rather than execute interactive browser tests, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. The API accepts parameters used by other screenshot APIs, with options including full-page capture, CSS selectors, device presets, custom CSS and JavaScript, and caching. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted as a visitor and removed before capture, along with supported consent banners, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server lets AI agents use screenshot, page-info, and PDF tools. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots.

Start with 1,000 free screenshots a month; no card required.

FAQ

Does KEDA replace Selenium Grid?

No. Grid routes WebDriver sessions and manages browser nodes; KEDA adjusts Kubernetes workload capacity from queue demand.

Can this scale more than one browser?

Yes. Configure a pool and matching trigger for each browser capability set you need to serve.

Does queue-aware scaling guarantee a test starts immediately?

No. Scheduling, image pulls, browser startup, node registration, cluster capacity, and KEDA polling still affect start time.

Should I choose KEDA or the Grid-native Kubernetes session factory?

Choose based on whether persistent shared node pools or ephemeral per-session Pods fit your operations. Validate compatibility and measure your own workload; published source material does not show a universal winner.

Sources