How to Load Balance Headless Browser Sessions
Control headless browser concurrency with bounded workers, reliable cleanup, and clear capacity signals across managed services and self-hosted fleets.
Load balance headless browser sessions by placing jobs in a queue, allowing only a bounded number of browser sessions to run at once, and releasing each slot as soon as its session ends—even when the job fails. A provider’s queue can absorb bursts, but an application-side limit still helps control load on target sites and makes your own queue delay visible.
A session is an active browser connection or browser context doing work. Concurrency is the number of sessions active simultaneously. Browserless defines it as “the maximum number of browser sessions that can run simultaneously on a Browserless instance.” Check the current limits of your chosen service or deployment before setting a cap. Browserless terminology
1. Set a capacity limit and queue jobs
Choose an application concurrency cap based on the capacity you intend to use. For a managed provider, that cap should fit within your current plan allowance and leave room for other consumers. For self-hosted browsers, start with a conservative cap and tune it using representative load tests; there is no portable sessions-per-CPU or sessions-per-GB rule that fits every page and browser workload.
A bounded worker pool or semaphore is a practical way to enforce the limit. Each worker acquires a slot before connecting to or launching a browser, runs one job, closes its session, and releases the slot in unconditional cleanup. A queue absorbs short bursts while the active-session count stays bounded.
- Enqueue incoming browser jobs rather than opening a session for every request immediately.
- Start work only when a worker slot is available.
- Connect to the browser endpoint or launch a local browser.
- Run the automation with a job-level timeout and any target-site pacing rules you need.
- Close the page, context, and browser connection in cleanup, including on exceptions.
- Record job outcome and queue delay, then let the worker take the next job.
Browserless documents automatic queuing and also recommends client-side concurrency limits to avoid overwhelming a target site. Treat provider-side queuing as burst handling, not a substitute for capacity planning: queued requests still affect latency and may interact with request timeouts or throughput limits. Confirm those details for your provider and plan. Concurrent sessions · Terminology and capacity
2. Implement a bounded Playwright worker pool
This Node.js example connects to a remote Playwright browser over a WebSocket endpoint and runs no more than MAX_CONCURRENT_SESSIONS jobs at once. Replace the endpoint and token with the values provided by your browser service. The queue is in memory for clarity; use a durable job queue if jobs must survive process restarts.
import { chromium } from 'playwright';
const endpoint = process.env.BROWSER_WS_ENDPOINT;
const maxConcurrent = Number(process.env.MAX_CONCURRENT_SESSIONS ?? 4);
const jobs = [
{ id: 'one', url: 'https://example.com' },
{ id: 'two', url: 'https://example.org' },
];
if (!endpoint) throw new Error('Set BROWSER_WS_ENDPOINT');
if (!Number.isInteger(maxConcurrent) || maxConcurrent < 1) {
throw new Error('MAX_CONCURRENT_SESSIONS must be a positive integer');
}
async function capture(job) {
let browser;
try {
browser = await chromium.connectOverCDP(endpoint);
// For CDP endpoints, use the default context when endpoint-level
// proxy or profile settings must carry through.
const context = browser.contexts()[0];
if (!context) throw new Error('Remote browser has no default context');
const page = await context.newPage();
await page.goto(job.url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
const title = await page.title();
console.log({ job: job.id, title });
await page.close();
} finally {
// Always disconnect this client connection, even if navigation fails.
if (browser) await browser.close();
}
}
let next = 0;
async function worker() {
while (true) {
const index = next++;
if (index >= jobs.length) return;
const job = jobs[index];
try {
await capture(job);
} catch (error) {
console.error(`Job ${job.id} failed`, error);
// Persist a retry or terminal failure here according to your policy.
}
}
}
await Promise.all(
Array.from({ length: Math.min(maxConcurrent, jobs.length) }, () => worker()),
);
Install the Playwright package and configure the remote endpoint using the browser provider’s current connection instructions. Connection APIs, endpoint URLs, browser versions, and session semantics vary by provider; verify them against the deployed versions. Browserless’s Playwright CDP examples advise using the default context when launch-level proxy or profile settings need to carry through, because a newly created context may not inherit them. Validate this integration detail for your endpoint.
Important cleanup detail
In this example, the CDP connection is closed in finally, even when page creation or navigation throws. For a direct Playwright browser connection, close the context and browser in a finally block as well. Follow the provider’s documented distinction between closing a session and merely disconnecting a client: make sure the operation actually releases the remote capacity. Browserless warns that failing to close remote sessions can exhaust concurrency. Browserless best practices
3. Choose where the browser fleet runs
| Decision | Managed browser service | Self-hosted fleet |
|---|---|---|
| Operations | The provider manages browser pool and runtime operations. | Your team operates deployment, capacity, health, and browser updates. |
| Control | Use provider endpoints and supported controls. | You control more of the deployment and configuration. |
| Capacity | Check current plan limits, queueing behavior, timeouts, and session semantics. | Set and operate concurrency in your deployment; measure before scaling. |
| Geography | Choose among supported provider regions and endpoint hosts. | Select infrastructure regions under your control. |
| Validation | Confirm quota, endpoint, timeout, and cleanup behavior. | Load-test worker size, scaling, health checks, updates, and cleanup. |
A managed service can reduce browser operations work; self-hosting gives your team deployment ownership. The documentation does not establish a general cost or performance break-even point, so compare using your workload, staffing, and operating requirements. Browserless documents both managed browser usage and deployment scaling options. Browsers as a Service
4. Pick a region and verify connection behavior
When latency matters, place the browser endpoint near the workload or the users who depend on its results. Check the provider’s current regional endpoint map and make sure the WebSocket URL in your environment targets the intended region. Browserless recommends nearby regions to reduce latency; available regions and hostnames can change. Connection URLs and endpoints
Measure the full job path, not just browser connection time: queue wait, connection setup, navigation, page readiness, extraction or screenshot work, and cleanup. A nearby region cannot fix slow target pages, overloaded workers, or an overly strict readiness condition.
5. Observe queue pressure and tune safely
Track the signals that explain whether your cap is too high, too low, or simply facing slow pages. These are operational recommendations; exact thresholds depend on your service and workload.
- Active sessions: compare current sessions with your application cap and provider capacity signals where available.
- Queued jobs and queue wait: a growing queue indicates offered work exceeds available completion capacity over that interval.
- Session duration: examine percentiles and outliers to find pages or waits that hold slots too long.
- Outcomes: separate navigation timeouts, browser crashes, target-site blocks, and application errors.
- Cleanup failures: alert on sessions that remain active after a job completes or fails.
- Target-site pressure: set per-domain limits or pacing when a single destination should receive less parallel traffic than your browser fleet can support.
Raise the cap gradually only when the fleet has headroom and the target sites tolerate the added parallelism. If queue latency rises while workers are saturated, add measured capacity or reduce offered work; do not assume adding workers will improve a target site’s response time.
6. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Concurrency or capacity errors | Sessions are still open, the application cap exceeds provider capacity, or retries create extra connections. | Ensure unconditional cleanup, reduce the cap, and check live plan limits and provider capacity signals. |
| Jobs wait too long in queue | Arrival rate exceeds completion rate, sessions are long, or the configured cap is too low for available capacity. | Inspect queue wait and session duration; tune workload, timeouts, or capacity based on measurements. |
| Connection timeout or rejected WebSocket | Wrong endpoint, region mismatch, invalid credentials, network policy, or provider-side saturation. | Copy the current endpoint and authentication format from provider documentation; check region, credentials, network egress, and service pressure. |
| Proxy or profile settings appear missing | A newly created Playwright context may not inherit launch-level settings on some CDP integrations. | Check the provider’s integration guidance and try its default context behavior; validate on the exact endpoint and library versions deployed. |
| Navigation repeatedly times out | Slow target, blocked request, unsuitable readiness condition, or timeout shorter than queue plus execution time. | Separate queue and navigation timing, choose a readiness condition appropriate to the page, and set bounded timeouts with a retry policy. |
| Browser appears busy after a failed job | Exception path skipped closing the page, context, or remote session. | Put cleanup in finally; verify that the provider considers the session released, not just locally disconnected. |
| Browser crashes under load | The tested cap may exceed capacity for your browser build, page mix, or resource profile. | Lower concurrency, reproduce with representative pages, and scale workers only after load testing. Do not use a generic sessions-per-CPU estimate. |
7. Performance, reliability, and cost
- Performance: concurrency helps process independent jobs in parallel, but too many sessions can increase queueing, resource contention, and pressure on target sites. Test representative pages, browser builds, contexts, and resource profiles.
- Reliability: use bounded timeouts, unconditional cleanup, and a retry policy that does not multiply load during an outage. Make jobs idempotent where retries are possible, and preserve failures for inspection.
- Cost: compare managed plan limits and billing terms with the operational cost of running and maintaining a self-hosted fleet. The available documentation does not establish a universal break-even point. Verify current vendor quotas and pricing before committing.
- Volatile limits: provider concurrency allowances, maximum session durations, endpoint hosts, and regional availability may change. Check the current provider documentation during deployment planning rather than relying on copied numeric tables.
8. Deployment checklist
- [ ] Define the active-session cap and where it is enforced.
- [ ] Confirm current provider quota, queue behavior, request timeout, and maximum session duration, or document the equivalent self-hosted settings.
- [ ] Verify the WebSocket endpoint, region, authentication, and supported browser and Playwright versions.
- [ ] Ensure every success, timeout, cancellation, and exception path releases the session.
- [ ] Track active sessions, queue depth and wait, session duration, failures, and cleanup issues.
- [ ] Set target-site concurrency or pacing rules independently of fleet capacity.
- [ ] Load-test representative workloads before raising concurrency or scaling workers.
- [ ] Decide how retries, process restarts, and queued jobs are handled.
9. Or skip the browser setup
If the job is to capture a website image or PDF, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed before the shot; known consent platforms, newsletter popups, and chat widgets are removed, and each of those steps can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing; response headers report the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for free and get 1,000 screenshots a month with no card.
10. FAQ
Should the queue live in my application or at the browser provider?
An application queue gives you control over admission, per-target limits, visibility into local wait time, and job persistence. A provider queue can absorb bursts. You can use both, while validating how each queue affects timeouts and total latency.
Can I use one global cap for every target?
A global cap protects your browser capacity, but per-domain limits may also be needed because different sites tolerate different request rates.
How many sessions should one worker run?
Set the number from load tests for your actual browser, page mix, and deployment. The cited documentation provides no portable sessions-per-worker or per-CPU sizing figure.
Does a remote browser connection always release capacity when my script exits?
Do not assume so. Follow the provider’s session lifecycle instructions and explicitly close the relevant remote session in cleanup.


