How to Improve Application Performance with an Open-Source Load Balancer
Learn how open-source load balancers improve latency, throughput, and resilience, then tune routing, health checks, connections, and capacity safely.
Direct answer: An open-source load balancer can improve application performance by spreading requests across multiple application instances, choosing a routing policy that matches request behavior, reusing upstream connections, and stopping traffic to failed servers. It cannot make slow application code fast by itself. Measure the current bottleneck, establish a representative baseline, change one variable at a time, and verify latency, throughput, errors, and resource use.
This guide covers NGINX Open Source, HAProxy, and Envoy concepts. The exact directives and features depend on the product version and edition, so validate examples against the software you deploy.
1. What a load balancer changes
A load balancer sits between clients and application instances. It accepts client connections, selects an upstream server, forwards the request, and returns the response. With more than one healthy instance, the system can use available CPU and memory across machines, handle more concurrent work, and continue serving when an instance fails. NGINX describes the goals as better resource utilization, throughput, latency, and fault tolerance.
| Potential improvement | How it happens | What can prevent it |
|---|---|---|
| Throughput | Requests run concurrently on several instances. | The application, database, network, or balancer is already saturated. |
| Latency | Work avoids overloaded instances and may use a closer or faster backend. | Queueing, slow code, lock contention, or an unsuitable routing policy. |
| Tail latency | Busy or unhealthy servers receive less work. | Health checks detect only process availability, not degraded requests. |
| Availability | Failed instances are removed from rotation. | Incorrect checks, slow failure detection, or no redundant balancer. |
| Connection overhead | Persistent upstream connections avoid repeated handshakes. | Excess idle connections, file-descriptor limits, or incompatible backend timeouts. |
2. Measure before tuning
Record a baseline during representative traffic. Include both the load balancer and every application tier.
- Request rate and throughput, separated by endpoint and request type.
- Median, p95, p99, and maximum latency.
- HTTP status and transport error rates.
- Active, idle, queued, and reused connections.
- Backend CPU, memory, garbage collection, database time, and saturation.
- Load-balancer CPU, memory, network, file descriptors, accept queues, and worker utilization.
- TLS handshake rate, request-body sizes, response sizes, and protocol mix.
Use realistic concurrency and include slow requests, large responses, authenticated traffic, TLS, persistence requirements, and heterogeneous backends. A higher request count is not a performance win if p99 latency or errors increase.
3. Choose a routing algorithm
Round robin
Round robin sends requests to servers in sequence and is NGINX’s default when no method is specified. It works well when instances have similar capacity and requests consume similar work. Equal request counts do not guarantee equal CPU time: one request may finish in milliseconds while another runs for seconds.
Least connections
Least connections selects the server with fewer active connections. It is useful when request duration varies and active connections are a reasonable proxy for current work. Long-lived streaming or WebSocket connections can distort the signal, so monitor actual backend load.
Least time
Some products can select using response-time measurements plus active-connection data. The timing signal may be time to first byte, full response, or full response while accounting for in-flight requests. Choose the signal that matches the user-visible objective and validate it under your workload.
Weights
Assign higher weights to larger instances when capacity differs. Treat weights as an initial distribution, not proof of proportional work. Confirm CPU, memory, queueing, and latency on each backend after deployment.
IP hash and affinity
IP-hash routing keeps a client on the same server when possible. It can help legacy session storage, but shared NAT addresses can concentrate many users on one instance and mobile clients can change addresses. Prefer shared session storage or signed, portable sessions when the application permits it.
Other policies
Envoy documents weighted round robin, Maglev, least-loaded, and random policies, with endpoints supplied by static configuration, DNS, or dynamic xDS. Select a policy based on routing requirements and measured behavior rather than a universal ranking.
4. Configure health checks and failure handling
A health check should represent the ability to serve the traffic you send. A TCP port check can pass while the application is deadlocked or unable to reach its database. Use a lightweight HTTP endpoint with a defined status and response contract, and avoid checks that trigger expensive business work.
NGINX Open Source documents passive checks: failed responses cause the proxy to avoid a server for a period, then live requests probe recovery. Its max_fails and fail_timeout settings control this behavior; setting max_fails to zero disables checks. Periodic active HTTP checks and dynamic group-management features are documented as NGINX Plus capabilities. Confirm edition support before relying on them.
Minimal NGINX configuration
http {
upstream app_pool {
# Round robin is the default.
server app1.internal:8080 max_fails=3 fail_timeout=10s;
server app2.internal:8080 max_fails=3 fail_timeout=10s;
server app3.internal:8080 max_fails=3 fail_timeout=10s;
}
server {
listen 80;
location / {
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_pass http://app_pool;
}
}
}
Test syntax with nginx -t, reload gracefully, and watch upstream status and latency. For least connections, add least_conn; inside the upstream block. For unequal capacity, add a weight such as server app1.internal:8080 weight=3;.
5. Reuse upstream connections safely
Keep-alive and connection pools reduce TCP and TLS setup overhead. They also retain idle sockets, memory, and file descriptors. If a backend closes a pooled connection unexpectedly, the next request can fail unless the proxy retries safely. Align idle timeouts, maximum concurrent streams, retry rules, and circuit breakers with backend capacity.
Envoy connection pools can reuse endpoint connections and multiplex HTTP/2 streams, subject to concurrent-stream limits and circuit breakers. HAProxy documentation describes several http-reuse modes; more aggressive reuse can reduce CPU work while consuming more idle connections and increasing failure risk when backend behavior conflicts. Verify directive availability in the exact HAProxy edition and version.
6. Compression and caching
Compression can reduce transfer time for clients on slow or high-latency links, at the cost of CPU. Compress text formats selectively, avoid recompressing already compressed media, and check response-size and CPU changes. HAProxy documents an in-memory cache that avoids repeat transfers while objects remain valid; it is a helper, not a replacement for a full caching layer. Ensure cache keys, authorization, invalidation, and privacy rules are correct before enabling it.
7. Benchmark changes under realistic load
- Capture a baseline for each important endpoint and request class.
- Reproduce the same traffic mix with the same concurrency, payload sizes, protocol versions, and backend count.
- Change one setting, such as the balancing method or keep-alive limit.
- Run long enough to include cache warm-up, connection reuse, garbage collection, and steady state.
- Compare p50, p95, p99, errors, throughput, and saturation on every tier.
- Repeat with one backend slow, unavailable, or returning errors.
Do not copy hardware-specific kernel or buffer values from a tuning article. HAProxy’s tuning guidance emphasizes that file descriptors, queues, buffers, connection limits, and reuse settings trade CPU, memory, and capacity. Set them from measured traffic and monitor after every change.
8. Capacity and high availability
If the balancer is the bottleneck, add capacity to the balancer tier and design for its failure. Active/active and active/standby arrangements address different availability and operational needs. Account for connection state, health-check coordination, DNS or anycast behavior, certificate distribution, and draining during maintenance. A single highly tuned balancer is still a single failure domain.
9. Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| Latency rises after adding servers | Round robin sends long requests evenly by count, not by work; or a shared database is saturated. | Compare least connections and weighted policies, then inspect backend and database saturation. |
| One backend receives most traffic | IP affinity, unequal weights, stale DNS, or long-lived connections. | Inspect routing state and connection age; remove unnecessary affinity and correct weights. |
| Intermittent 502/503 responses | Backend resets pooled sockets, health thresholds are too low, or connect/read timeouts are mismatched. | Align timeouts, inspect reset logs, tune reuse conservatively, and verify retry safety. |
| Checks pass but users see errors | The check tests only a port or shallow process state. | Expose a cheap endpoint that verifies required dependencies and expected status. |
| File-descriptor exhaustion | Too many idle client or upstream connections. | Raise limits only with capacity evidence; cap idle pools and monitor descriptors. |
| CPU increases after enabling compression | Compression work exceeds network savings. | Limit MIME types and sizes, choose an appropriate compression level, and compare p99 latency. |
| HTTP/2 streams fail under load | Concurrent-stream or circuit-breaker limits conflict with backend capacity. | Set limits with backend concurrency measurements and watch queueing. |
| Deployments cause dropped requests | Instances are removed without connection draining. | Drain existing connections, stop new assignments, then terminate after an appropriate grace period. |
| Performance differs from a published example | Different OS, version, TLS mix, hardware, or request distribution. | Use external figures only as context; benchmark your deployment. |
10. Cost and reliability considerations
Open-source software licensing does not remove infrastructure costs. Budget for redundant balancers, compute, bandwidth, TLS certificates, observability, on-call work, and capacity headroom. HAProxy and NGINX open-source editions have different feature boundaries from HAProxy Enterprise and NGINX Plus; choose a paid edition only when its documented operational features are required. No universal hardware size is responsible without concurrency, throughput, TLS, topology, and failure requirements.
11. A practical way to observe application pages
After tuning a load balancer, teams often need repeatable screenshots of key routes to check visual regressions, redirects, login states, or cache behavior. You can run a browser yourself with Playwright or Selenium, wait for the page, handle consent banners, hide overlays, and save an image. That approach gives control but adds browser binaries, concurrency limits, cleanup, and failure handling to your system.
12. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the verdict and billing status with X-Page-Verdict and X-Billed headers.
See the complete option reference in the ScreenshotNeo documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const fs = require('node:fs/promises');
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-element capture, dark mode, device presets, custom viewports, retina scale, PDF paper settings and page ranges, HTML/CSS rendering, custom JavaScript and CSS, clicks, selector or network-idle waits, request blocking, custom headers and cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
13. Frequently asked questions
Will a load balancer always make an application faster?
No. It improves distribution and resilience when multiple instances have spare capacity. A slow query, lock, dependency, or saturated network remains slow behind a balancer.
Which algorithm should I start with?
Start with round robin for similar, short-lived requests. Test least connections when request durations vary, and use weights for unequal instances. Keep the policy that improves your measured tail latency without raising errors or saturation.
Do I need active health checks?
Not always. Passive checks can detect failures from real traffic. Active checks are useful when you need periodic probing before a user request reaches an instance, but availability depends on the product edition and deployment.
Should I enable the most aggressive connection reuse?
No. Reuse can reduce setup cost while increasing idle resource use and reset-related failures. Tune it with connection, descriptor, memory, and error measurements.
Is caching at the load balancer enough?
Usually not for complex applications. A small in-memory helper can avoid repeat transfers, but authorization, invalidation, object size, and freshness often require a dedicated cache design.
How do I prove a tuning change helped?
Run the same representative workload before and after the change, then compare throughput, p95 and p99 latency, errors, and saturation on the balancer and backends. Include failure and recovery tests.


