How to Debug an Infinite Loop in Node.js Production Code
Find out whether Node.js is stuck in a synchronous loop, capture evidence safely, locate the hot function, and ship a reliable fix.

Direct answer: treat an apparent infinite loop as an incident first. Confirm whether the process is CPU-bound or waiting on I/O, preserve the affected process and deployment context, capture a diagnostic report when your environment allows it, and collect a CPU profile or flamegraph. Then inspect the hottest stack frames and reproduce the behavior with the same class of input. A profile shows where time was sampled; source inspection proves whether a loop cannot terminate.
A synchronous loop that never yields runs on Node.js’s JavaScript execution thread. While it runs, callbacks, timers, socket events, and incoming requests cannot be processed. Clinic.js describes the event loop this way: “The event loop is single-threaded: only one operation is processed at a time.” A function that schedules setTimeout and then returns is different: the callback runs on a later event-loop turn after the synchronous function completes. See the official Clinic.js documentation and Doctor guide.
1. Decide whether this is really an infinite loop
Several failures look identical to a caller: requests stop completing, health checks time out, and latency rises. Their causes differ, so start with observable symptoms.

| Signal | Likely direction | What to check |
|---|---|---|
| Sustained high CPU in one process | Synchronous JavaScript, a tight native call, or repeated computation | CPU profile, hot stacks, recent code and input changes |
| Low CPU while requests wait | Slow or unavailable asynchronous dependency | Database, HTTP, DNS, queue, timeout and connection telemetry |
| Memory grows continuously | Unbounded retention, queue growth, or a loop producing objects | Heap data, allocation paths and backlog size |
| Only one route or job stalls | Input-dependent path or route-specific synchronous work | Request parameters, payload size and recent feature flags |
Do not label every expensive operation an infinite loop. A parser walking a very large document, a retry loop with a long bound, recursive traversal of a cyclic graph, and a loop that eventually terminates can all consume CPU for a long time. Use “infinite loop” only after the termination logic is shown to fail or the process remains in the same path under a bounded reproduction.
2. Establish scope and preserve incident context
- Record the first observed time, affected service version, instance or pod, routes or jobs, and the input class involved.
- Compare affected and healthy instances. A single hot process suggests a bad request, corrupted state, or uneven workload; a fleet-wide event suggests a deployment or configuration change.
- Check recent releases, feature flags, parser changes, retry-policy changes and dependency updates.
- Follow your incident procedure before sending signals, pausing traffic or restarting a process. The correct action depends on your process manager, container runtime and service-level policy.
Capture request IDs, deployment identifiers and coarse timestamps before collecting detailed diagnostics. Avoid adding high-volume synchronous logging while the event loop is already blocked; logging can increase the work and obscure the original path.
3. Capture a Node.js diagnostic report
Node.js diagnostic reports are designed for development, testing and production problem determination. A report can include JavaScript and native stacks, heap information, libuv handles, platform details and resource data. Read the version-specific Node.js diagnostic report documentation before relying on an option in production.
You can configure report generation at process start or trigger it programmatically. The exact signal, file destination and startup flags vary by Node.js version and deployment, so use the mechanism supported by your runtime and operational policy. A minimal in-process trigger looks like this:
const report = require('node:process/report');
// Call from a controlled diagnostic endpoint or incident handler.
const filename = report.writeReport();
console.error(`Diagnostic report written to ${filename}`);
Protect any endpoint that can write a report. Reports may contain environment details, file paths, stack data and resource metadata. Store them according to your data-retention and access rules, and transfer them off the affected host only through approved channels.
4. Profile CPU time and event-loop symptoms
A CPU profile samples stacks during a time window. In a flamegraph, wide frames represent a large share of sampled CPU time and repeated application frames are useful clues. They are not proof of non-termination: a profile can capture a legitimate but expensive operation. Clinic.js Flame documents CPU collection and flamegraph analysis; Clinic.js also supports collection-only workflows so data can be visualized away from the server. Confirm compatibility with your exact Node.js and operating-system versions before using these tools in an incident.

Profile a representative reproduction
Reproducing outside production is usually the safest way to collect a longer profile. Keep the payload shape, feature flags and dependency behavior close to the failing case.
// loop-repro.js
function transform(items) {
let index = 0;
while (index < items.length) {
// Replace with the suspected production operation.
items[index] = JSON.parse(JSON.stringify(items[index]));
index += 1;
}
return items;
}
const items = Array.from({ length: 10000 }, (_, index) => ({ index }));
console.log(transform(items).length);
Run the profiler according to its current documentation, collect a bounded interval, and inspect the result off-box when possible. If your team uses Visual Studio Code, its JavaScript profiling documentation explains how to open .cpuprofile files and inspect CPU views.
Use event-loop evidence carefully
High CPU combined with delayed timers and request callbacks is consistent with a non-yielding synchronous path. Low CPU combined with pending promises points elsewhere. Correlate event-loop delay metrics, CPU, request latency and dependency timings rather than using one metric as a verdict.
5. Read the hot stack and inspect termination logic
Start at the widest application frame, then follow callers until you reach the request, job or message handler that supplied the input. Inspect these failure patterns:
- Control variable never changes:
while (i < n)executes without incrementingi, or mutation occurs only on a branch that is never reached. - Wrong comparison direction: a decrementing counter is compared with
<, or a signed value wraps unexpectedly. - Retry without a bound: a failed operation immediately retries forever instead of applying a maximum attempt count and backoff.
- Recursive traversal of cyclic data: parent links or graph edges revisit an already-seen node.
- Unexpectedly large input: a parser, sort, regular expression or nested traversal is finite but too expensive for a request.
- Repeated synchronous work per request: a cache miss causes the same CPU-heavy calculation to run concurrently for every caller.
Add a termination invariant to the fix. For a counter loop, state what changes each iteration and why it must cross the bound. For graph traversal, track visited nodes. For retries, define maximum attempts, total elapsed time and a terminal error.
function retry(operation, { maxAttempts = 5, delayMs = 100 } = {}) {
let attempt = 0;
return (async () => {
while (attempt < maxAttempts) {
attempt += 1;
try {
return await operation(attempt);
} catch (error) {
if (attempt >= maxAttempts) throw error;
await new Promise(resolve => setTimeout(resolve, delayMs * attempt));
}
}
throw new Error('unreachable: retry loop exhausted without returning');
})();
}
6. Mitigate the live incident, then fix the code
Use the service’s playbook to shed the affected workload, isolate the instance, roll back a suspect deployment, or replace an unhealthy process. A restart can restore capacity but destroys the most useful live state, so capture approved evidence first when the process is stable enough. If the loop is caused by one input, quarantine that input or route while you prepare the code fix.
For CPU-heavy work that is valid but too expensive for the request thread, bound the work, process in chunks and yield between chunks, or move the computation to a worker process or queue. Yielding does not repair a broken termination condition; it only prevents one task from monopolizing the event loop while bounded work proceeds.
async function processInBatches(items, batchSize = 500) {
for (let offset = 0; offset < items.length; offset += batchSize) {
const batch = items.slice(offset, offset + batchSize);
processBatch(batch);
// Give timers and I/O a chance to run.
await new Promise(resolve => setImmediate(resolve));
}
}
Roll out the correction with a representative reproduction, bounded load and the same runtime version as production. Confirm that CPU, event-loop delay, request latency and error rates return to their normal service-specific ranges.
7. Troubleshooting checklist
| Problem | Cause | Fix |
|---|---|---|
| Profiler shows no obvious loop | The sample window missed an intermittent path, or the bottleneck is native or asynchronous. | Repeat during the symptom, correlate with inputs, and inspect dependency timings and diagnostic-report stacks. |
| CPU is normal but requests hang | Waiting on a dependency, lock, socket or promise. | Inspect timeout, connection-pool, DNS, database and upstream metrics before changing loop code. |
| Restart “fixes” the issue temporarily | State, cache contents or a recurring input triggers the path again. | Preserve evidence before restart and reproduce with the same state or payload class. |
| Diagnostic report cannot be written | Unsupported runtime option, permissions, read-only filesystem or policy restriction. | Check the Node.js version documentation, destination permissions and deployment policy; use an approved alternate capture path. |
| Adding logs makes the outage worse | Synchronous formatting or excessive output adds event-loop work and I/O pressure. | Use sampled, structured signals and capture a bounded profile instead. |
| Fix passes unit tests but fails in production | Production input size, flags or dependency behavior differs. | Replay sanitized production-shaped inputs and verify under realistic concurrency. |
8. Performance, reliability and cost considerations
Live capture and reproduction answer different questions. A live diagnostic report gives broad state—stacks, handles, heap and platform data—but may expose sensitive information and has operational risk. A reproduction profile gives focused CPU attribution with more control, but it can miss production-only state. Collection-only profiling lets you visualize away from the server. None of the cited documentation establishes a universal overhead percentage, so measure in a staging-like environment and keep capture windows bounded.
Make diagnostics repeatable: record runtime and tool versions, profile duration, workload shape and whether the process was healthy or degraded. Keep report files and profiles linked to the incident timeline. For recurring jobs, add explicit maximum work, attempt and elapsed-time limits so a future bug fails closed instead of consuming the request thread indefinitely.
9. Or skip the browser setup
If your incident work includes collecting screenshots of dashboards, error pages or reproduction URLs, ScreenshotNeo provides a direct API request instead of maintaining a browser. It accepts a URL and returns PNG, JPEG, WebP or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', buffer);
ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, device presets and custom viewports, dark mode, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.
10. FAQ
Does 100% CPU prove an infinite loop?
No. It proves the process is doing substantial CPU work. A finite but expensive parser, sort or serialization path can produce the same symptom. Use a profile and source inspection to establish termination behavior.
Should I use a diagnostic report or a CPU profile first?
Use the safest capture supported by your incident policy. Reports provide broad runtime state; profiles focus on CPU attribution. If the process is stable, collect both with bounded duration and controlled handling.
Can adding setTimeout fix the loop?
Only if the work is already bounded and you are deliberately yielding between batches. A timeout does not fix a condition that never changes.
When should work move to a worker?
Move valid, CPU-heavy work when bounding and chunking still threatens request latency. Keep cancellation, limits and error handling explicit in the worker protocol.
What should I retain after the incident?
Keep the sanitized reproduction, runtime and tool versions, profile or report references, triggering input class, code change and verification results. That record shortens the next investigation without retaining unnecessary production data.


