How to Fix Memory Growth When Screenshotting HTML Pages in Java
Diagnose whether Java heap, native memory, or the browser is growing, then apply lifecycle, cache, output, and page-size fixes.
Memory growth during a screenshot loop is a symptom, not a diagnosis. First identify whether the increasing measurement belongs to the Java heap, native or off-heap JVM memory, or a separately running browser process. Then match the fix to that owner: close the correct resources, stop retaining screenshot buffers, reduce history or page scope, and investigate browser-side DOM growth when the JVM is stable.
A larger heap or RSS number after each capture does not by itself prove a leak. The useful signal is the retained live set after comparable garbage-collection points and a repeatable workload.
1. Identify which process and memory pool is growing
Java-hosted browsers such as HtmlUnit keep parsing, DOM, JavaScript, cookies, and networking inside the JVM. Playwright and Selenium usually control a separate browser process. A Java heap dump can reveal Java object retention, but it cannot explain all browser-native memory. Browser DevTools cannot explain every Java object retained by your application.
| Observation | Likely owner | Next check |
|---|---|---|
| Used heap remains higher after full GC | Java objects, image buffers, queues, page state | Heap dumps and GC-root paths |
| RSS grows while heap is stable | Native JVM memory, driver, or browser process | Process-level metrics and browser task manager |
| Only the browser process grows | DOM nodes, JavaScript objects, listeners, open pages or contexts | Browser memory and heap snapshots |
| Growth tracks image dimensions | Screenshot bytes, Base64 copies, encoders | Output representation and retention lifetime |
Build a repeatable measurement loop
- Capture the same representative URL repeatedly with the same viewport, scale, full-page setting, and concurrency.
- Record capture count, page dimensions, screenshot mode, output format, and whether results are held, queued, encoded, or persisted.
- Record Java heap used after comparable GC points, process RSS or native memory, and browser or renderer memory separately.
- Take measurements after warm-up. Heap committed and RSS do not need to fall immediately after every iteration.
Oracle’s Java 21 troubleshooting guide recommends heap dumps for leak analysis and describes Flight Recorder heap statistics for finding object growth over time. Use Oracle’s memory-leak procedure as the baseline.
# Replace 12345 with the Java process ID
jcmd 12345 GC.heap_dump /tmp/capture-before.dmp
# Run a known number of captures, then dump again
jcmd 12345 GC.heap_dump /tmp/capture-after.dmp
Compare the dumps by retained class and path to a GC root. Useful commands and tools include jcmd, jmap, JConsole, and -XX:+HeapDumpOnOutOfMemoryError. Do not increase -Xmx as the first fix; a larger heap can hide retention while increasing recovery time.
2. Close resources at the correct lifecycle boundary
Audit every object that owns state: browser, browser context or session, page or tab, driver, HTTP client, streams, image buffers, and application queues. The exact close method depends on the library version, so follow its matching API documentation.
HtmlUnit: close WebClient and limit history only when appropriate
HtmlUnit’s FAQ uses the wording “HtmlUnit appears to be leaking memory; what’s the deal?” Its documented checks are to use a current version and close WebClient, preferably with try-with-resources. The getting-started guide shows that a client owns browser state across page loads.
import com.gargoylesoftware.htmlunit.WebClient;
import com.gargoylesoftware.htmlunit.html.HtmlPage;
public final class HtmlUnitCapture {
public static void main(String[] args) throws Exception {
try (WebClient client = new WebClient()) {
// Use these only when back-navigation/history is unnecessary.
client.getOptions().setHistoryPageCacheLimit(0);
client.getOptions().setHistorySizeLimit(0);
HtmlPage page = client.getPage("https://example.com");
// Render or screenshot page here, then release page references.
System.out.println(page.getTitleText());
}
}
}
Setting both history limits to zero is a conditional reduction in retained history, not a universal leak cure. Keep history when the workflow needs back-navigation or page restoration. See the HtmlUnit FAQ and getting-started guide.
Playwright Java: close pages and contexts, and avoid unnecessary byte copies
import com.microsoft.playwright.*;
import java.nio.file.Paths;
public class PlaywrightCapture {
public static void main(String[] args) {
try (Playwright pw = Playwright.create()) {
Browser browser = pw.chromium().launch();
BrowserContext context = browser.newContext();
Page page = context.newPage();
page.navigate("https://example.com");
page.screenshot(new Page.ScreenshotOptions()
.setPath(Paths.get("shot.png"))
.setFullPage(false)
.setScale("css"));
page.close();
context.close();
browser.close();
}
}
}
Playwright’s in-memory screenshot API returns a byte[]; the path API writes directly to disk. If downstream code does not need bytes, use a path and avoid retaining arrays or converting them to Base64. The Page API documents setScale("css") and setScale("device"). Device scale produces one pixel per device pixel and can make high-DPI images twice as large or more than CSS scale.
Selenium Java: inspect the selected output type
import org.openqa.selenium.OutputType;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
public class SeleniumCapture {
public static void main(String[] args) {
WebDriver driver = new ChromeDriver();
try {
driver.get("https://example.com");
// Choose one representation and release it after processing.
java.io.File file = ((org.openqa.selenium.TakesScreenshot) driver)
.getScreenshotAs(OutputType.FILE);
System.out.println(file.getAbsolutePath());
} finally {
driver.quit();
}
}
}
TakesScreenshot also supports byte-array and Base64 outputs. Inspect where those values are stored and remove references after persistence or processing. See the Selenium TakesScreenshot API.
3. Stop retaining screenshot data in your application
A screenshot can be large even when the browser is healthy. A loop that appends every byte[], Base64 string, stream, or result object to a list will grow the Java heap by design. Base64 also expands binary data, and repeated conversions create temporary copies.
- Write directly to a file or object store when byte access is unnecessary.
- Process one result, flush it, and remove the reference before the next capture.
- Bound producer and consumer queues; apply back-pressure instead of accumulating an unbounded backlog.
- Do not keep page objects, response bodies, or exception objects in per-capture logs.
- Use a bounded cache with an explicit byte limit and eviction policy.
Path output = Paths.get("shots", captureId + ".png");
Files.createDirectories(output.getParent());
page.screenshot(new Page.ScreenshotOptions().setPath(output));
// Do not add a byte[] copy to a long-lived collection.
4. Bound screenshot dimensions and scope
Full-page capture covers the entire scrollable page. A page with long feeds, canvases, or repeated lazy-loaded sections can produce a very large bitmap and trigger more browser work. Prefer the smallest scope that meets the requirement:
| Requirement | Choice | Memory implication |
|---|---|---|
| One component | Element screenshot | Usually fewer pixels and less page stitching |
| Visible screen | Viewport screenshot | Bounded dimensions |
| Entire document | Full-page screenshot | Can grow with document height |
| Retina output | Device scale | More pixels and larger buffers |
| Normal density | CSS scale | Smaller output when acceptable |
Set a maximum viewport and document height where your product requirements allow it. Inspect lazy-loading behavior: full-page capture may load images that are not present in a viewport capture. Compare PNG, JPEG, and WebP output according to fidelity and size needs.
5. Investigate browser-side growth separately
If Java heap remains stable but the browser process grows, inspect the page rather than taking more heap dumps. Chrome’s memory guide recommends Task Manager, memory timelines, heap snapshots, detached DOM-tree inspection, and retained-reference analysis.
- Close pages and contexts after each isolated job.
- Look for detached DOM nodes, global arrays, timers, event listeners, and continually appended application data.
- Check whether the capture loop opens a new tab or context without closing the previous one.
- Use a fresh context for untrusted or highly stateful pages when isolation matters.
6. Treat untrusted pages as a resource-control problem
Pathological pages can consume CPU, memory, network bandwidth, or time without any library defect. HtmlUnit states that parsing, DOM processing, JavaScript, and networking run in the hosting JVM and recommends limits for time, memory, CPU, page size, and requests. Apply equivalent controls around other stacks:
- Set navigation and rendering timeouts.
- Reject or truncate pages above a maximum response or rendered size.
- Limit concurrent captures and queue depth.
- Block unnecessary resource types, ads, trackers, or third-party requests where supported.
- Run hostile workloads in an isolated process or container with operating-system limits.
Read HtmlUnit’s security guidance before processing untrusted content.
7. A diagnostic checklist
- Reproduce the growth with a fixed URL set and fixed capture settings.
- Measure post-GC Java live set, JVM native memory, process RSS, and browser RSS separately.
- Capture heap dumps at two comparable capture counts.
- Compare retained classes and GC-root paths.
- Audit closure for browser, context, page, driver, client, streams, and queues.
- Check screenshot representation: file, bytes, Base64, copies, and retention duration.
- Reduce full-page scope, device scale, dimensions, and concurrency one variable at a time.
- Use browser memory tools when only the browser process grows.
- Re-run the same workload after each change and compare the post-warm-up live set.
8. Common errors and fixes
| Symptom | Cause | Fix |
|---|---|---|
OutOfMemoryError: Java heap space |
Retained screenshot data, page objects, history, or queues | Heap-dump comparison; close owners; bound queues; write files directly |
| Heap looks flat but RSS rises | Native memory, driver, or separate browser process | Measure processes separately and inspect browser memory tools |
| Growth only on full-page captures | Large document, lazy images, or stitching buffers | Use element or viewport scope, CSS scale, and bounded dimensions |
| Growth after every HtmlUnit navigation | WebClient not closed or history retained | Use current HtmlUnit, try-with-resources, and conditionally set both history limits to zero |
| Large spikes with Base64 | Encoding expansion and duplicate strings | Use file output or bytes, process immediately, and clear references |
| Browser tabs accumulate | Pages or contexts are never closed | Close at job completion; verify with browser task manager |
| Timeouts and huge pages | Untrusted or pathological content | Set time, size, request, CPU, memory, and concurrency limits |
9. Performance, reliability, and cost considerations
Measure throughput together with memory. Lowering scale or page scope reduces pixels but may change visual requirements. Reusing a browser can reduce startup cost while retaining more state; creating isolated contexts can improve cleanup at additional startup cost. Choose based on measured warm-up behavior and required isolation.
Keep concurrency below the point where browser processes, image buffers, and queues compete for memory. Persist outputs incrementally so a failed job does not force the application to retain all prior screenshots. For long-running workers, record capture count and restart only when evidence shows unavoidable native growth; a restart is a containment measure, not a diagnosis.
10. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF, so the Java service does not need to manage a browser lifecycle for this capture path. See the ScreenshotNeo API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
public class ScreenshotNeoCapture {
public static void main(String[] args) throws Exception {
String url = "https://stripe.com";
String endpoint = "https://api.screenshotneo.com/v1/shot?access_key=YOUR_API_KEY&url="
+ java.net.URLEncoder.encode(url, java.nio.charset.StandardCharsets.UTF_8);
HttpRequest request = HttpRequest.newBuilder(URI.create(endpoint)).GET().build();
HttpResponse response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofByteArray());
Files.write(Path.of("shot.webp"), response.body());
}
}
Before capture, cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
11. FAQ
Does a rising heap number prove a leak?
No. Check the live retained set after comparable GC points and identify retained objects before changing heap size.
Should every screenshot loop create a new browser?
Not necessarily. Reuse can reduce startup work, while isolated contexts can simplify state cleanup. Measure the lifecycle that matches your workload.
When should HtmlUnit history limits be zero?
Only when back-navigation and history restoration are not required. The setting reduces retained history; it is not a general leak fix.
Why can a small webpage produce a large screenshot?
Device scale, full-page height, lazy-loaded images, canvases, and image format can increase pixel count and buffer size.
Can a heap dump explain Chrome memory?
No. A heap dump describes Java objects. Inspect the separate browser process with browser memory tools when browser memory grows.


