ScreenshotNeo

BlogHow-to

How to Fix Memory Growth When Screenshotting HTML Pages in Java

Diagnose whether Java heap, native memory, or the browser is growing, then apply lifecycle, cache, output, and page-size fixes.

By the ScreenshotNeo team1 October 20269 min read

Memory growth during a screenshot loop is a symptom, not a diagnosis. First identify whether the increasing measurement belongs to the Java heap, native or off-heap JVM memory, or a separately running browser process. Then match the fix to that owner: close the correct resources, stop retaining screenshot buffers, reduce history or page scope, and investigate browser-side DOM growth when the JVM is stable.

A larger heap or RSS number after each capture does not by itself prove a leak. The useful signal is the retained live set after comparable garbage-collection points and a repeatable workload.

1. Identify which process and memory pool is growing

Java-hosted browsers such as HtmlUnit keep parsing, DOM, JavaScript, cookies, and networking inside the JVM. Playwright and Selenium usually control a separate browser process. A Java heap dump can reveal Java object retention, but it cannot explain all browser-native memory. Browser DevTools cannot explain every Java object retained by your application.

Observation Likely owner Next check
Used heap remains higher after full GC Java objects, image buffers, queues, page state Heap dumps and GC-root paths
RSS grows while heap is stable Native JVM memory, driver, or browser process Process-level metrics and browser task manager
Only the browser process grows DOM nodes, JavaScript objects, listeners, open pages or contexts Browser memory and heap snapshots
Growth tracks image dimensions Screenshot bytes, Base64 copies, encoders Output representation and retention lifetime

Build a repeatable measurement loop

  1. Capture the same representative URL repeatedly with the same viewport, scale, full-page setting, and concurrency.
  2. Record capture count, page dimensions, screenshot mode, output format, and whether results are held, queued, encoded, or persisted.
  3. Record Java heap used after comparable GC points, process RSS or native memory, and browser or renderer memory separately.
  4. Take measurements after warm-up. Heap committed and RSS do not need to fall immediately after every iteration.

Oracle’s Java 21 troubleshooting guide recommends heap dumps for leak analysis and describes Flight Recorder heap statistics for finding object growth over time. Use Oracle’s memory-leak procedure as the baseline.

# Replace 12345 with the Java process ID
jcmd 12345 GC.heap_dump /tmp/capture-before.dmp
# Run a known number of captures, then dump again
jcmd 12345 GC.heap_dump /tmp/capture-after.dmp

Compare the dumps by retained class and path to a GC root. Useful commands and tools include jcmd, jmap, JConsole, and -XX:+HeapDumpOnOutOfMemoryError. Do not increase -Xmx as the first fix; a larger heap can hide retention while increasing recovery time.

2. Close resources at the correct lifecycle boundary

Audit every object that owns state: browser, browser context or session, page or tab, driver, HTTP client, streams, image buffers, and application queues. The exact close method depends on the library version, so follow its matching API documentation.

HtmlUnit: close WebClient and limit history only when appropriate

HtmlUnit’s FAQ uses the wording “HtmlUnit appears to be leaking memory; what’s the deal?” Its documented checks are to use a current version and close WebClient, preferably with try-with-resources. The getting-started guide shows that a client owns browser state across page loads.

import com.gargoylesoftware.htmlunit.WebClient;
import com.gargoylesoftware.htmlunit.html.HtmlPage;

public final class HtmlUnitCapture {
  public static void main(String[] args) throws Exception {
    try (WebClient client = new WebClient()) {
      // Use these only when back-navigation/history is unnecessary.
      client.getOptions().setHistoryPageCacheLimit(0);
      client.getOptions().setHistorySizeLimit(0);

      HtmlPage page = client.getPage("https://example.com");
      // Render or screenshot page here, then release page references.
      System.out.println(page.getTitleText());
    }
  }
}

Setting both history limits to zero is a conditional reduction in retained history, not a universal leak cure. Keep history when the workflow needs back-navigation or page restoration. See the HtmlUnit FAQ and getting-started guide.

Playwright Java: close pages and contexts, and avoid unnecessary byte copies

import com.microsoft.playwright.*;
import java.nio.file.Paths;

public class PlaywrightCapture {
  public static void main(String[] args) {
    try (Playwright pw = Playwright.create()) {
      Browser browser = pw.chromium().launch();
      BrowserContext context = browser.newContext();
      Page page = context.newPage();
      page.navigate("https://example.com");
      page.screenshot(new Page.ScreenshotOptions()
          .setPath(Paths.get("shot.png"))
          .setFullPage(false)
          .setScale("css"));
      page.close();
      context.close();
      browser.close();
    }
  }
}

Playwright’s in-memory screenshot API returns a byte[]; the path API writes directly to disk. If downstream code does not need bytes, use a path and avoid retaining arrays or converting them to Base64. The Page API documents setScale("css") and setScale("device"). Device scale produces one pixel per device pixel and can make high-DPI images twice as large or more than CSS scale.

Selenium Java: inspect the selected output type

import org.openqa.selenium.OutputType;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;

public class SeleniumCapture {
  public static void main(String[] args) {
    WebDriver driver = new ChromeDriver();
    try {
      driver.get("https://example.com");
      // Choose one representation and release it after processing.
      java.io.File file = ((org.openqa.selenium.TakesScreenshot) driver)
          .getScreenshotAs(OutputType.FILE);
      System.out.println(file.getAbsolutePath());
    } finally {
      driver.quit();
    }
  }
}

TakesScreenshot also supports byte-array and Base64 outputs. Inspect where those values are stored and remove references after persistence or processing. See the Selenium TakesScreenshot API.

3. Stop retaining screenshot data in your application

A screenshot can be large even when the browser is healthy. A loop that appends every byte[], Base64 string, stream, or result object to a list will grow the Java heap by design. Base64 also expands binary data, and repeated conversions create temporary copies.

  • Write directly to a file or object store when byte access is unnecessary.
  • Process one result, flush it, and remove the reference before the next capture.
  • Bound producer and consumer queues; apply back-pressure instead of accumulating an unbounded backlog.
  • Do not keep page objects, response bodies, or exception objects in per-capture logs.
  • Use a bounded cache with an explicit byte limit and eviction policy.
Path output = Paths.get("shots", captureId + ".png");
Files.createDirectories(output.getParent());
page.screenshot(new Page.ScreenshotOptions().setPath(output));
// Do not add a byte[] copy to a long-lived collection.

4. Bound screenshot dimensions and scope

Full-page capture covers the entire scrollable page. A page with long feeds, canvases, or repeated lazy-loaded sections can produce a very large bitmap and trigger more browser work. Prefer the smallest scope that meets the requirement:

Requirement Choice Memory implication
One component Element screenshot Usually fewer pixels and less page stitching
Visible screen Viewport screenshot Bounded dimensions
Entire document Full-page screenshot Can grow with document height
Retina output Device scale More pixels and larger buffers
Normal density CSS scale Smaller output when acceptable

Set a maximum viewport and document height where your product requirements allow it. Inspect lazy-loading behavior: full-page capture may load images that are not present in a viewport capture. Compare PNG, JPEG, and WebP output according to fidelity and size needs.

5. Investigate browser-side growth separately

If Java heap remains stable but the browser process grows, inspect the page rather than taking more heap dumps. Chrome’s memory guide recommends Task Manager, memory timelines, heap snapshots, detached DOM-tree inspection, and retained-reference analysis.

  • Close pages and contexts after each isolated job.
  • Look for detached DOM nodes, global arrays, timers, event listeners, and continually appended application data.
  • Check whether the capture loop opens a new tab or context without closing the previous one.
  • Use a fresh context for untrusted or highly stateful pages when isolation matters.

6. Treat untrusted pages as a resource-control problem

Pathological pages can consume CPU, memory, network bandwidth, or time without any library defect. HtmlUnit states that parsing, DOM processing, JavaScript, and networking run in the hosting JVM and recommends limits for time, memory, CPU, page size, and requests. Apply equivalent controls around other stacks:

  • Set navigation and rendering timeouts.
  • Reject or truncate pages above a maximum response or rendered size.
  • Limit concurrent captures and queue depth.
  • Block unnecessary resource types, ads, trackers, or third-party requests where supported.
  • Run hostile workloads in an isolated process or container with operating-system limits.

Read HtmlUnit’s security guidance before processing untrusted content.

7. A diagnostic checklist

  1. Reproduce the growth with a fixed URL set and fixed capture settings.
  2. Measure post-GC Java live set, JVM native memory, process RSS, and browser RSS separately.
  3. Capture heap dumps at two comparable capture counts.
  4. Compare retained classes and GC-root paths.
  5. Audit closure for browser, context, page, driver, client, streams, and queues.
  6. Check screenshot representation: file, bytes, Base64, copies, and retention duration.
  7. Reduce full-page scope, device scale, dimensions, and concurrency one variable at a time.
  8. Use browser memory tools when only the browser process grows.
  9. Re-run the same workload after each change and compare the post-warm-up live set.

8. Common errors and fixes

Symptom Cause Fix
OutOfMemoryError: Java heap space Retained screenshot data, page objects, history, or queues Heap-dump comparison; close owners; bound queues; write files directly
Heap looks flat but RSS rises Native memory, driver, or separate browser process Measure processes separately and inspect browser memory tools
Growth only on full-page captures Large document, lazy images, or stitching buffers Use element or viewport scope, CSS scale, and bounded dimensions
Growth after every HtmlUnit navigation WebClient not closed or history retained Use current HtmlUnit, try-with-resources, and conditionally set both history limits to zero
Large spikes with Base64 Encoding expansion and duplicate strings Use file output or bytes, process immediately, and clear references
Browser tabs accumulate Pages or contexts are never closed Close at job completion; verify with browser task manager
Timeouts and huge pages Untrusted or pathological content Set time, size, request, CPU, memory, and concurrency limits

9. Performance, reliability, and cost considerations

Measure throughput together with memory. Lowering scale or page scope reduces pixels but may change visual requirements. Reusing a browser can reduce startup cost while retaining more state; creating isolated contexts can improve cleanup at additional startup cost. Choose based on measured warm-up behavior and required isolation.

Keep concurrency below the point where browser processes, image buffers, and queues compete for memory. Persist outputs incrementally so a failed job does not force the application to retain all prior screenshots. For long-running workers, record capture count and restart only when evidence shows unavoidable native growth; a restart is a containment measure, not a diagnosis.

10. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF, so the Java service does not need to manage a browser lifecycle for this capture path. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;

public class ScreenshotNeoCapture {
  public static void main(String[] args) throws Exception {
    String url = "https://stripe.com";
    String endpoint = "https://api.screenshotneo.com/v1/shot?access_key=YOUR_API_KEY&url="
        + java.net.URLEncoder.encode(url, java.nio.charset.StandardCharsets.UTF_8);
    HttpRequest request = HttpRequest.newBuilder(URI.create(endpoint)).GET().build();
    HttpResponse response = HttpClient.newHttpClient()
        .send(request, HttpResponse.BodyHandlers.ofByteArray());
    Files.write(Path.of("shot.webp"), response.body());
  }
}

Before capture, cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

11. FAQ

Does a rising heap number prove a leak?

No. Check the live retained set after comparable GC points and identify retained objects before changing heap size.

Should every screenshot loop create a new browser?

Not necessarily. Reuse can reduce startup work, while isolated contexts can simplify state cleanup. Measure the lifecycle that matches your workload.

When should HtmlUnit history limits be zero?

Only when back-navigation and history restoration are not required. The setting reduces retained history; it is not a general leak fix.

Why can a small webpage produce a large screenshot?

Device scale, full-page height, lazy-loaded images, canvases, and image format can increase pixel count and buffer size.

Can a heap dump explain Chrome memory?

No. A heap dump describes Java objects. Inspect the separate browser process with browser memory tools when browser memory grows.