ScreenshotNeo

BlogHow-to

How to Fix Headless Chrome Crashes When Printing Large HTML Files to PDF on Heroku

Diagnose Chrome startup failures, R14 memory events, cache problems and H12 timeouts when generating large PDFs with Puppeteer on Heroku.

By the ScreenshotNeo team1 October 20269 min read

Large HTML-to-PDF jobs on Heroku usually fail for one of five reasons: Chrome cannot start, the browser cache or shared libraries are missing, the dyno runs out of memory, the request exceeds Heroku’s router windows, or browser/page objects are leaked between jobs. Classify the log first, then apply the fix that matches the failure.

The reliable baseline is a maintained Chrome for Testing buildpack, Puppeteer launched with --headless and --no-sandbox, one render at a time while you measure RSS, strict page and browser cleanup, and an asynchronous worker for documents that can exceed the web request window.

1. Identify the failure signature

Log or symptom Likely cause First action
Chrome failed to launch, missing shared library Unsupported Chrome installation or missing system dependencies Use the Chrome for Testing buildpack and verify the executable on PATH.
Failed to move to new namespace or sandbox error Chrome sandbox cannot run inside the dyno Launch with --no-sandbox.
Executable not found after upgrading Puppeteer Puppeteer v19+ cache location changed Move /app/.cache/puppeteer into the slug cache and clear the Heroku build cache.
R14 Process exceeded dyno memory and is paging to swap Reduce concurrency and document complexity; measure RSS before resizing the dyno.
R15 or SIGKILL Memory exceeded the dyno limit by a large margin Stop concurrent jobs, reduce the render, or move to a larger dyno or managed renderer.
H12 while Chrome is still busy Synchronous PDF generation exceeded Heroku router timing Queue the job and render it in a worker.
Memory rises after every successful job Pages, browsers, or application objects are not released Close objects in finally blocks and recycle the browser after a bounded number of jobs.

Record the complete log line, Puppeteer version, Chrome version, dyno type, document size, page count, asset count, render duration, RSS, swap, and browser process count. That record prevents treating a router timeout as a Chrome crash.

2. Install Chrome correctly on Heroku

Puppeteer’s official troubleshooting guide recommends the Puppeteer Heroku buildpack and says Heroku launches need --no-sandbox mode (Puppeteer troubleshooting). The maintained Chrome for Testing buildpack installs Chrome and Chromedriver and places them on PATH. Its documented baseline flags are --headless and --no-sandbox; some workloads also use --disable-gpu or --remote-debugging-port=9222.

Do not start a new deployment with the deprecated heroku/google-chrome buildpack. Heroku’s documentation points new applications to Chrome for Testing, which does not inject the old shim flags.

Buildpack setup

heroku buildpacks:clear
heroku buildpacks:add heroku/nodejs
heroku buildpacks:add https://github.com/heroku/heroku-buildpack-chrome-for-testing
heroku config:set PUPPETEER_SKIP_DOWNLOAD=true

git push heroku main

Use the buildpack order required by its current documentation. After deployment, open a one-off dyno and verify the executable:

heroku run bash
which chrome
chrome --version
node --version
npm ls puppeteer

3. Handle Puppeteer v19+ browser caching

Puppeteer v19 changed its browser cache behavior. The Puppeteer Heroku buildpack README documents a heroku-postbuild pattern that moves /app/.cache/puppeteer into the slug’s local .cache directory. Choose one cache strategy, confirm that the browser exists in the deployed slug, and redeploy after clearing stale build artifacts.

{
  "scripts": {
    "heroku-postbuild": "mkdir -p .cache && cp -R /app/.cache/puppeteer .cache/"
  }
}

If logs show a missing executable or library immediately after a buildpack or Puppeteer change:

  1. Run heroku plugins only if you use a cache-clearing plugin supported by your team’s deployment process.
  2. Clear the Heroku build cache using your established Heroku workflow.
  3. Redeploy without changing application code.
  4. Check which chrome, chrome --version, and the cache directory again.

4. Launch Puppeteer with safe defaults

const express = require('express');
const puppeteer = require('puppeteer');

const app = express();
app.use(express.json({ limit: '2mb' }));

const chromeArgs = [
  '--headless',
  '--no-sandbox',
  '--disable-setuid-sandbox'
];

async function renderPdf(html) {
  const browser = await puppeteer.launch({
    headless: true,
    args: chromeArgs,
    timeout: 30_000
  });

  let page;
  try {
    page = await browser.newPage();
    await page.setContent(html, {
      waitUntil: 'networkidle0',
      timeout: 60_000
    });
    return await page.pdf({
      format: 'A4',
      printBackground: true,
      preferCSSPageSize: true,
      margin: { top: '20mm', right: '16mm', bottom: '20mm', left: '16mm' },
      timeout: 120_000
    });
  } finally {
    if (page) await page.close().catch(() => {});
    await browser.close().catch(() => {});
  }
}

app.post('/pdf', async (req, res) => {
  if (typeof req.body.html !== 'string' || req.body.html.length === 0) {
    return res.status(400).json({ error: 'html must be a non-empty string' });
  }

  try {
    const pdf = await renderPdf(req.body.html);
    res.type('application/pdf').send(pdf);
  } catch (error) {
    console.error('pdf-render-failed', error);
    res.status(500).json({ error: 'PDF generation failed' });
  }
});

const port = process.env.PORT || 3000;
app.listen(port, () => console.log(`listening on ${port}`));

--disable-setuid-sandbox is commonly paired with --no-sandbox. Add --disable-gpu only when logs or measurements show a GPU initialization problem. Add --remote-debugging-port=9222 only when you need remote debugging; it is not a general crash fix. Avoid --single-process as a blanket recommendation because it changes Chrome’s process model and can make failures less isolated.

5. Control memory during large renders

Heroku emits R14 when a Node process exceeds its dyno memory quota and begins paging; R15 indicates a much larger exceedance and the process can be killed (Node memory use, error codes). Chrome, the Node heap, decoded images, fonts, PDF buffers, and temporary objects all share the dyno’s memory.

Limit concurrency

let active = 0;
const queue = [];
const MAX_CONCURRENT_RENDERS = 1;

function schedule(job) {
  return new Promise((resolve, reject) => {
    queue.push({ job, resolve, reject });
    drain();
  });
}

async function drain() {
  if (active >= MAX_CONCURRENT_RENDERS || queue.length === 0) return;
  active++;
  const item = queue.shift();
  try {
    item.resolve(await item.job());
  } catch (error) {
    item.reject(error);
  } finally {
    active--;
    drain();
  }
}

Start with one simultaneous page and one browser. Increase only after repeated measurements show headroom. Large images, web fonts, SVG filters, canvases, huge tables, and client-side chart libraries can dominate memory even when the HTML source is small.

Measure every job

function memoryMb() {
  const m = process.memoryUsage();
  return {
    rss: Math.round(m.rss / 1024 / 1024),
    heapUsed: Math.round(m.heapUsed / 1024 / 1024),
    external: Math.round(m.external / 1024 / 1024)
  };
}

async function measuredRender(render) {
  const started = Date.now();
  const before = memoryMb();
  try {
    return await render();
  } finally {
    console.log(JSON.stringify({
      before,
      after: memoryMb(),
      durationMs: Date.now() - started
    }));
  }
}

Run the same large document repeatedly with concurrency one. If RSS does not return near its previous baseline, inspect event listeners, retained HTML, buffers, pages, and browsers. Recycle a browser after a bounded number of jobs when measurements show retained growth. Always close each page and close the browser during worker shutdown.

6. Avoid Heroku router timeouts

Heroku gives a web request an initial 30-second window to return response data. After the first response, each byte resets a rolling 55-second inactivity window (HTTP routing). A blocking page.pdf() can therefore produce H12 even while Chromium is healthy.

For large documents, use this flow:

  1. Accept the request and validate the input.
  2. Persist a job record and return a job ID quickly.
  3. Render in a worker process or worker dyno.
  4. Store the PDF in object storage or another durable location.
  5. Expose job status and a download URL.
  6. Expire completed files and failed jobs according to your retention policy.

If a synchronous response is unavoidable, set explicit navigation and PDF timeouts, keep the document small enough for your measured limits, and do not confuse periodic progress bytes with a fix for memory pressure. The worker pattern is more reliable because it removes the router deadline from the render operation.

7. Reduce document pressure

  • Remove unused images, fonts, scripts, and tracking resources from print-only HTML.
  • Resize source images before embedding them; CSS dimensions do not reduce decoded bitmap memory.
  • Split very long reports into independently rendered sections and combine PDFs in a separate step.
  • Wait for the specific data needed by the document instead of waiting indefinitely for every third-party request.
  • Use print CSS to disable animations, video, sticky effects, and expensive shadows.
  • Set a maximum input size and reject jobs that cannot fit your measured memory budget.
await page.addStyleTag({ content: `
  *, *::before, *::after {
    animation: none !important;
    transition: none !important;
  }
  video, iframe, .live-widget, .chat-launcher {
    display: none !important;
  }
` });

8. Troubleshooting checklist

Chrome will not start

Check which chrome, the Chrome version, buildpack order, and the complete launch error. Install the maintained buildpack, use --headless --no-sandbox, and verify that required shared libraries are present. Do not assume --disable-gpu fixes a missing library.

Sandbox failure

Heroku dynos commonly require --no-sandbox. If the error persists, confirm the deployed process is using the flags you configured rather than a local Chromium executable.

Missing browser after a Puppeteer upgrade

Check the v19+ cache location, apply the documented postbuild copy, clear the Heroku build cache, and redeploy. Ensure the cache exists inside the slug rather than only in the build environment.

R14 or R15

Set concurrency to one, log RSS and swap, close pages and browsers in finally, and repeat the same job. If one job still exceeds the quota, reduce the document or select a larger dyno. A larger dyno does not repair a leak.

H12 or H15

Measure navigation and PDF duration separately. If the render outlasts the router windows, move it to a worker queue. H12 is a request-lifecycle problem even when Chrome eventually produces a valid PDF.

Fonts or non-Latin text are missing

Puppeteer’s Heroku guidance notes that extra fonts may be needed for Chinese, Japanese, or Korean output. Install the fonts through a supported build process and verify the rendered glyphs in a smoke test.

Blank or incomplete pages

Wait for the application’s data-ready selector or a deliberate readiness signal. Use networkidle0 only when the page can become idle; analytics, websockets, and polling can prevent it. Capture console and page error events, and test third-party assets from the dyno.

9. Smoke-test the deployed dyno

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({
    headless: true,
    args: ['--headless', '--no-sandbox'],
    timeout: 30_000
  });
  try {
    const page = await browser.newPage();
    await page.setContent('

Heroku PDF smoke test

ok

', { waitUntil: 'load' }); await page.pdf({ path: '/tmp/smoke.pdf', format: 'A4' }); console.log('wrote /tmp/smoke.pdf'); await page.close(); } finally { await browser.close(); } })();

Run it on the deployed dyno before testing the full report. If this fails, fix Chrome installation or flags first. If it succeeds but the real document fails, compare memory, asset count, fonts, page count, and timing.

10. Choose where rendering should run

Option Best when Main cost or risk
Chrome in the dyno Measured documents fit memory and run in an asynchronous worker You own Chrome, fonts, caches, upgrades, cleanup, and crash recovery.
Larger Heroku dyno The application is healthy and measured peak memory is the limiting factor Higher runtime cost; it does not fix leaks or router design.
Managed browser or PDF API You need isolation, burst capacity, or less browser operations work External dependency, data-transfer concerns, and vendor cost or terms.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. It is useful when the output is a page image or PDF and you do not want to maintain Chrome on Heroku. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Every response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for the complete options. The basic request is:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());

ScreenshotNeo supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper settings and page ranges, HTML/CSS rendering, custom JavaScript and CSS, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work.

Free accounts include 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

What is the first change to make after an R14?

Set render concurrency to one and measure RSS, swap, duration, and browser process count. Then fix lifecycle cleanup before choosing a larger dyno.

Can I keep PDF generation in a web dyno?

Only when measured jobs fit the memory quota and complete inside Heroku’s request windows. Otherwise enqueue the work and run it in a worker.

Does --disable-gpu solve Chrome crashes?

It can help with a GPU initialization error, but it does not fix missing libraries, cache problems, memory leaks, or router timeouts.

Is there a universal maximum HTML size?

No authoritative source in this guide publishes one. Establish a limit from your own memory and duration measurements.