How to Run Arbitrary HTML5 Securely with Puppeteer
Run untrusted HTML5 with Puppeteer using Chrome sandboxing, container isolation, network controls, timeouts, and safer production defaults.
Run arbitrary HTML5 as if it were hostile code. Keep Chrome’s sandbox enabled, execute the browser inside a disposable container or OS-level boundary, restrict outbound network access, pass no credentials, and enforce resource and wall-clock limits. Puppeteer’s process separation and request filtering add defense in depth, but neither replaces Chrome’s sandbox nor host isolation.
Puppeteer’s own documentation says Chrome uses multiple sandbox layers and that running without a sandbox is strongly discouraged. If you see No usable sandbox!, fix the host or move the job to an isolated runtime; do not add --no-sandbox for arbitrary content. Read Puppeteer’s sandbox troubleshooting guidance.
1. Security model for arbitrary HTML5
What can untrusted HTML do?
HTML5 can execute JavaScript, load remote resources, open connections, consume CPU and memory, trigger browser bugs, and attempt requests to internal services. A page may also contain infinite loops, huge canvases, large downloads, WebAssembly, redirects, popups, or code designed to keep the browser busy.
Use several boundaries
- Chrome sandbox: keep the sandbox enabled and run a current Chrome build.
- OS or container isolation: run each job in a disposable worker with a minimal filesystem and no sensitive mounts.
- Network policy: deny internal address ranges, cloud metadata endpoints, and destinations that the job does not need. Enforce this outside Puppeteer.
- Process limits: apply CPU, memory, process-count, disk, and wall-clock limits.
- Credential separation: do not reuse an authenticated profile or expose secrets in environment variables, files, cookies, or headers.
Site Isolation gives Chromium another defense-in-depth layer by placing different sites in separate sandboxed processes and limiting the sensitive data each process receives. It is not a substitute for the host boundary. See the Chromium Site Isolation documentation.
2. A safer Puppeteer implementation
The following Node.js program accepts HTML from standard input, renders it in a new browser context, writes a PNG, and applies a small application-level request allowlist. The allowlist is an additional guardrail; your container or host firewall must remain the enforcement layer.
import puppeteer from 'puppeteer';
import fs from 'node:fs/promises';
const html = await fs.readFile(0, 'utf8');
const output = process.env.OUTPUT ?? 'render.png';
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 15000);
const maxHtmlBytes = 2 * 1024 * 1024;
if (Buffer.byteLength(html, 'utf8') > maxHtmlBytes) {
throw new Error('HTML input exceeds the configured size limit');
}
// Permit only destinations explicitly required by this render.
// Keep this list short and configure network egress outside the process too.
const allowedHosts = new Set([
'example.com',
'www.example.com'
]);
const browser = await puppeteer.launch({
headless: true,
// Do not add --no-sandbox for untrusted HTML.
timeout: 30000,
args: [
'--disable-dev-shm-usage',
'--disable-features=IsolateOrigins,site-per-process'
]
});
try {
const page = await browser.newPage();
await page.setViewport({ width: 1280, height: 900, deviceScaleFactor: 1 });
await page.setJavaScriptEnabled(true);
await page.setDefaultNavigationTimeout(timeoutMs);
await page.setDefaultTimeout(timeoutMs);
await page.setRequestInterception(true);
page.on('request', request => {
const type = request.resourceType();
const url = request.url();
// Data and blob URLs are local to the document. Restrict other traffic.
if (url.startsWith('data:') || url.startsWith('blob:')) {
return request.continue();
}
let parsed;
try {
parsed = new URL(url);
} catch {
return request.abort('blockedbyclient');
}
const safeScheme = parsed.protocol === 'https:';
const safeHost = allowedHosts.has(parsed.hostname);
const blockedType = new Set(['media', 'font']).has(type);
if (!safeScheme || !safeHost || blockedType) {
return request.abort('blockedbyclient');
}
return request.continue();
});
await page.setContent(html, {
waitUntil: 'networkidle0',
timeout: timeoutMs
});
await page.screenshot({
path: output,
type: 'png',
fullPage: true,
captureBeyondViewport: true
});
} finally {
await browser.close();
}
console.log(`Wrote ${output}`);
Install and run it with:
npm install puppeteer
printf '%s' '<!doctype html><h1>Hello</h1>' | node render.mjs
# OUTPUT=page.png TIMEOUT_MS=10000 node render.mjs
The example includes --disable-dev-shm-usage, which can help in small containers whose shared-memory mount is limited. It does not disable the Chrome sandbox. The two Site Isolation related flags are intentionally omitted from production hardening: do not disable Site Isolation for untrusted content. Remove those flags if you copy the example unchanged.
3. Configure the browser and host correctly
Keep Chrome and Puppeteer compatible
Puppeteer releases are tied to browser revisions. Keep the package and its managed browser current together, and validate compatibility when updating. The Puppeteer FAQ explains this relationship.
Do not “fix” sandbox errors with a flag
A No usable sandbox! launch error means the runtime cannot provide a usable Chrome sandbox. Common causes include missing user namespaces, restrictive kernel settings, or container security profiles. Correct the host configuration, use a supported container runtime, or move rendering to a worker with the required sandbox capabilities. Puppeteer documents --no-sandbox only for content that is absolutely trusted.
Choose headless mode for behavior, not security
Regular headless Chrome is Puppeteer’s default. The separate chrome-headless-shell binary can be more performant for some automation workloads but does not match regular Chrome completely. Neither mode is a security boundary. Choose based on required HTML5 behavior, compatibility, and operational support. See Puppeteer’s headless modes guide.
4. Isolation and network controls
Container or OS boundary
- Use a short-lived worker or recycle the browser after each job or small batch.
- Run as a non-root user where the runtime supports Chrome’s sandbox requirements.
- Mount only a writable temporary directory; do not mount source trees, SSH keys, package-manager credentials, or host sockets.
- Drop unnecessary Linux capabilities and use a restrictive seccomp/AppArmor/SELinux profile appropriate for your image.
- Set memory, CPU, PID, temporary-disk, and maximum-output limits.
Network egress
Start with no outbound access, then allow only the origins required by the rendering job. Block loopback, private IPv4 and IPv6 ranges, link-local addresses, Unix sockets, internal DNS, and cloud metadata endpoints. Resolve and validate destinations through your network layer as well as in application code; DNS rebinding and redirects can bypass a simple hostname check.
Puppeteer’s experimental Chrome URL allowlist can add browser-level filtering in supported Chrome versions, but its API documentation explicitly calls it “not a complete network sandbox.” Keep container or OS-level policy as the enforcement layer. See the Puppeteer ConnectOptions reference.
5. Limits that prevent denial of service
| Control | Why it matters | Typical implementation |
|---|---|---|
| Input size | Prevents oversized HTML and embedded data | Reject input above a documented byte limit |
| Navigation timeout | Stops pages waiting forever | setDefaultNavigationTimeout and an outer job deadline |
| Memory and CPU | Limits huge canvases, WebAssembly, and busy loops | Container cgroups or an equivalent OS quota |
| Process count | Limits fork and worker abuse | Container PID limit |
| Output size | Prevents enormous screenshots or PDFs | Check file size and image dimensions before returning |
| Browser lifetime | Reduces persistence between jobs | Close the context and recycle the browser regularly |
Apply an outer supervisor timeout as well as Puppeteer’s page timeout. If the deadline expires, terminate the worker process; a JavaScript timeout alone may leave a stuck Chromium process alive.
6. Handling HTML5 features safely
- JavaScript: enable it only when the page requires it. Treat every script as untrusted code.
- WebAssembly: account for CPU and memory usage; disable it at the network or browser policy layer if unnecessary.
- Canvas and WebGL: cap viewport and output dimensions. GPU exposure adds complexity; use a runtime policy suited to your workload.
- Workers and service workers: clear the context after each job and prevent persistent profiles.
- Downloads and popups: reject unexpected requests and avoid granting download directories or extensions.
- Storage: use a fresh incognito context; never expose a logged-in profile, bearer token, or production cookie jar.
7. Troubleshooting
No usable sandbox!
Cause: the host or container does not expose the sandbox mechanisms Chrome needs.
Fix: correct namespace or security-profile configuration, use a supported image/runtime, or move the job to an isolated worker. Do not add --no-sandbox for arbitrary HTML.
Pages hang until the timeout
Cause: a never-ending script, blocked resource, long network request, service worker, or page that never reaches the chosen lifecycle event.
Fix: enforce an outer deadline, reduce the allowed network, use a smaller resource budget, and choose a lifecycle condition that matches your capture. Log the last requested URL and resource type.
Images or fonts are missing
Cause: request interception or the network policy blocked the origin, MIME type, redirect, or font resource.
Fix: inspect aborted requests, explicitly allow the required HTTPS origins, and verify that redirects remain within the allowlist.
Chrome crashes in a container
Cause: insufficient memory, a small /dev/shm, incompatible libraries, or a process limit.
Fix: increase the worker’s memory and shared-memory allocation where appropriate, or use --disable-dev-shm-usage as an operational workaround. Keep the sandbox enabled and inspect Chrome’s stderr output.
Unexpected access to internal services
Cause: application-level hostname checks missed redirects, alternate IP representations, DNS rebinding, or non-HTTP protocols.
Fix: enforce egress at the container or host network layer, deny private and metadata ranges, and validate every connection path outside Puppeteer.
8. Performance and reliability
- Reuse a browser only for a small, controlled batch; create a fresh incognito context for every job.
- Recycle the browser after a bounded number of jobs to limit leaks and accumulated state.
- Use regular headless Chrome when compatibility matters; evaluate
chrome-headless-shellonly when its behavior fits your HTML5 workload. - Keep screenshots bounded by viewport and pixel dimensions. Full-page captures can consume substantially more memory than viewport captures.
- Record browser version, Puppeteer version, job duration, blocked requests, exit reason, and output size.
- Retry only infrastructure failures. Do not blindly retry deterministic script errors or blocked destinations.
Security controls can add latency through isolation startup, network policy checks, and browser recycling. Measure these costs with your own HTML and container limits; the source material does not provide a universal benchmark.
9. Cost and operational trade-offs
Self-hosting gives control over browser versions, network policy, and data location, but you operate Chrome updates, sandbox-compatible hosts, isolation, monitoring, capacity, and incident response. Browser-level request filtering is easier to deploy but weaker than an OS or container boundary. Treat the latter as mandatory for arbitrary content.
10. Or skip the browser setup
ScreenshotNeo provides a website screenshot API when you need an image or PDF without operating Puppeteer workers. Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options. A basic request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots, and every feature is available on every plan. Create a free ScreenshotNeo account.
11. FAQ
Is Puppeteer itself a sandbox?
No. Puppeteer is an automation client that operates off-process from the browser. It does not replace Chrome’s sandbox or OS/container isolation. See the official FAQ.
Can I safely render arbitrary HTML with --no-sandbox?
No. For arbitrary or untrusted HTML, fix the runtime or use an isolated worker with Chrome’s sandbox enabled.
Does headless mode make malicious HTML safe?
No. Headless is a browser mode. Security comes from Chrome sandboxing, host isolation, network policy, credential separation, and resource limits.
Should I allow all network requests and rely on Puppeteer?
No. Puppeteer’s URL filtering is an extra guardrail, not a complete network sandbox. Enforce the destination policy at the container or host network layer.
When should I recycle the browser?
Recycle it after each job when isolation is the priority, or after a small bounded batch when startup cost matters. Always create a fresh context and never reuse secrets-bearing state.


