ScreenshotNeo

BlogHow-to

How to Make Puppeteer setContent Load Static File Requests

Learn why setContent() does not load sibling files, and how to serve assets, use absolute URLs, intercept requests, and diagnose failures.

By the ScreenshotNeo team1 October 20268 min read

Short answer: page.setContent() sets the supplied HTML markup. It is not a disk-file loader and does not automatically make sibling CSS, JavaScript, images or fonts available. For an existing static directory, serve that directory over HTTP and use page.goto(). For generated markup, use absolute asset URLs or add styles and scripts directly. Use request interception only when you need to rewrite, fulfill or block requests.

This guide shows complete Puppeteer examples, explains readiness and URL resolution, covers file:// edge cases, and includes fixes for the failures developers usually see.

1. What setContent() actually does

The Puppeteer page.setContent() API accepts an HTML string and places that markup in the page. Its documented lifecycle option defaults to load. The method does not document a directory argument, a disk-file loader, or a static asset base path.

Given this HTML:

await page.setContent(`
  <!doctype html>
  <html>
    <head>
      <link rel="stylesheet" href="styles.css">
    </head>
    <body>
      <img src="images/logo.png" alt="Logo">
      <script src="app.js"></script>
    </body>
  </html>
`);

the browser must resolve styles.css, images/logo.png, and app.js against a meaningful document URL. A string passed to setContent() does not turn your project folder into an HTTP site.

2. Choose the approach that matches your input

Situation Recommended method Why
An existing folder contains HTML and sibling assets Serve the folder and call page.goto() HTTP gives relative URLs a predictable directory base.
HTML is generated as a string Keep setContent(); use absolute URLs or inline/add resources No local directory needs to be discovered.
You must rewrite, block or synthesize requests Enable request interception Handlers can continue, abort or respond to each request.
A workflow specifically requires local files Use file:// only after validating your exact Chromium build File origins can be opaque and linked-resource behavior varies.

3. Serve a static directory and navigate to it

This is the reliable choice for a static site. The server’s document root becomes the asset base, so links such as css/site.css and images/hero.webp resolve normally.

Minimal Node.js server

import http from 'node:http';
import fs from 'node:fs/promises';
import path from 'node:path';
import { fileURLToPath } from 'node:url';
import puppeteer from 'puppeteer';

const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), 'public');
const mime = {
  '.html': 'text/html; charset=utf-8',
  '.css': 'text/css; charset=utf-8',
  '.js': 'text/javascript; charset=utf-8',
  '.json': 'application/json; charset=utf-8',
  '.png': 'image/png',
  '.jpg': 'image/jpeg',
  '.jpeg': 'image/jpeg',
  '.webp': 'image/webp',
  '.svg': 'image/svg+xml',
  '.woff2': 'font/woff2'
};

const server = http.createServer(async (req, res) => {
  try {
    const requestPath = new URL(req.url, 'http://127.0.0.1').pathname;
    const relative = decodeURIComponent(requestPath === '/' ? '/index.html' : requestPath);
    const filePath = path.resolve(root, `.${relative}`);
    if (!filePath.startsWith(`${root}${path.sep}`) && filePath !== root) {
      res.writeHead(403).end('Forbidden');
      return;
    }
    const data = await fs.readFile(filePath);
    res.writeHead(200, {
      'Content-Type': mime[path.extname(filePath)] ?? 'application/octet-stream',
      'Cache-Control': 'no-store'
    });
    res.end(data);
  } catch {
    res.writeHead(404).end('Not found');
  }
});

await new Promise(resolve => server.listen(0, '127.0.0.1', resolve));
const { port } = server.address();

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto(`http://127.0.0.1:${port}/index.html`, { waitUntil: 'load' });
await page.screenshot({ path: 'static-page.png', fullPage: true });
await browser.close();
server.close();

For production or CI, you can replace the small server with any static HTTP server. Keep the server alive until the page has loaded every resource needed by the capture.

Important details

  • Include the URL scheme: http://127.0.0.1:PORT/index.html, not just a path.
  • Make sure the server exposes the exact URL requested by each src, href, and font declaration.
  • Use a loopback bind address when the browser and server run in the same job.
  • Prevent directory traversal and return the correct content type.

4. Keep setContent() for generated markup

If the HTML is assembled in memory, use absolute URLs for external assets:

await page.setContent(`
  <!doctype html>
  <html>
    <head>
      <link rel="stylesheet" href="https://example.test/styles.css">
    </head>
    <body>
      <img src="https://example.test/image.png" alt="">
    </body>
  </html>
`, { waitUntil: 'load' });

For local or generated resources, inline them or use page.addStyleTag() and page.addScriptTag(), which accept a URL or content.

await page.setContent('<main id="app">Hello</main>');
await page.addStyleTag({ content: '.ready { color: darkgreen; }' });
await page.addScriptTag({ content: 'document.querySelector("#app").classList.add("ready");' });

If you need relative URLs while using setContent(), add a base element that points to an HTTP directory you control:

await page.setContent(`
  <base href="http://127.0.0.1:4173/">
  <link rel="stylesheet" href="css/site.css">
  <img src="images/photo.png" alt="">
`);

The base URL must be reachable by Chromium. A base element does not make files available by itself.

5. Intercept requests only when you need custom delivery

Request interception lets you modify, fulfill or block traffic. Once enabled, every request pauses until your handler calls continue(), respond(), abort(), or the request completes from cache.

await page.setRequestInterception(true);
page.on('request', async request => {
  try {
    const url = new URL(request.url());
    if (url.pathname === '/config.json') {
      await request.respond({
        status: 200,
        contentType: 'application/json',
        body: JSON.stringify({ mode: 'capture' })
      });
      return;
    }
    if (url.hostname === 'ads.example.test') {
      await request.abort('blockedbyclient');
      return;
    }
    await request.continue();
  } catch {
    if (!request.isInterceptResolutionHandled()) await request.abort();
  }
});

await page.setContent(`
  <script>
    fetch('/config.json').then(r => r.json()).then(console.log);
  </script>
`);

Keep the handler small and deterministic. Guard against a request being resolved by another listener, and ensure every branch resolves the request. One forgotten request can make navigation appear to hang.

6. Wait for the state you actually need

The load event means the document’s load lifecycle fired. It does not prove that a client-side application finished rendering after load. Puppeteer’s setContent() wait options support documented lifecycle values; choose a condition that represents the output you will use.

await page.setContent(html, { waitUntil: 'load' });
await page.waitForSelector('[data-rendered="true"]');
await page.screenshot({ path: 'ready.png' });

Other useful patterns include:

// Wait for a specific response.
const responsePromise = page.waitForResponse(response =>
  response.url().endsWith('/api/data.json') && response.ok()
);
await page.goto('http://127.0.0.1:4173/index.html', { waitUntil: 'domcontentloaded' });
await responsePromise;

// Wait for an application condition.
await page.waitForFunction(() => window.app?.status === 'ready');

Prefer a selector, response, or application state over an arbitrary delay when correctness matters. Use a delay only when the page has no observable readiness signal.

7. Troubleshooting common failures

Symptom Likely cause Fix
CSS or images return 404 Relative URLs have no useful HTTP base, or the server path is wrong Use page.goto() to a served directory, add a valid <base>, and inspect the final request URL.
Page hangs after enabling interception A request was never resolved Call continue(), respond(), or abort() on every branch.
Screenshot is taken before content appears load fired before asynchronous rendering completed Wait for the rendered selector, response, or application state.
Fonts are missing Font URL is not exposed, blocked, or used before it finishes Verify the font request, serve the file, and wait for document.fonts.ready.
Scripts run but relative fetches fail The script’s requests resolve against an unexpected document URL Use absolute endpoints or a reachable base URL; inspect page.url() and request events.
file:// works locally but fails in CI File-scheme origins and restrictions vary by browser/build Serve the directory over loopback HTTP and test with the deployed Chromium version.
HTTPS assets fail from a local page Certificate, network, or mixed-content policy issue Inspect console and response errors; use valid certificates or a controlled local endpoint.

Log the URLs that matter

page.on('requestfailed', request => {
  console.error('REQUEST FAILED', request.url(), request.failure());
});
page.on('response', response => {
  if (!response.ok()) console.error('HTTP', response.status(), response.url());
});
page.on('console', message => console.log('BROWSER', message.type(), message.text()));

When debugging, compare the URL in the HTML with the URL in the request event. That quickly separates a bad base path from a missing file, a browser policy issue, or an application error.

8. Performance, reliability and cost considerations

  • Reuse the browser: launch one browser process and create pages per job. Repeated launches add startup work.
  • Keep the local server close: loopback serving avoids external DNS and network variability.
  • Wait narrowly: a specific readiness signal prevents unnecessary idle time while avoiding incomplete captures.
  • Cache stable assets: browser caching can reduce repeated transfers, but invalidate it when generated files change.
  • Control interception: interception adds handler work to every request. Enable it only for routes you must change.
  • Set explicit timeouts: fail a job with a useful error instead of waiting indefinitely for an unavailable asset.
  • Make jobs repeatable: pin the Puppeteer/Chromium version, use deterministic test data, and record failed request URLs.
  • Resource costs: self-hosting consumes CPU, memory, browser processes and bandwidth. Measure your own workload; the API documentation does not provide a universal screenshot benchmark.

9. Or skip the browser setup

If your goal is simply a clean screenshot or PDF of a URL, ScreenshotNeo provides a website screenshot API and MCP server. It handles the browser setup and accepts one GET request.

See the ScreenshotNeo API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups and chat widgets are removed before the shot.
  • Bot checks, blank pages, failed loads and cache hits are never billed; response headers identify the page verdict and billing status.
  • An MCP server lets Claude, Cursor and other MCP clients take screenshots with take_screenshot, inspect pages with get_page_info, and create PDFs with capture_pdf.
  • There are 1,000 free screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account and start with 1,000 screenshots a month at no charge.

10. FAQ

Can setContent() load an HTML file path?

No. Read the file yourself and pass its contents, or serve its directory and navigate to an HTTP URL.

Should I use networkidle0 with setContent()?

The documented setContent() wait options do not use networkidle0 or networkidle2. Wait for the selector, response, or state your capture requires.

Is a base element enough to load local assets?

Only if its URL points to a server that exposes those assets. It changes URL resolution; it does not publish files.

When is request interception justified?

Use it for deliberate rewriting, blocking, or synthetic responses. For ordinary static files, serving the directory is simpler and less error-prone.

Why is HTTP usually preferable to file://?

Browsers can treat file documents as opaque origins, and linked-file behavior can vary. Loopback HTTP gives predictable URL and origin behavior.