ScreenshotNeo

BlogHow-to

How to Convert HTML to PNG in Bulk

Convert many HTML files or URLs to consistent PNGs with Playwright, controlled concurrency, retries, and a hosted ScreenshotNeo option.

By the ScreenshotNeo team29 September 202611 min read

How to Convert HTML to PNG in Bulk

Direct answer: render every HTML input in a real browser, wait for the page to reach the state you need, and save a PNG screenshot for each item. Playwright’s Page API provides the navigation and screenshot operations; bulk conversion is a loop around that per-page API, with your own scheduling, filenames, retries, and logging. The same approach works for local HTML files and public URLs.

A PNG is a browser rendering at a chosen viewport, device scale, and page state. It is not a transformation of HTML source text that bypasses layout, fonts, images, CSS, or JavaScript. Decide first whether each output should show the visible viewport, the entire scrollable document, or one selected element. Playwright documents all three capture patterns and PNG output options in its screenshot guide and Page screenshot API.

Choose the capture contract before writing code

Batch jobs are predictable when every input follows the same capture contract. Record these decisions in configuration rather than scattering them through the worker code.

Decision Options When to use
Input Local file, file URL, or HTTP(S) URL Local files for reports and templates; URLs for published pages
Region Viewport, fullPage, or element Viewport for cards; full page for documents; element for a component
Viewport Fixed width and height in CSS pixels Use fixed dimensions when outputs must be comparable
Scale css or dpr css keeps one image pixel per CSS pixel; dpr follows device pixel ratio and creates larger files
Background Opaque or transparent Transparency is useful for compositing and applies to PNG, not JPEG
Readiness Load event, selector, delay, or network idle Choose a condition that represents finished content for your pages
Motion Allow, disable, or mask animations Disable motion for repeatable snapshots and visual comparisons

Set up a Playwright bulk converter

  1. Install Node.js and create a project directory.
  2. Install Playwright and its browser binaries.
  3. Put input files in an html/ directory, or prepare a text file containing URLs.
  4. Run one browser process, create a limited number of pages, and write a unique PNG path for every input.
mkdir html-to-png
cd html-to-png
npm init -y
npm install playwright
npx playwright install chromium

The browser is expensive to start, so the example below launches it once and reuses it. It also limits concurrency, creates deterministic names, logs failures, retries transient errors, and supports viewport, full-page, transparent, and element captures.

Bulk conversion is a controlled loop around browser rendering, waiting, and one output file per input.
Bulk conversion is a controlled loop around browser rendering, waiting, and one output file per input.

Complete Node.js example for local HTML files

Save this as bulk-html-to-png.js. The script accepts files from html/, or paths passed on the command line. For each file it opens a page, waits for fonts and a configurable readiness selector, disables CSS animations, and writes a PNG to png/.

const fs = require('node:fs/promises');
const path = require('node:path');
const { pathToFileURL } = require('node:url');
const { chromium } = require('playwright');

const INPUT_DIR = path.resolve('html');
const OUTPUT_DIR = path.resolve('png');
const CONCURRENCY = Number(process.env.CONCURRENCY || 4);
const RETRIES = Number(process.env.RETRIES || 2);
const VIEWPORT = {
  width: Number(process.env.WIDTH || 1440),
  height: Number(process.env.HEIGHT || 900)
};
const FULL_PAGE = process.env.FULL_PAGE !== 'false';
const READY_SELECTOR = process.env.READY_SELECTOR || '';
const TRANSPARENT = process.env.TRANSPARENT === 'true';
const ELEMENT_SELECTOR = process.env.ELEMENT_SELECTOR || '';

function safeName(file) {
  return path.basename(file, path.extname(file))
    .replace(/[^a-z0-9._-]+/gi, '-')
    .replace(/^-+|-+$/g, '') || 'page';
}

async function inputsFromArgs() {
  const args = process.argv.slice(2);
  if (args.length) return args.map(item => path.resolve(item));
  const names = await fs.readdir(INPUT_DIR);
  return names.filter(name => /\\.html?$/i.test(name))
    .sort()
    .map(name => path.join(INPUT_DIR, name));
}

async function capture(browser, file) {
  const page = await browser.newPage({
    viewport: VIEWPORT,
    deviceScaleFactor: 1,
    colorScheme: process.env.COLOR_SCHEME === 'dark' ? 'dark' : 'light'
  });

  try {
    await page.goto(pathToFileURL(file).href, {
      waitUntil: 'load',
      timeout: Number(process.env.TIMEOUT || 60000)
    });
    await page.addStyleTag({
      content: `*, *::before, *::after {
        animation: none !important;
        transition: none !important;
        caret-color: transparent !important;
      }`
    });
    await page.evaluate(() => document.fonts?.ready);
    if (READY_SELECTOR) {
      await page.waitForSelector(READY_SELECTOR, { state: 'visible', timeout: 30000 });
    }
    if (process.env.DELAY_MS) {
      await page.waitForTimeout(Number(process.env.DELAY_MS));
    }

    const output = path.join(OUTPUT_DIR, `${safeName(file)}.png`);
    const options = {
      path: output,
      type: 'png',
      fullPage: FULL_PAGE,
      animations: 'disabled',
      omitBackground: TRANSPARENT
    };
    if (ELEMENT_SELECTOR) {
      const element = page.locator(ELEMENT_SELECTOR);
      await element.screenshot(options);
    } else {
      await page.screenshot(options);
    }
    return output;
  } finally {
    await page.close();
  }
}

async function withRetry(browser, file) {
  let lastError;
  for (let attempt = 0; attempt <= RETRIES; attempt++) {
    try {
      return await capture(browser, file);
    } catch (error) {
      lastError = error;
      if (attempt < RETRIES) await new Promise(r => setTimeout(r, 1000 * (attempt + 1)));
    }
  }
  throw lastError;
}

async function main() {
  await fs.mkdir(OUTPUT_DIR, { recursive: true });
  const files = await inputsFromArgs();
  if (!files.length) throw new Error('No HTML files found');
  const browser = await chromium.launch();
  let next = 0;
  let failed = 0;

  async function worker() {
    while (true) {
      const index = next++;
      if (index >= files.length) return;
      const file = files[index];
      try {
        const output = await withRetry(browser, file);
        console.log(`OK  ${file} -> ${output}`);
      } catch (error) {
        failed++;
        console.error(`FAIL ${file}: ${error.message}`);
      }
    }
  }

  await Promise.all(Array.from(
    { length: Math.min(CONCURRENCY, files.length) }, worker
  ));
  await browser.close();
  if (failed) process.exitCode = 1;
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run a full-page batch with four concurrent pages:

node bulk-html-to-png.js

Useful variations:

# Capture only the viewport
FULL_PAGE=false node bulk-html-to-png.js

# Capture one component from every file
ELEMENT_SELECTOR='.invoice' node bulk-html-to-png.js

# Wait for a rendered component and use a dark color scheme
READY_SELECTOR='[data-rendered="true"]' COLOR_SCHEME=dark node bulk-html-to-png.js

# Use a slower, safer batch on a small machine
CONCURRENCY=1 RETRIES=3 node bulk-html-to-png.js

# Convert explicitly named files
node bulk-html-to-png.js html/one.html html/two.html

Converting a list of web URLs

For remote pages, replace the file URL with an HTTP URL and keep the same scheduling pattern. Store one URL per line in urls.txt, then use this compact worker. Do not assume that a successful navigation means the application has finished rendering; wait for a page-specific selector or a deliberate delay.

const fs = require('node:fs/promises');
const { chromium } = require('playwright');

const urls = (await fs.readFile('urls.txt', 'utf8'))
  .split(/\\r?\\n/).map(x => x.trim()).filter(Boolean);
const browser = await chromium.launch();
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });

for (const [index, url] of urls.entries()) {
  const page = await context.newPage();
  try {
    await page.goto(url, { waitUntil: 'load', timeout: 60000 });
    await page.waitForLoadState('networkidle', { timeout: 30000 }).catch(() => {});
    await page.screenshot({
      path: `png/url-${String(index + 1).padStart(4, '0')}.png`,
      type: 'png',
      fullPage: true,
      animations: 'disabled'
    });
  } finally {
    await page.close();
  }
}
await browser.close();

Full-page, viewport, and element screenshots

Viewport capture

page.screenshot({ fullPage: false }) captures the current viewport. Use this for thumbnails, above-the-fold previews, and fixed-size social images. The output is limited by the viewport dimensions you set on the browser context or page.

Full-page capture

fullPage: true captures the complete scrollable document. Long pages can create very tall and memory-heavy PNGs. Lazy-loaded images may not appear unless scrolling or page-specific JavaScript triggers them; make the page load those assets before capture and wait for an appropriate selector.

Element capture

Use page.locator('.selector').screenshot() when the output should contain a chart, invoice, product card, or other component. A missing selector is an error, so treat it as a validation failure and include the URL or filename in your logs.

Dimensions, scale, transparency, and repeatability

CSS pixels describe layout. With a CSS scale, one output pixel corresponds to one CSS pixel. A device-pixel scale follows the device pixel ratio and increases pixel dimensions and file size. Choose one deliberately for the entire batch; mixing scales makes visual comparisons and downstream processing confusing.

Set deviceScaleFactor on the context when you need high-density output. Use omitBackground: true for transparent PNGs. Transparency does not apply to JPEG, so PNG is the correct format for cutouts or compositing.

Rendering can vary with browser version, operating system fonts, font loading, current time, random data, network responses, and animations. For repeatable output:

  • Pin the Playwright and browser versions used by the worker.
  • Install the fonts your templates require, or bundle web fonts with the page.
  • Disable animations and transitions, as the example does.
  • Wait for document.fonts.ready and a page-specific ready marker.
  • Freeze dynamic data where possible and use a fixed timezone and locale.
  • Keep viewport, color scheme, scale, and browser settings consistent.

Playwright documents screenshot styling and animation controls, while its screenshot API notes that browser and host differences can still affect pixels. Treat pixel-perfect equality across different environments as a separate engineering problem.

Concurrency, retries, and batch reliability

Concurrency is a resource setting, not a universal Playwright limit. The reviewed documentation describes the per-page screenshot API and does not establish a general bulk throughput guarantee. Start with one or a few pages, observe memory and CPU, then increase gradually.

  • Unique outputs: derive names from a stable ID or index. Never let two workers write the same path.
  • Bounded concurrency: a queue prevents hundreds of pages from opening at once and exhausting memory.
  • Retries: retry timeouts and transient network errors with backoff. Do not endlessly retry invalid HTML, a missing selector, or a consistently blocked URL.
  • Resumability: skip an output only when it exists and passes a validation check such as nonzero size; otherwise recapture it.
  • Logs: record input, output, elapsed time, viewport, browser version, and error message.
  • Isolation: use a separate browser context when cookies, authentication, locale, or permissions must not leak between jobs.

Performance and cost considerations

Browser startup, page navigation, JavaScript execution, font loading, and full-page stitching all contribute to run time. Reusing one browser and limiting page concurrency usually reduces setup overhead, while full-page captures and high device scales increase memory and output size. Network idle is convenient but can delay forever on pages that keep connections open; combine a timeout with a selector or bounded delay.

For an in-house workflow, your direct costs are compute, bandwidth, storage, and engineering time. Local HTML avoids network variability but still needs browser resources. Remote URLs add DNS, TLS, third-party scripts, rate limits, and access-control concerns. Keep temporary PNGs on storage with a retention policy appropriate to the data.

Troubleshooting common failures

Symptom Likely cause Fix
Browser executable missing Playwright package installed without browser binaries Run npx playwright install chromium in the deployment image.
Blank or incomplete image Capture ran before fonts, data, or client rendering finished Wait for document.fonts.ready, a ready selector, or a bounded delay.
Lazy images absent Images load only after scrolling or intersection events Trigger the page’s loading behavior, scroll in steps, then wait for image completion.
Timeout at network idle Analytics, sockets, or polling keep the network active Prefer a selector and timeout; use network idle only where it is meaningful.
Element not found Selector changed, content is conditional, or the frame is different Verify the selector, wait for it, and handle an expected absence explicitly.
Different pixels between runs Fonts, animations, time, random data, browser, or host changed Pin the environment, disable motion, fix locale/timezone, and control data.
Out-of-memory errors Too many concurrent pages or very tall full-page PNGs Lower concurrency, capture elements or viewports, and process inputs in smaller batches.
Access denied or bot challenge The remote site requires authentication or blocks automation Use permitted credentials and headers, or obtain permission from the site owner. Do not attempt to bypass access controls.

Or skip the browser setup

ScreenshotNeo provides a hosted website screenshot API when you want one request per URL instead of maintaining browser binaries and workers. It accepts the URL and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted and removed before the capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the full parameter list in the ScreenshotNeo API documentation. This is a runnable cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

For bulk jobs, send one request per URL from your queue, or use ScreenshotNeo’s bulk capture option for up to 100 URLs per call. You can also choose full-page capture with lazy images loaded, a CSS selector for one element, dark mode, any viewport or one of 12 device presets, retina scale, custom CSS and JavaScript, clicks, waits, blocked ads or resource types, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed image links, asynchronous jobs with signed webhooks, usage reporting, and an OpenAPI specification. The parameter names used by other screenshot APIs also work, which can simplify migration.

ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots, with every feature available on every plan. Create a free ScreenshotNeo account.

Security and privacy checklist

  • Do not place API keys in HTML, client-side JavaScript, public repositories, or image URLs unless you intentionally use a signed link.
  • Decide whether pages may contain personal, financial, or confidential data before sending them to a hosted renderer.
  • For authenticated pages, use short-lived credentials and a dedicated account with the minimum required access.
  • Sanitize filenames derived from URLs or user input and keep output directories outside executable paths.
  • Respect robots, terms, access controls, and ownership of every remote page you capture.
A hosted renderer can remove common overlays before capturing the page.
A hosted renderer can remove common overlays before capturing the page.

FAQ

Can I convert HTML to PNG without a browser?

Not reliably for modern pages. HTML must be laid out with CSS and often execute JavaScript, load fonts, and fetch images. A browser renderer produces the visual result that users see.

Should I use full-page or viewport screenshots?

Use full-page for complete documents and viewport for fixed-size previews. Use an element screenshot when the deliverable is one component.

Why are my PNGs much larger than expected?

Device-pixel scale, large dimensions, transparency, and full-page height all increase pixel count. Reduce scale or capture only the required region.

How should I handle one failed item in a batch?

Log the input and error, retry transient failures with a limit, and continue the queue. Exit nonzero or produce a failure report so the batch is not mistaken for a complete run.

Can the same workflow produce PDFs?

Playwright has separate PDF functionality, while ScreenshotNeo’s API supports PDF options such as paper size, margins, landscape mode, and page ranges.