ScreenshotNeo

BlogHow-to

How to Convert HTML to PDF in an AWS Lambda Function

Render HTML as a PDF in AWS Lambda with Puppeteer and Chromium. Compare ZIP and container packaging, configure storage, and handle common rendering issues.

By the ScreenshotNeo team4 October 20268 min read

Yes. You can convert HTML to PDF in AWS Lambda by packaging a compatible headless Chromium browser with a browser automation library such as Puppeteer. The function launches the browser, loads HTML, calls the browser’s PDF-generation method, and returns the PDF bytes or saves them to durable storage such as Amazon S3. Lambda supports HTML-to-PDF file processing as a use case, and AWS has documented a Puppeteer and headless Chrome container pattern. AWS’s example is for screenshots, however; PDF rendering with it is an engineering approach to validate, not an AWS-tested conversion recipe. AWS Lambda file processing · AWS Puppeteer container example

1. Choose how to package Chromium

Chromium has native operating-system dependencies and a substantial footprint. Choose a package that matches the Lambda runtime and CPU architecture, and test the exact browser and Puppeteer versions together.

Choice Use it when Trade-offs
ZIP archive or layer Your browser distribution and dependencies fit your packaging workflow. Build for the Lambda Linux environment and architecture. ZIP deployment has its own package constraints; Lambda’s console upload threshold is 50 MB locally, with larger archives uploaded from S3.
Container image You need direct control over browser and operating-system dependencies. Build and publish an image to ECR. Lambda container images can be up to 10 GB uncompressed. AWS’s Puppeteer example uses this approach.

Lambda does not let you switch an existing function between ZIP and image package types. Create a new function if you later need the other type. For a browser-heavy deployment, an image is often the more straightforward starting point. This is a packaging recommendation based on the dependency footprint, not an AWS mandate.

2. Build a Lambda container with Puppeteer

The following Node.js example illustrates the handler and PDF creation flow. It assumes the image includes a Lambda-compatible Chromium executable and compatible Puppeteer installation. The AWS example establishes the browser-in-container pattern; select and verify a current browser distribution for your runtime and architecture before deployment. Do not copy an old blog image tag without checking current Lambda runtime support.

Create package.json:

{
  "name": "lambda-html-to-pdf",
  "version": "1.0.0",
  "type": "module",
  "dependencies": {
    "puppeteer-core": "<pin-a-compatible-version>"
  }
}

Replace the placeholder with a pinned version verified alongside your chosen Chromium package. The Chromium package often provides its own launcher path and launch arguments; use those rather than assuming a system-wide chromium binary. The illustrative handler below expects CHROMIUM_EXECUTABLE_PATH to be set by the image.

// index.mjs
import puppeteer from 'puppeteer-core';

export const handler = async (event) => {
  const html = event?.html;
  if (typeof html !== 'string' || html.length === 0) {
    return { statusCode: 400, body: 'Provide a non-empty html string.' };
  }

  let browser;
  try {
    browser = await puppeteer.launch({
      executablePath: process.env.CHROMIUM_EXECUTABLE_PATH,
      headless: true,
      args: ['--no-sandbox', '--disable-setuid-sandbox']
    });
    const page = await browser.newPage();
    await page.setContent(html, { waitUntil: 'networkidle0', timeout: 30000 });
    await page.emulateMediaType('print');
    const pdf = await page.pdf({
      format: 'A4',
      printBackground: true,
      preferCSSPageSize: true
    });

    return {
      statusCode: 200,
      headers: { 'content-type': 'application/pdf' },
      isBase64Encoded: true,
      body: Buffer.from(pdf).toString('base64')
    };
  } catch (error) {
    console.error('HTML-to-PDF conversion failed', error);
    return { statusCode: 500, body: 'PDF conversion failed.' };
  } finally {
    if (browser) await browser.close();
  }
};

This handler is an adaptable implementation sketch, not a tested AWS sample. For an HTTP API that expects a binary response, configure binary media handling and response encoding appropriately. For larger PDFs, prefer saving the output to S3 and returning a reference or signed URL rather than sending the bytes through the invocation response.

Create a Dockerfile using a currently supported Lambda Node.js base image. Install the pinned application dependencies and a Chromium build compatible with that base image, set the executable path, copy the handler, and use the Lambda runtime entry point. AWS base images include the runtime interface client; if you use a different base image, include a Lambda runtime interface client as described in AWS’s container image documentation. The exact Chromium installation commands depend on the chosen distribution, so there is no safe universal install command to copy across architectures and base images.

Build for the same architecture configured on the function. Lambda deployment binaries must match that architecture. Verify the image starts in Lambda and that Chromium can launch before relying on PDF output.

3. Return a PDF or store it in S3

For small synchronous conversions, returning PDF bytes can be convenient, but invocation and API response size limits may make this unsuitable for larger documents. For durable or larger outputs, write the generated bytes to S3 and return the object key or a time-limited signed URL. AWS’s file-processing guidance uses /tmp for temporary work and S3 for durable handling.

If writing through a temporary file, use /tmp, then upload the file to S3 and remove it when done. Lambda’s filesystem is otherwise read-only. Writable ephemeral storage is configurable from 512 MB to 10,240 MB in 1 MB increments. Account for Chromium temporary files, downloaded assets, and the generated PDF together when choosing a size.

4. Make HTML and print output predictable

  • Set print CSS: use @page for page size and margins, and break-before, break-after, or break-inside to control pagination.
  • Include backgrounds: printBackground: true includes CSS backgrounds that Chromium may otherwise omit.
  • Wait for required content: network idle is useful for static pages, but pages with polling or persistent connections may never become idle. Prefer a specific readiness selector or application signal for dynamic content.
  • Load fonts and images deliberately: package fonts where possible, await image decoding when necessary, and confirm remote assets are reachable from the Lambda network.
  • Choose a source-loading method: page.setContent is suitable for supplied markup. To render a URL, use page.goto and an explicit wait condition. Relative asset URLs need a base URL or absolute paths.
  • Sanitize untrusted input: arbitrary HTML can execute JavaScript and make network requests. Restrict outbound access, validate input, and avoid passing secrets into pages that users control.

5. Configure resources and reliability

There is no sourced universal memory, timeout, or conversion-speed setting for Chromium PDF work. Measure representative documents instead of borrowing the 256 MB and 15-second values from AWS’s PDF-encryption sample; that sample does not render HTML in a browser.

  • Memory: benchmark your own mix of page count, image size, fonts, and JavaScript. More memory also changes the CPU available to Lambda, which can affect browser startup and rendering time.
  • Timeout: allow for cold start, browser launch, asset fetching, rendering, and storage upload. Bound navigation and readiness waits so a slow external asset cannot consume the full invocation.
  • Temporary storage: raise ephemeral storage when browser files, assets, and output approach the configured capacity.
  • Concurrency: each concurrent invocation may launch a browser and consume memory and temporary space. Set concurrency in line with downstream capacity and cost limits.
  • Browser lifecycle: close the browser in a finally block. Reusing a browser across warm invocations may reduce startup work, but requires careful page cleanup and recovery after crashes; validate this pattern before adopting it.
  • Retries: retry transient failures selectively. A deterministic malformed document or incompatible binary will not be fixed by retrying. Use idempotent output keys if jobs can be invoked again.

6. Troubleshooting

Symptom Likely cause Fix
Executable not found The browser path differs from the packaged distribution or the binary was not copied into the image. Inspect the image contents, set the actual executable path, and verify it in the Lambda environment.
Browser fails to launch Missing native libraries, incompatible architecture, unsupported runtime, or incorrect launch arguments. Rebuild for the configured architecture and Lambda-compatible Linux environment; pin compatible browser and automation versions.
Works locally, fails in Lambda Local OS dependencies or CPU architecture differ from the deployment environment. Build and validate using the same base image and architecture as the function.
PDF is blank or missing images Rendering began before content or assets loaded, relative paths lack a base URL, or outbound requests are blocked. Wait for a meaningful selector or image readiness, use absolute asset URLs or a base URL, and check network access.
Fonts or layout differ Font files are unavailable or print media rules change the page. Package required fonts, wait for font readiness, inspect print CSS, and set page size and margins explicitly.
Invocation times out Slow assets, unbounded page readiness, a large document, or insufficient configured time. Set navigation and rendering timeouts, reduce unnecessary remote assets, measure the workload, and tune Lambda timeout and memory.
Disk full or write error Code writes outside writable storage or ephemeral storage is too small. Use /tmp and increase configured ephemeral storage to cover peak temporary and output files.
Response is truncated or rejected The PDF is too large for the chosen invocation or API response path. Store the PDF in S3 and return a reference instead of inline bytes.

7. Cost notes

Lambda cost depends on invocation count, configured memory, execution duration, and associated services such as S3 and ECR. Browser startup, external asset loading, and page complexity affect duration; the available sources provide no benchmark for conversion speed or cost. Measure your templates and traffic, and use S3 for outputs that should persist. Keep deployment image size and browser updates in your operational plan.

Or skip the browser setup

If the source is a live web page and you need a screenshot or PDF capture workflow without packaging Chromium, ScreenshotNeo is a website screenshot API and MCP server. Its PDF endpoint accepts a URL in one request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf

See the ScreenshotNeo API documentation for PDF options and authentication. ScreenshotNeo removes cookie banners, popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, no card required.

FAQ

Does Lambda itself convert HTML to PDF?

Lambda runs your code; you provide the rendering engine and its dependencies. A headless browser such as Chromium performs the HTML rendering and PDF generation.

Does AWS’s Puppeteer example create PDFs?

The cited AWS example demonstrates Puppeteer and headless Chrome for screenshots in a container image. PDF creation is an implementation extrapolation, so validate the browser package and output in your environment.

Can I convert a remote webpage instead of an HTML string?

Yes. Navigate Chromium to the page URL, wait for the required content, and ensure the Lambda function can reach the page and its assets. Consider access controls and SSRF risks when accepting URLs from users.

Do I need S3?

No, not for every small synchronous response. S3 is useful when the file needs durable storage or is too large to return conveniently from the invocation.