ScreenshotNeo

BlogHow-to

How to Generate Website Screenshots in an AWS Lambda Function

Package a headless browser with Lambda, capture a page, and save the image to S3. Includes a Puppeteer example, deployment options, and troubleshooting.

By the ScreenshotNeo team4 October 202610 min read

To generate a website screenshot in AWS Lambda, package a headless browser with your function, launch it to open the target URL, write the screenshot to Lambda’s writable /tmp directory, then upload it to durable storage such as Amazon S3. AWS’s example uses Puppeteer and headless Chrome in a Lambda container image; it also demonstrates a separate fan-out function for asynchronously capturing a list of URLs. Read the AWS example.

This guide uses Node.js and Puppeteer. The code shows the capture and S3 upload flow; browser packaging and launch settings must match the exact browser build, Lambda runtime, and architecture you deploy. Verify those versions together before shipping.

1. Choose a Lambda package format

Lambda supports both ZIP deployment packages and container images. A ZIP can work if your browser binary and its dependencies fit and are packaged correctly. A container image gives you control over the browser dependencies and build process, which is why AWS uses one for its Puppeteer example. AWS documents both packaging options.

Choice Use it when Plan for
ZIP package Your browser build and all required files fit your deployment approach. Packaging the browser and native libraries, and validating them in Lambda’s runtime environment.
Container image You need more control over operating-system packages and browser dependencies. Building and publishing an image, keeping it within Lambda’s limits, and ensuring it works with Lambda’s runtime interface.

Lambda container images can be up to 10 GB uncompressed. Their filesystem is read-only except for /tmp. Lambda’s configurable temporary storage ranges from 512 MB to 10,240 MB in 1 MB increments. Browser extraction, caches, and temporary output need to fit in writable storage. See Lambda container image requirements and configure ephemeral storage.

2. Create the screenshot function

The handler below accepts a URL, opens it with Puppeteer, captures a PNG, and uploads the bytes to S3. It uses the AWS SDK for JavaScript v3. The example expects Puppeteer to be installed and a compatible browser to be included in the image. If your selected browser package supplies its own executable path or launch arguments, adapt the launch configuration to that package’s current documentation.

Project files

package.json
{
  "name": "lambda-website-screenshot",
  "version": "1.0.0",
  "type": "module",
  "dependencies": {
    "@aws-sdk/client-s3": "^3.0.0",
    "puppeteer": "^24.0.0"
  }
}

Pin package versions in your lockfile for repeatable builds. This package declaration alone does not guarantee that the installed browser and native libraries are compatible with your Lambda image; validate the complete image as described below.

// index.mjs
import puppeteer from 'puppeteer';
import { S3Client, PutObjectCommand } from '@aws-sdk/client-s3';
import { randomUUID } from 'node:crypto';

const s3 = new S3Client({});
const bucket = process.env.SCREENSHOT_BUCKET;

export const handler = async (event) => {
  if (!bucket) throw new Error('SCREENSHOT_BUCKET is required');

  // For a public endpoint, add strict URL validation and destination controls.
  const url = event?.url;
  if (typeof url !== 'string' || !/^https?:\/\//i.test(url)) {
    return { statusCode: 400, body: JSON.stringify({ error: 'Provide an http(s) URL' }) };
  }

  let browser;
  try {
    browser = await puppeteer.launch({
      headless: true,
      args: ['--no-sandbox', '--disable-setuid-sandbox'],
      // Set executablePath here if your packaged browser requires it.
    });
    const page = await browser.newPage({ viewport: { width: 1365, height: 768 } });
    await page.goto(url, { waitUntil: 'networkidle2', timeout: 45000 });
    const image = await page.screenshot({ type: 'png', fullPage: true });

    const key = `screenshots/${randomUUID()}.png`;
    await s3.send(new PutObjectCommand({
      Bucket: bucket,
      Key: key,
      Body: image,
      ContentType: 'image/png'
    }));
    return { statusCode: 200, body: JSON.stringify({ bucket, key }) };
  } finally {
    if (browser) await browser.close();
  }
};

networkidle2 is one navigation policy, not a guarantee that every page has finished rendering: pages with persistent network traffic may never become idle, and delayed content may appear afterward. For those pages, use a bounded delay or wait for a selector that identifies the content you need. Do not wait indefinitely.

Container image outline

Use an AWS Lambda base image for your chosen supported Node.js runtime, then install Puppeteer and the browser dependencies required by the selected build. The Dockerfile below shows the Lambda image structure; the browser installation step is deliberately package-specific. Follow the browser package’s current installation instructions and test the resulting image against Lambda’s read-only filesystem and default user.

# Dockerfile — complete the browser installation for your pinned build
FROM public.ecr.aws/lambda/nodejs:22

WORKDIR ${LAMBDA_TASK_ROOT}
COPY package*.json ./
RUN npm ci
COPY index.mjs ./

# Install/package a Chromium build and native libraries compatible with
# this base image, Node runtime, and Lambda architecture.
# Configure Puppeteer executablePath/args if the chosen package requires it.

CMD ["index.handler"]

AWS requires container images to implement the Lambda Runtime API; AWS base images include the Lambda runtime components. The container must run with a read-only filesystem apart from /tmp, and the default Lambda user must be able to read the files it needs. Do not assume an arbitrary local Chrome installation will run unchanged in Lambda. Review AWS’s image requirements.

3. Configure and deploy

  1. Create an S3 bucket for the output and set SCREENSHOT_BUCKET to its name in the Lambda function configuration.
  2. Give the function execution role permission to write objects to that bucket. Restrict the permission to the bucket and key prefix the function uses.
  3. Build the image for the Lambda architecture you configure, publish it to Amazon ECR, and create or update a Lambda function from that image.
  4. Set memory, timeout, and ephemeral storage based on representative pages, image dimensions, and the browser package. Keep enough timeout headroom for variable page loads.
  5. Invoke the function with an event such as {"url":"https://example.com"}. Check the returned S3 key and inspect the image.

For a container-based function, a SAM configuration can set key runtime parameters. This is a configuration fragment; set the image build metadata and role for your project before deploying.

Resources:
  ScreenshotFunction:
    Type: AWS::Serverless::Function
    Properties:
      PackageType: Image
      Timeout: 90
      MemorySize: 2048
      EphemeralStorage:
        Size: 1024
      Environment:
        Variables:
          SCREENSHOT_BUCKET: your-output-bucket
      Policies:
        - S3WritePolicy:
            BucketName: your-output-bucket

The memory, timeout, and ephemeral-storage values above are starting configuration examples, not universal recommendations. Tune them using your own pages and browser build. Lambda exposes settings for memory, timeout, temporary storage, networking, and filesystems in its function configuration.

4. Select capture behavior deliberately

Need Puppeteer setting or approach Trade-off
Viewport screenshot fullPage: false (default) Captures the current viewport dimensions.
Whole document fullPage: true Can produce a tall, larger image and use more memory and temporary space.
Wait for page readiness page.goto with an appropriate waitUntil Network-idle conditions may not suit pages with long polling, ads, or streaming.
Wait for app content await page.waitForSelector('.content') Use a selector meaningful to the page; handle timeout if it never appears.
Wait a fixed interval await new Promise(r => setTimeout(r, 1500)) Simple but can be wasteful or too short; keep the delay bounded.
Different viewport page.setViewport({ width, height, deviceScaleFactor }) Viewport affects responsive layout; device scale affects pixel dimensions.
JPEG or WebP output page.screenshot({ type: 'jpeg', quality: 80 }) or type: 'webp' where supported by the browser build Check format support and update S3 ContentType to match.

To capture a specific element, locate it and use its screenshot method: await page.locator('.report').screenshot({ path: '/tmp/report.png' }). To save an image to disk before upload, use a unique path under /tmp, never the function code directory. Uploading the returned screenshot bytes directly avoids an extra local copy.

5. Persist screenshots and handle batches

Lambda’s local temporary storage is not durable output storage. Upload the captured bytes to S3 or another persistent destination during the invocation. AWS’s screenshot example uses S3. Use unique object keys so concurrent calls do not overwrite one another, and set content type metadata to match the actual format.

For a list of URLs, AWS’s example uses a separate fan-out function that asynchronously invokes a screenshot worker for each URL. That pattern separates request coordination from browser work. Add your own limits on accepted batch size and concurrent invocations, and track each result so an individual failed page can be retried without losing successful outputs. AWS presents this as an example architecture, not a fixed throughput or completion-time guarantee.

6. Security considerations for URL capture

Accepting a URL means the function may make outbound requests to the supplied destination. Do not expose the sample handler as a public arbitrary-URL capture endpoint without establishing controls. For production, define an allowlist or equivalent destination policy, validate scheme and hostname, and consider redirects and DNS resolution when enforcing it. Avoid forwarding secrets or privileged cookies to arbitrary pages. Limit output size and execution time, and scope the Lambda role to the required S3 write path. Consult current AWS security guidance for the controls appropriate to your deployment; the example code is not a complete untrusted-URL defense.

7. Performance, reliability, and cost

  • Measure with real pages. Browser startup, page scripts, network activity, image dimensions, and browser version all affect duration and memory. The reviewed AWS sources provide no universal screenshot latency or throughput figure.
  • Tune memory and timeout together. Browser work can be CPU- and memory-intensive. Set a timeout above the slowest expected navigation and capture time, then monitor timeouts and memory use. Lambda has configurable memory and timeout settings; the appropriate values are workload-specific.
  • Manage temporary files. Keep extraction and cache data in /tmp, use unique names, and remove files your code creates. Size ephemeral storage for browser files and any intermediate outputs.
  • Close the browser in a finally block. This releases processes after successful captures and errors. Reusing an execution environment can reduce repeated setup work, but ensure pages and browser state do not leak between requests.
  • Account for the whole workflow. Lambda invocation and duration, memory configuration, container image storage, S3 storage and requests, and any orchestration or logging contribute to the bill. Calculate cost from your region, configured resources, traffic, retries, and retention; no workload-independent per-screenshot cost is claimed here.
  • Expect variable outcomes. A target may be slow, unavailable, rate-limited, or render differently in headless Chromium. Return a useful failure status, log a request identifier and safe diagnostic details, and retry only errors likely to be transient.

8. Troubleshooting

Symptom Likely cause Fix
Browser executable not found The browser was not included in the image, or Puppeteer expects it at a different path. Package a browser build compatible with the runtime and architecture; configure executablePath if required.
Shared library or launch error Required native browser dependencies are missing or incompatible with the image. Install the dependencies for the selected browser build in the container and test the built image in a Lambda-compatible environment.
Permission denied writing files The code writes outside /tmp or assumes the container filesystem is writable. Use /tmp for temporary files; upload durable output to S3.
Function times out during navigation The page is slow, waits never settle, or timeout is too short. Choose a wait condition suited to the page, wait for a specific selector where possible, and tune the Lambda timeout based on observed upper-bound workloads.
Screenshot is blank or missing content Capture happened before client-side rendering, or the page’s content depends on a later event. Wait for the content selector or a bounded delay; inspect navigation errors and page console output.
Image file is missing from S3 Upload failed, role lacks bucket permission, or the code returned before upload completion. Await the S3 request, grant scoped object-write access, and log the bucket/key and upload failure.
Image is cut off or unexpectedly tall Viewport dimensions or full-page behavior do not match the intended output. Set the viewport explicitly and choose fullPage intentionally; check for fixed-position elements and very long documents.
Works locally, fails in Lambda Local and deployed architectures, operating systems, browser builds, permissions, or writable paths differ. Build for the target Lambda architecture and validate the final container image with Lambda’s filesystem and user constraints.

9. Or skip the browser setup

If your application needs screenshots but you do not want to package and maintain Chromium in Lambda, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns an image or PDF; see the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets are removed. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

FAQ

Can a Lambda function return the screenshot directly instead of saving it to S3?

It can return image bytes if the caller and invocation path are designed to receive them, but output size and integration limits matter. For durable, shareable results, save the image to S3 and return its bucket and key or an authorized link.

Can I use Playwright instead of Puppeteer?

Possibly, if you package a browser build compatible with the Lambda runtime and architecture and satisfy its dependencies. Validate the current combination; the reviewed sources do not establish a current universal Lambda compatibility winner.

Should I use a ZIP or a container image?

Both are supported. Choose based on how you want to package and maintain the browser and its system dependencies. AWS’s Puppeteer example uses a container image.

What is a suitable timeout or memory size?

There is no single correct setting for every site. Measure representative slow pages, image sizes, and concurrency, then adjust memory, timeout, and temporary storage based on those results.