ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Websites Hosted in India Using AWS Lambda

Build a Lambda screenshot service with Puppeteer and Chromium, save results to S3, and choose a region by measuring real targets.

By the ScreenshotNeo team4 October 20269 min read

Direct answer: Run a Lambda function with a Lambda-compatible Chromium build and a matching browser automation library, navigate to the target URL, wait for the content your capture needs, take a screenshot, and save the resulting bytes to S3. The website being hosted in India does not automatically mean Lambda must run in an India region. Test candidate regions against the actual sites and compare completion rate, render time, and image quality.

AWS has documented a Puppeteer and headless Chrome screenshot architecture that stores captures in S3. Its example is historical, so treat the architecture as a pattern and verify current runtime, browser, and dependency compatibility before deployment. [AWS screenshot architecture]

1. Choose the browser and packaging approach

Lambda does not provide a general desktop browser for your function. You need a Chromium build that works in the chosen Lambda runtime and an automation library whose version is compatible with that browser.

Approach Good fit Check before deployment
Puppeteer with Chromium in a Lambda container image You want the documented AWS screenshot-to-S3 pattern and control over bundled system dependencies. The AWS example is historical. Confirm the runtime, Chromium, and Puppeteer versions you choose are maintained and compatible.
Playwright with a Lambda-compatible Chromium package Your team already uses Playwright or needs its viewport, element, and full-page screenshot APIs. The playwright-aws-lambda package example documents older runtime support; evaluate its current maintenance and compatibility rather than assuming it works on a current runtime.

AWS Lambda container images can package browser dependencies. AWS documents a maximum uncompressed image size of 10 GB. If you use a custom runtime or a non-AWS base image, include the Lambda Runtime Interface Client as required by the image documentation. Keep the browser and automation-library versions pinned and update them deliberately. [Lambda container image requirements]

For a ZIP deployment, package a Lambda-compatible Chromium artifact and all required dependencies within the applicable deployment limits. A generic Puppeteer install does not guarantee that a suitable browser binary is present or executable in Lambda. Check the package documentation for the runtime and architecture it supports.

2. Create the function and S3 destination

  1. Create an S3 bucket for captures. Decide object naming, access permissions, encryption, and retention based on your workload. Do not make the bucket public just to retrieve screenshots; use an authorized application path or controlled access.
  2. Give the Lambda execution role permission to write only to the destination bucket or prefix it needs. Add the minimum permissions for any logging or orchestration the function uses.
  3. Build a deployment artifact containing the function, browser, automation library, and required system libraries. For container deployment, follow AWS’s image format and runtime interface requirements.
  4. Configure a bounded function timeout and memory allocation based on measurements for your pages. As project-specific guidance, chrome-aws-lambda says to allocate at least 512 MB and recommends 1600 MB or more; this is not an AWS minimum or a universal sizing benchmark. [chrome-aws-lambda project]

The AWS architecture uses a Lambda function to capture a page and store the image in S3. A separate function can fan out over a list of URLs when the workload needs batch processing. [AWS architecture example]

3. Capture a page with Puppeteer

The following handler shows the core flow. It assumes your deployment provides puppeteer-core and a compatible Chromium module that exposes an executable path and launch arguments. Package APIs vary, so adapt those two values to the Chromium build you select. Configure BUCKET_NAME and grant the function permission to write there.

const puppeteer = require('puppeteer-core');
const chromium = require('@sparticuz/chromium');
const { S3Client, PutObjectCommand } = require('@aws-sdk/client-s3');

const s3 = new S3Client({});
const bucket = process.env.BUCKET_NAME;

exports.handler = async (event) => {
  const url = event?.queryStringParameters?.url ?? event?.url;
  if (!url) {
    return { statusCode: 400, body: 'Provide a url' };
  }

  let parsed;
  try {
    parsed = new URL(url);
  } catch {
    return { statusCode: 400, body: 'Invalid url' };
  }
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    return { statusCode: 400, body: 'Only http and https URLs are supported' };
  }

  let browser;
  try {
    browser = await puppeteer.launch({
      args: chromium.args,
      defaultViewport: { width: 1440, height: 1000 },
      executablePath: await chromium.executablePath(),
      headless: true,
    });

    const page = await browser.newPage();
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
    // Replace or supplement this with a page-specific readiness signal when needed.
    await page.waitForTimeout(1000);
    const image = await page.screenshot({ type: 'png', fullPage: true });

    const key = `screenshots/${Date.now()}.png`;
    await s3.send(new PutObjectCommand({
      Bucket: bucket,
      Key: key,
      Body: image,
      ContentType: 'image/png',
    }));

    return {
      statusCode: 200,
      headers: { 'content-type': 'application/json' },
      body: JSON.stringify({ bucket, key }),
    };
  } catch (error) {
    console.error('Screenshot capture failed', error);
    return { statusCode: 502, body: 'Screenshot capture failed' };
  } finally {
    if (browser) await browser.close();
  }
};

This example uses domcontentloaded plus a short bounded delay as a starting point, not a guarantee that every site has finished rendering. Prefer a known selector or application-specific readiness signal for dynamic pages. Avoid returning arbitrary target URLs from an unauthenticated public function without controls: that can let callers make your function request unintended destinations.

4. Navigate and wait for the right content

The right wait condition depends on the page. A document response can finish before client-side rendering, images, or data requests finish. Use a specific readiness signal where possible:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.waitForSelector('[data-ready="true"]', { timeout: 10000 });

If the target has no reliable selector, use a bounded delay and accept that it is an approximation. Network-idle waits can be unsuitable for sites that keep analytics, polling, or streaming connections open. Set navigation and readiness timeouts so one slow page cannot consume the entire invocation.

For authenticated pages, establish only the cookies or headers required for authorized access. Sites may also use bot defenses, geofencing, DNS, TLS, or application checks; Lambda location alone does not ensure access. Test the actual target with authorization.

5. Choose viewport, element, or full-page capture

Pick the capture scope based on how the image will be used:

  • Viewport: capture the visible browser area. Set viewport width and height before navigation if responsive layout matters.
  • Element: capture a specific component when you need a chart, product card, or other region. Locate the element and use the automation library’s element screenshot API.
  • Full page: capture the document’s scrollable page. Lazy-loaded content may need scrolling or another explicit load strategy before capture.

Playwright documents viewport, element, and full-page screenshots. [Playwright screenshot modes]

// Playwright examples after page navigation
await page.screenshot({ path: '/tmp/viewport.png' });
await page.locator('.report-card').screenshot({ path: '/tmp/card.png' });
await page.screenshot({ path: '/tmp/full-page.png', fullPage: true });

Keep browser version, operating environment, viewport, device scale, fonts, and page state consistent when comparing captures over time. Rendering can vary across environments and settings. [Playwright screenshot documentation]

6. Select a Lambda region for India-hosted sites

The available sources do not establish a best region or provide an India-specific latency or reliability comparison. An India-hosted URL does not by itself require an India-based Lambda. Choose candidate regions that meet your account, data, and operational constraints, then run the same authorized URL set from each.

  1. Use representative targets, including pages with different scripts, media, and authentication needs.
  2. Record end-to-end render time, success or failure, and whether the screenshot contains the expected content.
  3. Repeat captures to account for variation and compare results under the same browser build and settings.
  4. Select the region based on measured results alongside operational and data requirements.

Do not rank regions by presumed proximity alone: DNS routing, the site’s hosting and delivery setup, and page dependencies can all affect the result. This is measurement guidance; the sources provide no regional benchmark. [AWS Lambda runtimes and regions documentation]

7. Handle output, scaling, and cost

Output and retention

Write the screenshot buffer directly to S3 with a content type that matches the selected format. Use stable keys if captures should replace an earlier result, or unique keys if history matters. Set object permissions and retention deliberately, and consider an S3 lifecycle policy for temporary captures. The AWS sample establishes S3 as a destination but does not prescribe a retention period.

Batching and concurrency

For a list of URLs, use bounded fan-out rather than starting unlimited browser work at once. AWS’s example describes using a second function to fan out over target URLs. Account for Lambda concurrency, downstream site limits, S3 request volume, and retries. Make output keys and retry behavior idempotent enough that a retry does not create confusing duplicate results.

Performance and reliability

  • Browser startup and cold starts add work beyond the page’s own render time. Measure the complete invocation in the regions and package configuration you plan to use.
  • Large pages, full-page captures, heavy JavaScript, and high-resolution images use more memory and time. Set limits and test representative pages.
  • Close the browser in a finally path, including after navigation, screenshot, or upload errors.
  • Pin browser and automation versions, then schedule deliberate security and compatibility updates.
  • Record enough structured context to diagnose failures, such as a request identifier, target hostname, phase, and error class. Avoid logging credentials or sensitive page data.

Cost

Captures can incur Lambda execution and storage charges, along with any orchestration or related AWS service costs. Estimate the planned request volume, memory, duration, concurrency, and object retention against current AWS pricing before production; the archived sample cautions that charges may apply beyond the Free Tier. No India-specific cost or performance figure is established by the sources. [AWS Lambda pricing] [Amazon S3 pricing]

8. Troubleshooting

Symptom Likely cause What to check or change
Executable not found The deployment has no Chromium binary or the configured path is wrong. Verify the selected package’s executable path and that the binary is included in the artifact.
Browser fails to launch Browser and library versions, architecture, runtime, or system dependencies do not match. Use a compatible, pinned browser build and inspect launch logs. Rebuild the artifact for the Lambda runtime and architecture.
Function times out during navigation The site is slow, a dependency hangs, or the wait condition never resolves. Use bounded navigation and readiness timeouts, choose a suitable wait condition, and test the target from the function’s region.
Screenshot is blank or missing content Capture ran before client rendering or lazy-loaded content completed. Wait for a page-specific selector or scroll/load the needed content before capturing.
Images or fonts differ from local output Rendering environment, browser version, font availability, device scale, or timing differs. Keep the environment and capture settings consistent; verify external assets load before capture.
S3 upload returns access denied The execution role lacks permission for the bucket or key prefix, or bucket policy blocks it. Review the role and bucket policy and grant only the required write access.
Some India-hosted pages fail while others work Target-specific bot checks, authentication, DNS, TLS, geofencing, or application behavior. Inspect the failing target’s authorized access path and test from the intended region. Do not assume the region alone is the cause.
Memory errors on long pages Large documents, full-page capture, or heavy resources exceed the configured memory. Measure representative pages, adjust memory, reduce capture scope or scale where suitable, and bound page resource use.

9. Or skip the browser setup

If you want the screenshot without packaging Chromium or operating Lambda browser infrastructure, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month with no card.

10. FAQ

Does the website have to be hosted in India?

No. The same architecture can capture reachable websites in other locations. Access and rendering depend on the specific target and request path.

Can I use Playwright instead of Puppeteer?

Yes, with a compatible Lambda Chromium build and runtime. Validate package maintenance and version support before deploying.

Should I use a Lambda function for every URL?

For small workloads, one invocation per capture is straightforward. For batches, use controlled fan-out and account for concurrency, retries, and target-site limits.

Is a screenshot guaranteed to match a user’s browser?

No. Browser version, fonts, viewport, device scale, timing, and page state can change the rendered result. Keep those inputs consistent when repeatability matters.

Primary references