How to Speed Up Pyppeteer Page Loads on AWS Lambda
Measure Lambda startup, Chromium launch, navigation and readiness separately, then tune Pyppeteer waits, packaging, memory and concurrency.
Start by timing three separate phases: Lambda initialization, Chromium launch, and page navigation/readiness. Then make Pyppeteer wait only for the condition your job needs, reduce initialization work, reuse safe resources in warm environments, and tune memory and timeout from measurements. There is no universal speed-up percentage for Pyppeteer on Lambda: results depend on the runtime, Chromium build, page, network and workload.
1. Measure the slow phase first
A slow invocation can spend most of its time before page.goto() runs. AWS describes initialization as code download, runtime startup and initialization code, with package size and setup work among the contributors to latency. Instrument each phase and compare cold and warm invocations.
import asyncio
import json
import time
from pyppeteer import launch
def now_ms():
return time.perf_counter() * 1000
async def capture(url):
started = now_ms()
browser = await launch(
headless=True,
args=["--no-sandbox", "--disable-dev-shm-usage"],
)
browser_ready = now_ms()
page = await browser.newPage()
navigation_started = now_ms()
response = await page.goto(
url,
{
"waitUntil": "domcontentloaded",
"timeout": 30_000,
},
)
navigation_ready = now_ms()
await page.waitForSelector("main", {"timeout": 10_000})
content_ready = now_ms()
result = {
"status": response.status if response else None,
"launch_ms": round(browser_ready - started),
"navigation_ms": round(navigation_ready - navigation_started),
"readiness_ms": round(content_ready - navigation_ready),
"total_ms": round(content_ready - started),
}
await browser.close()
return result
print(asyncio.get_event_loop().run_until_complete(capture("https://example.com")))
Log the handler entry time as well, because the interval before handler execution belongs to the Lambda initialization phase. Record URL, success or failure, selected wait condition, output correctness and memory setting alongside duration.
2. Choose the earliest correct navigation condition
Pyppeteer goto() defaults to waitUntil='load'. Its documented choices include load, domcontentloaded, networkidle0 and networkidle2. The network-idle events require 500 ms with no more than the configured number of connections. See the Pyppeteer API reference.
| Condition | Use when | Risk |
|---|---|---|
domcontentloaded |
The required markup is available before images and other resources finish. | Content inserted after scripts or late requests may be missing. |
load |
The task needs the document load event and its dependent resources. | Unneeded images, fonts or third-party requests add delay. |
networkidle0 |
The page becomes completely quiet and that state means ready. | Polling, analytics or long-lived connections can prevent completion. |
networkidle2 |
A small number of ongoing connections is acceptable. | It still may not mean application data has rendered. |
For application pages, an explicit selector or JavaScript predicate often expresses readiness more accurately than a global network-idle event.
await page.goto(url, {
"waitUntil": "domcontentloaded",
"timeout": 30_000,
})
await page.waitForSelector("main article", {"timeout": 10_000})
await page.waitForFunction(
"() => document.querySelector('[data-ready=\"true\"]') !== null",
{"timeout": 10_000},
)
Confirm that the required data exists before returning. Lowering the wait condition without checking output can produce a faster but incorrect result. Pyppeteer documents a 30-second default navigation timeout and allows you to change it; a larger timeout prevents premature failure but does not make navigation faster.
3. Reduce Lambda initialization work
Keep imports and setup on the path that needs them
- Import only modules used by the handler.
- Move optional clients and expensive parsing out of module scope when most invocations do not need them.
- Remove unused dependencies from the deployment artifact.
- Measure browser-binary extraction or setup in your packaging model before optimizing it.
AWS identifies initialization code, dependency size and library or service setup as latency factors in its Lambda execution environment lifecycle documentation.
Package Chromium deliberately
Pyppeteer and Chromium must speak compatible DevTools protocols and match your Lambda runtime and architecture. Treat the browser binary as a versioned dependency. The chrome-aws-lambda repository shows a Puppeteer-oriented example and recommends at least 512 MB, with 1600 MB or more for its use case. That guidance does not prove compatibility with Pyppeteer or every current Lambda runtime. Verify maintenance status, architecture, package limits and protocol compatibility before adopting it, and do not copy launch flags blindly from a Node/Puppeteer example.
4. Reuse warm resources safely
Lambda may freeze and reuse an execution environment, so a later invocation can avoid some initialization. The environment can also be terminated without notice. AWS notes that /tmp contents may persist across reuse, but caches must tolerate missing or stale data.
import asyncio
from pyppeteer import launch
_browser = None
async def get_browser():
global _browser
if _browser is None or not _browser.process or _browser.process.poll() is not None:
_browser = await launch(
headless=True,
args=["--no-sandbox", "--disable-dev-shm-usage"],
)
return _browser
async def handler_async(event):
browser = await get_browser()
page = await browser.newPage()
try:
await page.goto(event["url"], {"waitUntil": "domcontentloaded", "timeout": 30_000})
await page.waitForSelector("main", {"timeout": 10_000})
return await page.title()
finally:
await page.close()
def lambda_handler(event, context):
return asyncio.get_event_loop().run_until_complete(handler_async(event))
Keep invocation-specific cookies, pages and user data isolated. Close pages in a finally block, detect a crashed browser, and recreate it. Do not assume a process survives the next request or that concurrent invocations share a safe page.
5. Tune memory, timeout and concurrency with measurements
Browser launch and rendering can be CPU-bound, while remote navigation is often network-bound. Compare several memory settings using representative URLs and record duration, failures and billed execution. AWS recommends reviewing Max Memory Used, using Lambda Power Tuning, and load-testing timeout choices in its Lambda best practices.
- Memory: test enough memory for Chromium stability, then compare cost and duration at higher settings.
- Timeout: set it above the slowest legitimate navigation plus readiness wait; a larger value only allows more time.
- Concurrency: test browser process limits, file descriptors and downstream site behavior under parallel invocations.
- Architecture: verify that the Chromium build and Python package support the selected architecture.
6. Use Provisioned Concurrency or SnapStart for startup predictability
Provisioned Concurrency pre-initializes execution environments and can reduce cold-start variability. It does not shorten the target website’s response time. SnapStart targets startup initialization for supported configurations. AWS documents configuration limitations, including runtime and storage restrictions, so check the current SnapStart documentation before choosing it. Neither feature replaces phase-level measurement.
7. A complete Pyppeteer handler pattern
import asyncio
import json
import os
import time
from pyppeteer import launch
_browser = None
async def browser_for_invocation():
global _browser
if _browser is None or _browser.process.poll() is not None:
_browser = await launch(
headless=True,
executablePath=os.environ.get("CHROMIUM_PATH"),
args=["--no-sandbox", "--disable-dev-shm-usage"],
)
return _browser
async def run(event):
url = event["url"]
wait_until = event.get("waitUntil", "domcontentloaded")
selector = event.get("selector", "main")
timeout = int(event.get("timeoutMs", 30_000))
started = time.perf_counter()
browser = await browser_for_invocation()
page = await browser.newPage()
try:
response = await page.goto(url, {"waitUntil": wait_until, "timeout": timeout})
if selector:
await page.waitForSelector(selector, {"timeout": min(timeout, 10_000)})
html = await page.content()
return {
"status": response.status if response else None,
"title": await page.title(),
"html": html,
"elapsedMs": round((time.perf_counter() - started) * 1000),
}
finally:
await page.close()
def lambda_handler(event, context):
return {
"statusCode": 200,
"headers": {"content-type": "application/json"},
"body": json.dumps(asyncio.get_event_loop().run_until_complete(run(event))),
}
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| High cold-start latency | Large artifact or expensive module-level initialization. | Measure the pre-handler phase, reduce dependencies and static work, and evaluate Provisioned Concurrency. |
goto() times out |
The chosen event never fires, the site is slow, or the network is blocked. | Test domcontentloaded, selector/function waits and network access separately. Increase timeout only when the longer wait is legitimate. |
| Missing dynamic content | Navigation completed before the application rendered data. | Wait for a specific selector or predicate and validate the returned content. |
| Network-idle wait never completes | Polling, analytics or streaming connections remain open. | Use a selector or function readiness condition. |
| Browser crashes or reports missing libraries | Incompatible Chromium package, architecture or insufficient memory. | Verify versions and architecture, inspect logs, and compare memory settings. |
| Warm invocations leak state | Pages, cookies or globals were reused across requests. | Create a fresh page per invocation, clear sensitive state and close pages in finally. |
| Works locally but fails in Lambda | Different sandbox, binary path, filesystem or runtime. | Use the Lambda-compatible executable, write temporary files only to /tmp, and test the packaged artifact. |
9. Performance, reliability and cost checklist
- Time handler entry, browser launch, navigation and readiness independently.
- Compare cold and warm runs across representative pages.
- Select the earliest wait condition that still guarantees correct output.
- Use explicit selectors or predicates for application readiness.
- Keep the deployment package and initialization path small.
- Reuse browsers only with crash detection and per-request page isolation.
- Load-test memory, timeout and concurrency settings.
- Record failures and output correctness with every timing sample.
- Account for Lambda duration, configured memory and any provisioned capacity when comparing cost.
Or skip the browser setup
If the goal is a clean screenshot rather than maintaining Chromium in Lambda, ScreenshotNeo provides a single API request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for options such as full-page or element capture, waits, custom CSS and JavaScript, headers, cookies, blocking, caching, signed links, async jobs, bulk capture and PDF output.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Should I always use domcontentloaded?
No. Use it only when the required content is present by that event, then verify with a selector or predicate.
Does more Lambda memory always make Pyppeteer faster?
No. It can provide more CPU, but remote navigation may dominate. Measure duration and cost at several settings.
Can I keep one browser open forever?
No. Warm environments are temporary and browsers can crash. Reuse when safe, detect failures and recreate the process.
Will Provisioned Concurrency speed up the target website?
It reduces startup variability. It does not change the remote site’s response or rendering time.


