How to Test Serverless Applications on AWS
Build a serverless testing strategy that starts with fast unit tests and verifies real AWS integrations, permissions, triggers, and workflows before release.
Test serverless applications at several levels: unit-test business logic quickly, use local tools for faster iteration, and deploy an isolated test stack to verify actual AWS integrations, permissions, triggers, and configuration. A local handler invocation can show how code responds to an event; it cannot prove that a deployed queue invokes the function or that its execution role grants the required permission.
AWS Prescriptive Guidance says, “Testing in the cloud is valuable for all phases of testing, including unit tests, integration tests, and end-to-end tests.” The practical goal is to choose each test for the evidence it can provide, then control the isolation and cost of cloud tests. AWS Prescriptive Guidance: Best practices for testing serverless applications
1. Build a testing pyramid for serverless
Serverless systems still benefit from unit, integration, and end-to-end tests. Their extra challenge is that important behavior lives in managed services and infrastructure configuration as well as function code.
| Test level | What it checks | Typical feedback | What it cannot establish alone |
|---|---|---|---|
| Unit | Business rules and transformations with controlled inputs and dependencies | Fast, runs near the code | AWS service behavior, deployed IAM, event-source wiring, or cloud configuration |
| Local function or API | Handler behavior and selected local runtime/API paths | Fast iteration, often with runtime-like containers | That the deployed trigger, role, quota, or service configuration works |
| Emulator | Selected service API interactions against an emulated environment | Useful middle layer for supported APIs | Exact AWS parity, production identity, IAM, or service quotas |
| Cloud integration | Real deployed service seams, permissions, and event delivery | Slower; requires an isolated environment | Every user journey or production-scale behavior unless explicitly covered |
| End-to-end | An application or workflow path through deployed components | Broadest integrated evidence, with more setup | All edge cases and load conditions not included in the scenario |
Use mocks to get broad, quick feedback on business logic. Retain cloud tests for claims that depend on AWS. For example, a mocked S3 success response does not prove the deployed Lambda role has s3:CreateBucket permission.
2. Make the handler easy to unit-test
Keep the Lambda handler as a thin adapter. It should validate and translate the incoming event, call ordinary application logic, and translate the result into the response format expected by its caller. Put business rules in functions that take ordinary values and dependencies explicitly.
# Python example: separate business logic from the Lambda adapter
def total_cents(items):
if not items:
raise ValueError("items must not be empty")
return sum(int(item["unit_price_cents"]) * int(item["quantity"]) for item in items)
def handler(event, context):
items = event.get("items")
if not isinstance(items, list):
return {"statusCode": 400, "body": "items must be a list"}
try:
total = total_cents(items)
except (KeyError, TypeError, ValueError):
return {"statusCode": 400, "body": "invalid items"}
return {"statusCode": 200, "body": str(total)}
Unit-test total_cents with valid, empty, malformed, and boundary inputs. Test the adapter separately for event validation and response shape. If the logic calls a database or AWS SDK, inject a small interface and use a fake or mock in unit tests; then verify the real integration in a cloud test.
Event cases to include
- Missing or incorrectly typed fields, empty collections, and invalid values.
- Optional fields omitted, explicitly null, or present with an empty value.
- Duplicate deliveries where the event source may retry; verify idempotency where required.
- Payloads near the limits and formats the application actually accepts.
- Failure paths: downstream dependency errors, timeouts, and validation failures.
3. Use local feedback deliberately
AWS SAM CLI can invoke functions locally and run a local API for rapid iteration. Its local execution uses Docker, which must be available. Runtime similarity is useful, but local execution is not a miniature copy of the complete AWS environment.
# Invoke a function from a local event file
sam local invoke OrderFunction --event events/order.json
# Start a local API for a SAM application
sam local start-api
Local function code can still call AWS APIs. Those calls may use configured AWS credentials and reach real resources. Use least-privilege credentials and nonproduction resources; keep test data separate from shared or production data. Review the SAM CLI documentation for setup and command options: AWS SAM CLI.
LocalStack is an optional emulator layer for selected AWS APIs. It can help exercise interactions without deploying every iteration, but API coverage and behavior can differ from AWS. Keep cloud checks for deployed identity, IAM, event-source configuration, quotas, and integrations whose exact behavior matters. LocalStack documentation
4. Deploy a test stack to check AWS contracts
Deploy an isolated test environment that uses the same relevant service types and configuration as the application. Exercise the actual entry point and check the downstream result. This is where you verify the seams between managed services, function code, IAM, and infrastructure.
- Deploy the application and dependencies using the infrastructure definition used by the project.
- Give the stack a developer, branch, or run identifier so concurrent tests do not share mutable resources.
- Initiate the real trigger: an HTTP request through API Gateway, a message on the SQS queue, an object uploaded to storage, an EventBridge event, or a workflow start.
- Observe invocation and downstream state, then assert the expected result and relevant failure behavior.
- Delete or reset the test data and clean up the stack when the run is complete.
Do not pass a hand-crafted SQS event directly to a handler and treat that as proof that the deployed SQS mapping works. Put a valid message on the deployed queue and confirm invocation and the expected downstream effect. Check the queue mapping, message constraints, visibility timeout, and execution-role permissions.
Cloud integration checklist
- Did the real trigger invoke the intended function with the expected event shape?
- Does the deployed execution role permit the required action on the intended resource?
- Are timeout, memory, retry, and event-source settings suitable for the operation?
- Do API, queue, storage, database, EventBridge, and workflow settings match the intended contract?
- Does the application handle duplicate delivery and downstream failures correctly?
- Can parallel test runs execute without overwriting one another’s data?
5. Test asynchronous work with bounded waits
For queues, events, and background workflows, success may occur after the initiating request returns. A test should wait for a downstream result within a defined deadline, rather than sleeping for an arbitrary long period or assuming the first poll will see the result.
# Pseudocode for an asynchronous integration test
run_id = unique_id()
publish_event({"run_id": run_id, "payload": test_payload})
deadline = now() + test_timeout
while now() < deadline:
result = read_result_for(run_id)
if result is not None:
assert result == expected_result
break
wait(short_poll_interval)
else:
fail("No downstream result before the test deadline")
remove_test_data(run_id)
Use a unique correlation or run ID in every test event and query only that run’s result. Set the timeout from the expected behavior and test environment, report a useful failure when it expires, and clean up both successful and failed runs. In shared accounts, isolate stacks by branch or developer and prevent concurrent runs from modifying the same test records.
6. Test Step Functions with supported methods
For state-machine logic, use the AWS Step Functions TestState API to test state behavior in isolation, then run cloud integration tests for the workflow’s actual service integrations and deployed configuration. AWS labels Step Functions Local unsupported and without feature parity; do not rely on it as a supported production-grade testing strategy.
See the official references for the TestState API and Step Functions Local limitations.
7. Add performance and release checks
Run performance tests in an environment that reflects the cloud services and configuration relevant to the workload. A local benchmark cannot establish how the full deployed system behaves under service limits, network conditions, or concurrency.
- Observe Lambda maximum memory use and initialization duration; adjust memory and initialization work based on measured behavior.
- Check relevant service quotas before load or concurrency tests, and request quota changes through the appropriate AWS process when needed.
- For VPC-connected functions, account for available subnet IP address capacity as concurrency changes.
- Separate performance testing from routine fast feedback so its setup and resource use are deliberate.
- Run cloud checks in CI before promoting to QA, staging, or production, and alert on expected spend.
AWS recommends cloud testing while also emphasizing isolation and cost control. Use dedicated or isolated environments where possible, least-privilege access, bounded test runs, unique resources, and cleanup. The exact spend depends on the services and duration used; the test design should make resource creation and teardown visible.
8. Comparison of testing approaches
| Approach | Speed | AWS fidelity | IAM and infrastructure validation | Isolation and setup |
|---|---|---|---|---|
| Mocks and unit tests | Fast | Low for managed-service behavior | No | Easy to isolate; low setup |
| SAM local | Fast iteration | Useful runtime feedback, partial environment | Limited; AWS API calls may reach real resources | Docker and careful credentials needed |
| Service emulator | Fast to moderate | Limited to implemented APIs and behavior | Does not establish deployed AWS identity or configuration | Requires emulator setup; isolate its data |
| Deployed cloud tests | Slower | Highest for the deployed services actually exercised | Yes, for tested roles and resources | Needs environment, cost controls, and cleanup |
9. Troubleshooting common failures
| Symptom | Likely cause | What to check |
|---|---|---|
| Local test passes, deployed call gets AccessDenied | The local identity has permissions the Lambda execution role lacks | Inspect the deployed role’s policies, resource scope, and the action used by the code |
| Handler test passes, queue message does not invoke Lambda | The deployed event-source mapping, queue policy, or trigger configuration is missing or invalid | Send a valid message to the actual queue; inspect mapping state, permissions, message shape, and visibility timeout |
| SAM local command cannot start the runtime | Docker is unavailable or not running, or the local runtime setup is incomplete | Start Docker, check local prerequisites, and confirm the function runtime configuration |
| SAM local unexpectedly reads or writes cloud data | Application code made AWS API calls using configured credentials | Use nonproduction resources, dedicated least-privilege credentials, and explicit test configuration |
| Async test flakes or times out | It assumes immediate completion, polls shared data, or has no run isolation | Use a unique run ID, bounded polling with a clear deadline, and an isolated result key |
| Emulator passes but AWS test fails | The required API behavior or identity/configuration differs from the emulator | Keep an AWS integration test for that seam; confirm service permissions and deployed settings |
| Load test fails before reaching expected concurrency | A service quota or VPC subnet address capacity may constrain the test | Review relevant quotas, Lambda configuration, and available subnet IPs |
| Cloud tests create unexpected spend or leave resources behind | Resources are shared, unbounded, or not cleaned up after failure | Use isolated stacks, cost alerts, bounded runs, and cleanup in both success and failure paths |
10. Browser checks for serverless web applications
If the serverless application exposes a web page, test browser-visible outcomes separately from backend contracts. A screenshot can help catch rendering changes, but it does not prove IAM permissions, queue delivery, or workflow correctness. For repeatable checks, use a stable test URL and control test data and environment.
DIY: capture a page with a browser
Playwright can launch a browser, navigate to a page, and save a screenshot. Install Playwright and its browser once in the project, then run this Node.js script. Replace the URL with the deployed test page.
npm install --save-dev playwright
npx playwright install chromium
// screenshot.mjs
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
await page.goto('https://your-test-site.example', { waitUntil: 'networkidle', timeout: 60000 });
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
Browser tests need a reachable environment, browser binaries, and a clear readiness condition. If the page keeps network requests open, use a selector or explicit application-ready signal instead of waiting indefinitely for network idle. Keep secrets out of source files and pass test credentials through the runner’s secret store.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request captures a URL as PNG, JPEG, WebP, or PDF. For a deployed test page, use the API instead of installing and managing a browser; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://your-test-site.example -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://your-test-site.example"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://your-test-site.example' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. These page captures complement backend tests; they do not verify AWS service integration or permissions.
Sign up for 1,000 free screenshots a month with no card.
FAQ
Do serverless applications need end-to-end tests?
Use them for critical deployed journeys and workflow behavior, alongside faster unit and integration checks. They provide broad evidence for the scenarios they exercise, not exhaustive coverage.
Can a mocked AWS SDK call replace a cloud integration test?
No. A mock checks your code against the behavior you programmed into the mock. It does not confirm real service behavior, deployed permissions, or infrastructure wiring.
Should every test run in AWS?
No. Keep fast, isolated logic tests close to the code. Run cloud tests where the result depends on deployed AWS behavior or configuration, and isolate and control those runs.
Does a screenshot test validate a serverless backend?
No. It can show browser-visible output for a page. Test backend permissions, triggers, and downstream effects through their actual service paths.


