ScreenshotNeo

BlogHow-to

Capture Website Screenshots in Python with Playwright and Upload Them to S3

Capture a viewport, full page, or element with Playwright, then upload the image to S3 with Boto3. Learn safe credential, storage, and access patterns.

By the ScreenshotNeo team4 October 20268 min read

Use Playwright Python to capture the page, then Boto3 to upload the result to Amazon S3. For a local-file workflow, pass the screenshot path to upload_file. For an in-memory workflow, pass the bytes returned by page.screenshot() through a readable binary file object to upload_fileobj. The upload and whether anyone can read the object are separate concerns: configure AWS credentials and permissions for the runtime, and keep S3 public access blocked unless you have a deliberate, reviewed sharing requirement.

1. Install Playwright and Boto3

This example uses Playwright’s synchronous API and Boto3’s standard credential provider chain. It does not put AWS access keys in source code.

python -m pip install playwright boto3
python -m playwright install chromium

Make sure the machine or container can launch Chromium. In a Linux container, install the browser dependencies as required by your deployment image. Configure AWS credentials using an appropriate provider for the environment, such as an IAM role attached to an AWS compute resource or environment-based credentials. On EC2, an attached IAM role can provide credentials through the provider chain.

2. Capture a full-page screenshot and upload it

Save the screenshot to a local path, then upload that filename to the bucket and object key. The browser is closed even if navigation or capture raises an exception.

from pathlib import Path

import boto3
from playwright.sync_api import sync_playwright

url = "https://example.com"
bucket = "your-bucket"
key = "screenshots/example.png"
output = Path("website.png")

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page(viewport={"width": 1440, "height": 900})
        page.goto(url, wait_until="load", timeout=30_000)
        page.screenshot(path=str(output), full_page=True, type="png")
    finally:
        browser.close()

s3 = boto3.client("s3")
s3.upload_file(
    str(output),
    bucket,
    key,
    ExtraArgs={"ContentType": "image/png"},
)
print(f"Uploaded s3://{bucket}/{key}")

upload_file accepts a local filename, bucket name, and object key. Boto3’s managed transfer handles large files by splitting them into chunks and uploading those chunks in parallel. The ExtraArgs parameter can carry supported upload options, including metadata; ContentType helps consumers interpret the object as a PNG.

3. Choose what to capture

Need Playwright approach Notes
Visible browser viewport page.screenshot(path="shot.png") Captures the current viewport by default.
Entire scrollable page page.screenshot(path="shot.png", full_page=True) Captures beyond the viewport. Very long pages can produce large images and take longer to render and upload.
One element page.locator("main").screenshot(path="main.png") Use a selector that identifies the intended element. The locator screenshot method can also return bytes when no path is given.
In-memory processing or transfer data = page.screenshot() Returns image bytes; wrap them in a readable binary stream for upload_fileobj.

Choose PNG for lossless output, or use JPEG or WebP where supported and suitable for your downstream use. Explicitly select the output type and use a matching key extension and content type. Screenshot output is only repeatable when relevant inputs are controlled: page content, viewport, device scale, fonts, animations, timing, cookies, and other state can change the result. Playwright screenshot options include animation handling; consult the official API for the current option set.

4. Upload screenshot bytes without a local image file

When no path is supplied, page.screenshot() returns bytes. upload_fileobj expects a readable binary file-like object, so wrap those bytes in io.BytesIO.

from io import BytesIO

import boto3
from playwright.sync_api import sync_playwright

url = "https://example.com"
bucket = "your-bucket"
key = "screenshots/example.png"

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page()
        page.goto(url, wait_until="load", timeout=30_000)
        image_bytes = page.screenshot(full_page=True, type="png")
    finally:
        browser.close()

s3 = boto3.client("s3")
s3.upload_fileobj(
    BytesIO(image_bytes),
    bucket,
    key,
    ExtraArgs={"ContentType": "image/png"},
)
print(f"Uploaded s3://{bucket}/{key}")

This avoids writing the screenshot to disk, though the image bytes still occupy memory. For very large full-page captures or a pipeline that needs retryable local artifacts, a temporary file can be a better handoff.

5. Configure credentials and access safely

Boto3 searches its configured credential providers. Use the identity and permissions appropriate to the runtime: for example, an attached role in AWS, or credentials supplied by your deployment platform. Avoid hard-coding access keys in scripts, container images, or repositories.

Grant the runtime only the permissions needed for its bucket and key prefix. The exact policy depends on your deployment and account controls, so there is no single universal policy to copy. Also distinguish these outcomes:

  • Upload succeeded: the runtime was authorized to write the object.
  • Object is readable by a consumer: a separate access policy or sharing mechanism permits that consumer to read it.

New S3 buckets, access points, and objects do not allow public access by default. S3 Block Public Access can override policies and ACL permissions that would otherwise allow public access. AWS’s Boto3 bucket creation documentation recommends keeping Block Public Access enabled and ACLs disabled for most modern use cases. Do not add a public ACL as a routine upload setting. If external access is required, choose and review a deliberate sharing mechanism and the relevant bucket or organization policies.

6. Make capture runs more reliable

  • Wait for the right page state. wait_until="load" waits for the load event, but pages that render content afterward may need a locator wait or an application-specific readiness condition.
  • Set timeouts. Navigation can hang on slow or unresponsive sites. Choose timeouts based on the workload and handle failures so one URL does not stop a batch.
  • Use stable selectors. For element captures, wait for the target locator and verify it resolves to the expected content before taking the shot.
  • Close resources. Close pages and browsers in cleanup paths. Reuse a browser process for a controlled batch when appropriate, while isolating page state between jobs.
  • Retry selectively. Retry transient navigation or upload failures with bounded attempts and backoff. Avoid unlimited retries, and make object keys deterministic or otherwise account for duplicate writes.
  • Control visual inputs. Set viewport and relevant device scale, wait for fonts or application content when necessary, and disable or finish animations when consistent output matters.
  • Watch image size. Full-page screenshots can be large. Consider element or viewport capture, a more compact format, or a downstream resize when the use case permits it.

For asynchronous Playwright, use async_playwright and await browser, navigation, and screenshot operations. Boto3 is synchronous; in an async service, run the upload outside the event loop or use an async-compatible transfer design rather than blocking other tasks.

7. Troubleshooting

Symptom Likely cause Fix
Browser executable or shared-library error Chromium or its system dependencies are missing from the environment. Install Chromium with Playwright’s browser installer and ensure the deployment image includes required OS libraries.
Navigation timeout The site is slow, keeps network activity open, or never reaches the selected lifecycle event. Set a suitable timeout, wait for a specific element or application-ready condition, and handle the failed URL separately.
Screenshot is blank or incomplete The page had not rendered the target content, a lazy section had not loaded, or a selector matched the wrong element. Wait for the actual content, scroll or otherwise trigger lazy content when needed, and verify the locator before capture.
NoCredentialsError No usable AWS credentials are available to Boto3. Configure the environment’s credential provider or attach the intended role; do not paste long-lived secrets into the code.
AccessDenied The active identity lacks permission for the requested bucket/key operation, or a bucket, organization, or endpoint policy blocks it. Confirm the identity and review the applicable IAM and S3 policies with the bucket owner or administrator.
Upload works but URL returns access denied The object is private; uploading does not make it public. Use an approved access mechanism or keep the object private. Review Block Public Access and policy requirements before enabling public access.
Image opens with the wrong type or download behavior The object metadata or key extension does not match the captured format. Set the correct ContentType and use a matching extension.
Memory use spikes A full-page image is large, or bytes and copies are retained in memory. Capture only the necessary region, write to a path, release byte buffers after upload, or process jobs with bounded concurrency.

8. Performance, reliability, and cost

End-to-end time includes browser startup, page loading, rendering, screenshot encoding, and S3 transfer. Reusing a browser for multiple captures can reduce startup overhead, while bounded concurrency prevents a batch from exhausting memory or overwhelming target sites. Boto3’s managed upload can parallelize large-file transfer; small screenshots generally have little to gain from multipart transfer.

Playwright and Boto3 do not make a visual capture deterministic by themselves. Dynamic content, personalization, consent state, timing, and browser rendering can vary. If images are used as evidence or comparison artifacts, record relevant capture settings and keep the page state controlled.

Costs depend on the compute environment, S3 storage and requests, and any data transfer. The research does not establish a universal price or performance benchmark. Set retention and lifecycle behavior appropriate to your use case, and avoid storing duplicate captures indefinitely if they are not needed.

9. Or skip the browser setup

If you need a screenshot without managing Playwright, Chromium, and the upload handoff, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its capture flow accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server gives Claude, Cursor, and other MCP clients screenshot tools.

See the ScreenshotNeo API documentation for request options. This Python example writes the returned image response to a file:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

The same endpoint can be called with cURL or Node.js:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is available on every plan. Sign up for the free plan to get started.

10. FAQ

Can I upload a Playwright screenshot directly to S3 without saving it locally?

Yes. Omit the screenshot path to get bytes, wrap them in io.BytesIO, and pass the readable object to upload_fileobj.

Does uploading with Boto3 make my screenshot public?

No. Upload permission and read access are separate. S3 objects are private by default under the standard protections described above.

How do I capture only one section of a page?

Use a locator for the section and call its screenshot() method after waiting for the intended element to be present.

Can I use this flow for PDFs?

Playwright has a separate page PDF capability for supported browser contexts; the examples here focus on image screenshots. Check the current Playwright API and your output and storage requirements before switching formats.

Should I use a path or bytes?

Use a path when you want a straightforward, inspectable artifact or need to limit memory use. Use bytes when in-memory processing or a direct file-like upload fits the job size and runtime.