ScreenshotNeo

BlogHow-to

Generate a Website Screenshot in Java with Playwright and Save It to S3

Capture a website with Playwright Java and upload the PNG to Amazon S3. Compare in-memory and file-based workflows, configure captures, and fix common failures.

By the ScreenshotNeo team4 October 202610 min read

Use Playwright Java to navigate to the page and call Page.screenshot(). For a direct upload, leave the screenshot path unset so Playwright returns the image as a byte[], then send those bytes to Amazon S3 with AWS SDK for Java 2.x putObject. This avoids an intermediate local file. If you also need a local artifact for debugging or retention, capture to a path and upload that file instead.

The examples below capture a full-page PNG. Playwright screenshots default to the current viewport; setFullPage(true) includes the full scrollable document. See the official Playwright Java screenshots guide and the Page API reference.

1. Add the Java dependencies and install a browser

Use Playwright Java and AWS SDK for Java 2.x. The following Maven dependencies use the current versions configured for your project; choose versions according to your dependency-management policy. The SDK modules shown are Playwright and S3.

<dependencies>
  <dependency>
    <groupId>com.microsoft.playwright</groupId>
    <artifactId>playwright</artifactId>
    <version>YOUR_PLAYWRIGHT_VERSION</version>
  </dependency>
  <dependency>
    <groupId>software.amazon.awssdk</groupId>
    <artifactId>s3</artifactId>
    <version>YOUR_AWS_SDK_VERSION</version>
  </dependency>
</dependencies>

After adding Playwright, install its browser binaries using the CLI for the Playwright version in your project. The official installation guide describes the Java setup, supported operating systems, and system requirements: Playwright Java installation. Browser launches are headless by default, which is suitable for servers and CI environments.

Provide AWS credentials through the AWS SDK’s normal credential-provider chain, and configure the AWS region. The bucket must already exist, and the calling identity needs permission to put objects at the chosen key. Avoid hard-coding credentials in source code.

2. Capture bytes and upload them directly

This complete example takes the target URL, bucket, and object key from environment variables. It captures the full page as a PNG byte array and uploads it with an explicit content type. The AWS SDK discovers credentials through its standard provider chain.

import com.microsoft.playwright.Browser;
import com.microsoft.playwright.Page;
import com.microsoft.playwright.Playwright;
import software.amazon.awssdk.regions.Region;
import software.amazon.awssdk.services.s3.S3Client;
import software.amazon.awssdk.services.s3.model.PutObjectRequest;
import software.amazon.awssdk.core.sync.RequestBody;

public class WebsiteScreenshotToS3 {
  public static void main(String[] args) {
    String url = requiredEnv("TARGET_URL");
    String bucket = requiredEnv("S3_BUCKET");
    String key = requiredEnv("S3_KEY");
    String regionName = System.getenv().getOrDefault("AWS_REGION", "us-east-1");

    try (Playwright playwright = Playwright.create();
         S3Client s3 = S3Client.builder()
             .region(Region.of(regionName))
             .build()) {
      Browser browser = playwright.chromium().launch();
      try {
        Page page = browser.newPage();
        page.navigate(url, new Page.NavigateOptions()
            .setWaitUntil(com.microsoft.playwright.options.WaitUntilState.DOMCONTENTLOADED)
            .setTimeout(60_000));

        byte[] png = page.screenshot(new Page.ScreenshotOptions()
            .setFullPage(true)
            .setType(com.microsoft.playwright.options.ScreenshotType.PNG));

        PutObjectRequest request = PutObjectRequest.builder()
            .bucket(bucket)
            .key(key)
            .contentType("image/png")
            .build();
        s3.putObject(request, RequestBody.fromBytes(png));

        System.out.printf("Uploaded %d bytes to s3://%s/%s%n", png.length, bucket, key);
      } finally {
        browser.close();
      }
    }
  }

  private static String requiredEnv(String name) {
    String value = System.getenv(name);
    if (value == null || value.isBlank()) {
      throw new IllegalArgumentException("Set environment variable " + name);
    }
    return value;
  }
}

Run it after setting TARGET_URL, S3_BUCKET, S3_KEY, and (optionally) AWS_REGION. The Java installation and upload snippets follow the documented Playwright and AWS APIs; adapt dependency versions and environment configuration to your deployment.

export TARGET_URL='https://example.com'
export S3_BUCKET='my-screenshot-bucket'
export S3_KEY='captures/example.png'
export AWS_REGION='us-east-1'
java WebsiteScreenshotToS3

The direct route holds the encoded image in process memory until the upload finishes. For ordinary screenshot sizes this is straightforward. If image sizes may be large or you need a local copy, use the file-based route below.

3. Choose bytes or a local file

Route Use it when Trade-off
Screenshot bytes, then RequestBody.fromBytes No local artifact is needed and the capture fits comfortably in memory. The screenshot byte array remains in memory during upload.
Screenshot to a path, then RequestBody.fromFile You need a local artifact for debugging, retention, or an existing file-based pipeline. You must manage disk space, unique paths, and cleanup.

This is a design comparison based on the APIs, not a performance benchmark. Measure in your deployment if throughput or memory is a concern.

Save to a path, then upload the file

import java.nio.file.Files;
import java.nio.file.Path;
import com.microsoft.playwright.Page;
import software.amazon.awssdk.core.sync.RequestBody;
import software.amazon.awssdk.services.s3.model.PutObjectRequest;

Path output = Path.of("target", "captures", "page.png");
Files.createDirectories(output.getParent());
page.screenshot(new Page.ScreenshotOptions()
    .setPath(output)
    .setFullPage(true)
    .setType(com.microsoft.playwright.options.ScreenshotType.PNG));

s3Client.putObject(
    PutObjectRequest.builder()
        .bucket(bucketName)
        .key(objectKey)
        .contentType("image/png")
        .build(),
    RequestBody.fromFile(output));

Use a unique local filename when captures can run concurrently. Delete temporary files after a successful upload if the application does not retain them. If the upload fails, decide whether to keep the file for diagnosis or retry and clean it later.

Use an input stream when needed

AWS SDK for Java 2.x supports stream-backed request bodies. When the stream length is known, supply the exact byte count. This helps the SDK send a well-defined request body; do not guess a length or pass a value that differs from the bytes available.

try (java.io.InputStream input = new java.io.ByteArrayInputStream(png)) {
  s3Client.putObject(request, RequestBody.fromInputStream(input, png.length));
}

For this byte-array example, RequestBody.fromBytes(png) is simpler. For file-backed data, prefer RequestBody.fromFile(path). Refer to AWS’s guidance on uploading streams with AWS SDK for Java 2.x and streaming operation differences.

4. Set the page state before capture

A screenshot is only as useful as the page state it captures. Choose the navigation wait condition based on the site’s behavior, then wait for a concrete selector or other condition that represents readiness. Do not assume that network idle means every application has finished rendering.

page.navigate(url, new Page.NavigateOptions()
    .setWaitUntil(com.microsoft.playwright.options.WaitUntilState.DOMCONTENTLOADED)
    .setTimeout(60_000));
page.locator("main article").waitFor();
byte[] png = page.screenshot(new Page.ScreenshotOptions().setFullPage(true));

Replace main article with a selector that is meaningful for the target site. If the target renders content after an API response or client-side transition, wait for the resulting content rather than using a fixed delay as the only readiness check. A delay can be useful for known animations or scheduled updates, but it adds time and may still capture too early.

5. Screenshot options that affect the result

Need Playwright option or approach Notes
Current visible area Default screenshot behavior Captures the viewport.
Entire scrollable page setFullPage(true) Captures beyond the viewport; very long pages can create large images.
Specific region setClip(...) Use a clip rectangle when the capture should cover a precise region.
Capture one element Use the locator screenshot API Useful for a chart, card, or component rather than the whole page.
Image format Screenshot type option PNG is lossless; JPEG and WebP have quality controls. Set the S3 contentType to match the format.
Scale CSS-pixel or device-pixel scale Device-pixel output can increase dimensions and storage size.
Dynamic regions Masks and stylesheet options Can cover or alter regions that change between captures or contain sensitive details.
Animation Animation handling option Choose how animations are handled when you need more repeatable output.

Playwright’s screenshot API documents clipping, format, scale, masks, stylesheet handling, and animation behavior. Consult the Java Page API for the exact option names available in your installed version. Screenshot pixels may vary across operating systems, fonts, browser builds, and page state; control the capture environment if visual comparisons require repeatability.

6. S3 key, access, and content type

  • Choose a stable or unique key. A stable key overwrites the previous object at that key. Include a timestamp, job identifier, or content-specific identifier if each capture must be retained.
  • Set the content type. Use image/png, image/jpeg, or image/webp to match the encoded screenshot. A mismatch can confuse browsers and downstream consumers.
  • Keep access private by default. Grant the application permission to upload to the required bucket and prefix. Use your existing access policy or a deliberate sharing mechanism; do not make a bucket public just to retrieve a screenshot.
  • Configure region and credentials explicitly for the environment. The sample takes region from AWS_REGION and uses the SDK credential chain. Workload roles are preferable to embedding long-lived secrets in code.

7. Troubleshooting

Symptom Likely cause Fix
Browser launch fails in CI or a container Browser binaries or required operating-system dependencies are missing. Install the browser for the Playwright version in use and satisfy its documented system requirements. Confirm the deployed image matches the supported OS setup.
Navigation times out The site is slow, long-lived, or waits on resources that do not settle. Use a suitable navigation condition, set an intentional timeout, and wait for the content selector the capture needs.
Screenshot is blank or incomplete Capture ran before client-rendered content appeared, or a consent overlay/auth wall blocked the page. Wait for the expected content and handle the target site’s legitimate consent or authentication flow.
Full-page capture is unexpectedly large or slow The page has a very long document or large images. Capture a clip or element if that meets the requirement; use an appropriate format and scale.
AccessDenied from S3 The active AWS identity lacks permission for the bucket or key prefix, or a policy denies the write. Check the active credential source, bucket/key, region, and object-write permissions.
Bucket not found or redirect/region error The bucket name or configured region does not match the bucket’s location. Confirm the bucket name and set the matching AWS region.
Uploaded file cannot be displayed Content type does not match the screenshot format, or the object is inaccessible to the viewer. Match contentType to the image encoding and verify access through the intended application path.
Stream upload reports a length or request-body issue The stream length is missing or inaccurate, or the stream cannot be replayed as expected. Supply the exact content length when known. Use bytes or a file-backed body when that better fits the data source.

8. Performance, reliability, and cost

For a single capture, browser startup and page rendering are usually part of the job’s runtime; the S3 request is a separate step. Reusing a browser process across multiple captures can avoid repeated launches, but manage pages and browser lifecycle carefully, and isolate jobs if they must not share state. These are operational considerations, not a benchmark claim.

Full-page screenshots, high device scale, and lossless formats can increase memory use, upload time, and S3 storage. Choose viewport versus full-page, image type, and scale to fit the use. A byte-array workflow requires memory for the encoded image; a file route shifts that artifact to disk and requires cleanup. For reliability, report capture and upload failures separately, use a deliberate retry policy for transient upload errors, and make object keys idempotent or unique according to whether retries should overwrite or retain captures.

Costs depend on your browser compute environment, runtime, and S3 storage and request usage. The research sources do not provide a benchmark or a cost estimate for a particular workload, so measure your job and estimate costs with your deployment’s actual image sizes and capture frequency.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. One GET request returns an image or PDF; the response can be saved and uploaded to S3 by your application. See the ScreenshotNeo API docs for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

In Java, make the same GET request with your HTTP client, write the returned bytes to a file or pass them to an S3 request body, and keep the API key in configuration. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo free and get 1,000 screenshots a month with no card.

FAQ

Can I upload the screenshot bytes directly to S3?

Yes. When you omit the screenshot path, Playwright returns a byte[]. Pass it to RequestBody.fromBytes with an S3 PutObjectRequest.

Does Playwright run headlessly on a server?

Yes. Browser launches are headless by default. Install the browser binaries and operating-system dependencies required by the Playwright version and deployment image.

Should I use PNG or JPEG?

Use PNG when lossless output matters. JPEG or WebP with an appropriate quality setting can reduce image size when the target use accepts lossy output. Keep the S3 content type aligned with the chosen encoding.

Will the same URL always produce identical pixels?

No. Page content, timing, browser version, fonts, and operating system can affect the rendered result. Control the environment and wait for a meaningful ready state when you need repeatable captures.