ScreenshotNeo

BlogHow-to

How to Export Apify Website Screenshots to Amazon S3

Store Apify screenshots as image files, then transfer their bytes to Amazon S3 with a clear key convention, safe permissions, and reliable retries.

By the ScreenshotNeo team4 October 20269 min read

Direct answer: store the screenshot bytes in an Apify key-value store, then read those bytes and upload them to an Amazon S3 object key with S3 PutObject. Apify’s S3 Uploader is documented for exporting dataset content; do not assume it copies binary screenshot records from a key-value store. Use a dataset for structured metadata or a manifest, and a separate upload step for each image.

Choose the right storage path

A screenshot is a binary file. Apify’s key-value stores support file-like records and are explicitly suitable for screenshots. Apify’s JavaScript SDK v1.3 example captures a screenshot buffer and writes it to a key-value store with the image/png content type. Check the current SDK syntax for the Actor version you use.

What you need Use
Keep PNG or JPEG bytes in Apify Key-value store record
Record URL, timestamp, status, and destination key Dataset row or manifest
Export structured dataset rows to S3 Apify S3 Uploader, after checking its current input and output layout
Export image bytes to S3 Read the key-value record and upload its body using S3 PutObject

Apify datasets are intended for sequential structured items and standard data exports. They are a useful companion for a manifest, but should not be treated as the screenshot file store. Check storage naming and retention before relying on an unnamed store; the dataset documentation’s unnamed-dataset expiry statement should not be assumed to describe key-value-store retention.

Capture and store a screenshot in Apify

The following illustrates the documented JavaScript SDK pattern. It assumes page is a Playwright page already opened by your Actor and url is the page URL. Adapt the imports and store setup to the SDK and Actor version in use.

const { KeyValueStore } = require('apify');

const url = 'https://example.com';
const page = /* your Actor's Playwright page */;
const store = await KeyValueStore.open('screenshots');

await page.goto(url, { waitUntil: 'networkidle' });
const image = await page.screenshot({ fullPage: true, type: 'png' });
const key = `screenshot-${Date.now()}.png`;

await store.setValue(key, image, { contentType: 'image/png' });
console.log(JSON.stringify({ url, key, contentType: 'image/png' }));

This is a pattern, not a guarantee that every current Actor exposes the same page object or uses the same import style. Choose an explicit store name when you need to identify the data later. Use a record key that is unique within that store; do not use an unnormalized source URL directly as a key.

Upload the image bytes to Amazon S3

For an automated export, read the record bytes and pass them as the S3 request body. This Node.js example uses AWS SDK for JavaScript v3. Install @aws-sdk/client-s3 in the Actor or service that performs the transfer, and provide AWS credentials through the environment or the runtime’s IAM role.

const { S3Client, PutObjectCommand } = require('@aws-sdk/client-s3');
const { KeyValueStore } = require('apify');

const bucket = process.env.S3_BUCKET;
const region = process.env.AWS_REGION;
const storeName = process.env.APIFY_STORE_NAME || 'screenshots';
const recordKey = process.env.APIFY_RECORD_KEY;
const objectKey = process.env.S3_OBJECT_KEY;

if (!bucket || !region || !recordKey || !objectKey) {
  throw new Error('Set S3_BUCKET, AWS_REGION, APIFY_RECORD_KEY, and S3_OBJECT_KEY');
}

const store = await KeyValueStore.open(storeName);
const image = await store.getValue(recordKey);
if (!image) throw new Error(`Screenshot record not found: ${recordKey}`);

const s3 = new S3Client({ region });
await s3.send(new PutObjectCommand({
  Bucket: bucket,
  Key: objectKey,
  Body: image,
  ContentType: 'image/png',
}));
console.log(`Uploaded s3://${bucket}/${objectKey}`);

The code expects the service to have access to the intended Apify store and AWS credentials. If your Actor reads a record through a different client or API, use that client to retrieve the bytes and keep the S3 PutObject request body as the image buffer. Preserve the actual content type, such as image/jpeg, when the source is JPEG.

One-off export with cURL

For a single record, download the record bytes to a local file using the Apify key-value-store record endpoint for your store, then upload that file with AWS CLI. The endpoint path includes your store and record key; use the store identifier and record key available from your Apify run or storage view. Authenticate the download with your Apify token without placing it in shell history where possible.

curl --fail --silent --show-error \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  "$APIFY_RECORD_URL" \
  -o screenshot.png

aws s3 cp screenshot.png "s3://$S3_BUCKET/$S3_OBJECT_KEY" \
  --content-type image/png

APIFY_RECORD_URL here is the record’s authorized download URL from your store context; it is a placeholder, not a fixed endpoint. If using a direct API URL, follow Apify’s current key-value-store API documentation for its route and authentication. The AWS CLI command uploads a local file; the S3 object remains private by default unless access is granted.

Python transfer example

For a standalone export service, read the binary response from an Apify record URL and pass its bytes to Boto3. Supply an authenticated record URL from your Apify integration and use credentials from the normal AWS credential chain.

import os
import requests
import boto3

record_url = os.environ["APIFY_RECORD_URL"]
bucket = os.environ["S3_BUCKET"]
object_key = os.environ["S3_OBJECT_KEY"]
content_type = os.getenv("SCREENSHOT_CONTENT_TYPE", "image/png")

response = requests.get(
    record_url,
    headers={"Authorization": f"Bearer {os.environ['APIFY_TOKEN']}"},
    timeout=(10, 90),
)
response.raise_for_status()
image_bytes = response.content
if not image_bytes:
    raise RuntimeError("The Apify record response was empty")

s3 = boto3.client("s3", region_name=os.getenv("AWS_REGION"))
s3.put_object(
    Bucket=bucket,
    Key=object_key,
    Body=image_bytes,
    ContentType=content_type,
)
print(f"Uploaded s3://{bucket}/{object_key}")

Install the dependencies with python -m pip install requests boto3. As with the cURL example, APIFY_RECORD_URL is supplied by your store workflow; consult the current Apify API documentation for how to construct it for the store and record you are using.

Design keys and metadata before batch export

An S3 key is the object name and path used for the upload. Pick a stable convention before exporting many records, for example:

screenshot-runs/{run-id}/{normalized-host}/{timestamp}.png
  • Include a run identifier when you need to group one Actor run’s output.
  • Normalize host names and remove or encode characters that your downstream tools may interpret specially.
  • Do not put secrets or sensitive query parameters from source URLs in object keys.
  • Decide whether repeated runs should overwrite a stable key or create a new timestamped object.
  • Keep a dataset manifest with source URL, Apify record key, S3 bucket and key, content type, capture time, and transfer status if downstream consumers need traceability.

Keep bucket access private unless your use case requires another access path. Grant only the permissions needed for the target bucket and prefix, protect both Apify and AWS credentials, and use a deliberate access mechanism for consumers. Confirm the object key and content type in S3, then check readability using the identity that is supposed to consume it.

Recurring and bulk exports

  1. Have the capture Actor write each screenshot to a named key-value store and emit a metadata row to a dataset.
  2. Run a transfer step that retrieves each record, maps it to a deterministic S3 key, and calls PutObject.
  3. Record success or failure per image in a manifest so a failed transfer can be retried without rerunning the capture.
  4. Retry transient network or service errors with bounded exponential backoff; do not retry permanent authorization or invalid-key errors without fixing their cause.
  5. Make retries idempotent by reusing the same destination key when the desired behavior is overwrite, or by assigning a unique run key when each capture should be retained.
  6. For dataset rows, the official Apify S3 Uploader may be appropriate. Verify its current actor inputs and resulting file layout, and use the binary bridge for screenshot records unless its current documentation explicitly confirms key-value-store support.

For large files or high-volume transfers, consider streaming the record to disk or through a stream instead of holding many images in memory. Keep concurrency bounded to avoid exhausting memory, sockets, or API quotas. This workflow’s exact throughput and cost depend on image size, volume, transfer retries, storage class, and AWS data access; no benchmark is assumed here.

Performance, reliability, and cost

  • Capture dominates time in many workflows: page load and screenshot rendering happen before S3 transfer. Reuse captured records for transfer retries instead of recapturing pages.
  • Bound concurrency: parallel uploads can reduce wall-clock time, but limit in-flight image buffers and tune to your Actor memory and network capacity.
  • Use timeouts and retries: separate record-read and S3-upload timeouts where possible. Retry transient failures with a cap; retain failure details for later recovery.
  • Keep the binary and metadata paths separate: this makes it possible to audit or reprocess metadata without changing image bytes.
  • Budget storage and requests: S3 storage, PUT requests, retrieval, and data transfer can contribute to cost according to your usage and AWS terms. Apify usage and storage are separate considerations. Review your account pricing and retention needs rather than assuming a fixed cost per screenshot.
  • Verify retention: define how long Apify records and S3 objects should remain, and confirm the applicable settings for each storage type.

Troubleshooting

Symptom Likely cause Fix
S3 object contains JSON or text instead of an image The transfer uploaded a metadata response or encoded representation rather than the binary record body. Read the record as bytes and pass the buffer directly as Body. Check the response content type and inspect the downloaded file.
Object opens as a download or has the wrong media type ContentType is missing or does not match the screenshot format. Set image/png or image/jpeg based on the stored image bytes.
AccessDenied from S3 The active AWS identity lacks permission for the bucket or target prefix, or a bucket policy denies the operation. Check the identity and bucket policy; grant only the required write permission for the intended destination.
Apify record not found Wrong store name, record key, or store/run context, or the record is no longer retained. Log the store and key at capture time, confirm the store exists, and check its configured retention.
Signature or authorization error reading Apify Missing, invalid, or incorrectly scoped Apify credentials, or an incorrect record URL. Use the correct token and current API route for the store record; keep tokens out of logs and source control.
Duplicate keys overwrite earlier screenshots The key convention is reused across runs. Add a run ID or timestamp if each capture must be preserved; retain a stable key only when overwrite is intentional.
Upload stalls or memory use grows Too many images are buffered or transferred concurrently, or network timeouts are unsuitable. Lower concurrency, set request timeouts, and stream or spool large images rather than retaining an entire batch in memory.

Or skip the browser setup

If your goal is to get clean page screenshots rather than preserve an existing Apify run, ScreenshotNeo provides a website screenshot API and MCP server. It returns PNG, JPEG, WebP, or PDF from one GET request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners are accepted like a visitor and removed along with supported newsletter popups and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents use tools such as take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. If you want those screenshots in S3, save the returned response bytes to an object using the upload pattern above. Sign up for 1,000 free screenshots a month, with no card.

FAQ

Can Apify upload screenshots directly to S3?

The official S3 Uploader is described as a dataset-to-S3 integration. For screenshot image records in a key-value store, use a transfer step that reads the bytes and uploads them to S3, unless current documentation explicitly adds direct key-value-store support.

Should I store screenshots in a dataset?

Use a key-value store for image bytes. Use a dataset for structured rows such as capture metadata or an export manifest.

Will S3 objects be publicly accessible after upload?

Objects are private by default unless access is granted. Choose and configure the access method required by the consumer.

Can I retry an upload without taking the screenshot again?

Yes, while the source record remains available. Keep the image in the Apify store until transfer success is recorded, and retry the same record with an intentional destination-key policy.