ScreenshotNeo

BlogHow-to

Save a Generated PDF to Amazon S3 in Python

Generate a PDF as bytes, upload it to Amazon S3 with Boto3, and handle metadata, permissions, memory use, and common failures.

By the ScreenshotNeo team29 September 202611 min read

Save a Generated PDF to Amazon S3 in Python

To save a generated PDF to Amazon S3 without writing a temporary file, get the finished document as bytes, wrap it in io.BytesIO, rewind the stream with seek(0), and pass it to Boto3’s upload_fileobj. Set ContentType to application/pdf so consumers can identify the object correctly. If the PDF already exists on disk, use upload_file with its path instead.

This guide covers the in-memory upload, a runnable example that generates a small PDF with ReportLab, the path-based alternative, upload options, permissions, failure handling, and practical notes on memory and reliability. The PDF-generation library and the S3 upload are separate steps: any library that can produce valid PDF bytes can feed the same upload code.

1. Install Boto3 and configure AWS access

Install Boto3 and, for the complete example below, ReportLab:

python -m pip install boto3 reportlab

Configure credentials using an AWS-supported mechanism before running the program. Common choices include an IAM role attached to an AWS compute environment, environment-based credentials, or a local shared AWS profile. Prefer short-lived role credentials where available; do not put access keys in source code or commit them to version control.

The identity used by the process needs permission to write to the target bucket and key. In a typical least-privilege policy, grant s3:PutObject only for the required object path. Additional permissions may be needed for particular bucket configurations, such as encryption or access-control settings. Check the bucket’s policies and encryption requirements as well as the identity policy if S3 returns an access error.

2. Generate PDF bytes and upload them from memory

Here is a complete example. It generates a small PDF in memory, uploads it, and prints the bucket and key only after the upload call returns successfully.

The upload handoff: finish the PDF, rewind the binary stream, and send it to the S3 object key.
The upload handoff: finish the PDF, rewind the binary stream, and send it to the S3 object key.
from io import BytesIO
import os

import boto3
from reportlab.pdfgen import canvas


def make_pdf_bytes() -> bytes:
    """Generate a minimal one-page PDF and return its bytes."""
    output = BytesIO()
    pdf = canvas.Canvas(output, pagesize=(612, 792))  # 8.5 x 11 inches at 72 points/inch
    pdf.setTitle("Generated report")
    pdf.drawString(72, 720, "Generated report")
    pdf.drawString(72, 690, "This PDF was generated in memory and uploaded to S3.")
    pdf.showPage()
    pdf.save()  # Finish the PDF before reading its bytes.
    return output.getvalue()


def upload_pdf_bytes(pdf_bytes: bytes, bucket: str, key: str) -> None:
    if not pdf_bytes:
        raise ValueError("PDF content is empty")

    stream = BytesIO(pdf_bytes)
    stream.seek(0)
    s3 = boto3.client("s3")
    s3.upload_fileobj(
        stream,
        bucket,
        key,
        ExtraArgs={"ContentType": "application/pdf"},
    )


if __name__ == "__main__":
    bucket = os.environ["PDF_BUCKET"]
    key = "reports/generated-report.pdf"
    content = make_pdf_bytes()
    upload_pdf_bytes(content, bucket, key)
    print(f"Uploaded s3://{bucket}/{key}")

Set the destination bucket in the environment and run the script:

export PDF_BUCKET=my-pdf-bucket
python upload_report.py

The important handoff is the call to pdf.save() before getvalue(). A PDF generator may buffer output until it is finalized; uploading before that point can result in incomplete or invalid content. Once generation is complete, getvalue() returns the document bytes. BytesIO presents those bytes as a binary file-like object, which is what upload_fileobj expects. AWS documents this method for readable file-like objects and specifies that the object must be in binary mode and return bytes.[AWS Boto3: upload_fileobj]

Why call seek(0)?

File-like objects have a current position. When code has read from or written to a stream, its position may be at the end. An uploader reading from that position would see no remaining bytes. Calling seek(0) resets the position to the start before upload. In the example, a newly created BytesIO(pdf_bytes) already starts at position zero, but the explicit rewind makes the upload function safe to adapt when it receives a stream that another step has already used.

3. Upload a PDF you have already saved to disk

If the generator already wrote the finished document to a local file, use Boto3’s path-based upload_file. It accepts the filename, bucket, and object key, plus optional transfer arguments. This avoids reading the entire file into a Python bytes value first.

Choose the upload helper based on the input: a file path for upload_file or a binary stream for upload_fileobj.
Choose the upload helper based on the input: a file path for upload_file or a binary stream for upload_fileobj.
import boto3

s3 = boto3.client("s3")
s3.upload_file(
    "./generated-report.pdf",
    "my-pdf-bucket",
    "reports/generated-report.pdf",
    ExtraArgs={"ContentType": "application/pdf"},
)
print("Uploaded s3://my-pdf-bucket/reports/generated-report.pdf")

AWS distinguishes the two managed transfer helpers by input: upload_file is for a filename, while upload_fileobj is for a readable file-like object.[AWS Boto3: uploading files]

Approach Use it when Trade-off
upload_fileobj(BytesIO(pdf_bytes), ...) The PDF is already bytes or is generated in memory. Simple handoff with no temporary file, but the document occupies memory.
upload_file(path, ...) The finished PDF is already on disk. Path-oriented and avoids making a second full in-memory copy.
Open a file in binary mode and call upload_fileobj You want a stream-oriented interface for an existing file. Keep the file open through the upload and open it with rb.

4. Set object metadata and transfer options

The managed transfer helpers accept ExtraArgs for supported S3 object settings. The most important setting for a PDF is often ContentType. Without it, a consumer may receive a generic content type instead of application/pdf. Additional supported values can include metadata and encryption-related settings, subject to the helper’s supported-argument list and the bucket’s policies. Check Boto3’s current documentation before passing less common arguments.[AWS Boto3: upload_fileobj parameters]

extra_args = {
    "ContentType": "application/pdf",
    "Metadata": {
        "document-kind": "monthly-report",
        "source": "report-generator",
    },
}

s3.upload_fileobj(stream, bucket, key, ExtraArgs=extra_args)

Metadata values are strings. Avoid placing secrets or personal data in object metadata: metadata may be visible to principals that can inspect the object. If you need a transfer progress callback, supply Callback; Boto3 calls it with transferred byte counts as the transfer progresses. A Boto3 transfer Config can tune transfer behavior such as multipart thresholds and concurrency. The managed transfer can use multipart upload and multiple threads when appropriate, so an application generally should not reimplement multipart handling for an ordinary upload.[AWS Boto3: managed transfers]

Progress callback example

class Progress:
    def __init__(self):
        self.transferred = 0

    def __call__(self, bytes_transferred):
        self.transferred += bytes_transferred
        print(f"Transferred {self.transferred} bytes", flush=True)

progress = Progress()
s3.upload_fileobj(
    stream,
    bucket,
    key,
    ExtraArgs={"ContentType": "application/pdf"},
    Callback=progress,
)

5. Choose a stable bucket and object key

The S3 object key is the path-like name within the bucket. Use a deterministic, appropriately scoped key if the same logical report should replace its prior version; use a unique identifier or timestamp when each generated report must be retained separately. Remember that a repeated upload to the same bucket and key addresses the same object name, so choose the overwrite behavior deliberately and apply bucket versioning or application-level retention according to your requirements.

Do not treat a key as a local filesystem path. Forward slashes are commonly used to group objects in the console, but S3 stores the complete key as an object name. Validate user-provided key components and avoid allowing untrusted input to choose arbitrary keys. If the object is private, return an application URL or a time-limited presigned URL according to your access model rather than making the object public by default.

6. Handle errors and only report success after upload

In production, catch the specific Boto3 or botocore exceptions your application can handle, log enough context to diagnose failures, and avoid logging credentials or sensitive document contents. A successful return from upload_fileobj is the point at which this simple workflow should report completion. If your application also records a database row, consider how it will reconcile a successful S3 write followed by a failed database update.

import logging

import boto3
from botocore.exceptions import BotoCoreError, ClientError

logger = logging.getLogger(__name__)


def safe_upload(pdf_bytes: bytes, bucket: str, key: str) -> None:
    if not pdf_bytes:
        raise ValueError("Refusing to upload an empty PDF")

    stream = BytesIO(pdf_bytes)
    stream.seek(0)
    try:
        boto3.client("s3").upload_fileobj(
            stream,
            bucket,
            key,
            ExtraArgs={"ContentType": "application/pdf"},
        )
    except ClientError:
        logger.exception("S3 rejected PDF upload to s3://%s/%s", bucket, key)
        raise
    except BotoCoreError:
        logger.exception("Boto3 could not complete PDF upload to s3://%s/%s", bucket, key)
        raise

Do not automatically retry every exception without considering the operation and the failure. Transient network or service failures may justify bounded retries with backoff; a permissions error or invalid bucket name will not be fixed by repeating the same request. For applications that retry, make sure the chosen key and overwrite semantics are acceptable if an earlier attempt may have completed even though its response was lost.

7. Memory, performance, reliability, and cost notes

  • Memory: Generating a PDF into BytesIO and calling getvalue() can keep the document in memory, and converting between stream and bytes can create additional copies. This is convenient for small and moderate documents. For large or high-concurrency jobs, generate to a temporary file or use a generator that writes to a suitable file-like stream, then upload from that stream. Set application limits so concurrent document generation cannot exhaust worker memory.
  • Upload behavior: Boto3’s managed transfer can use multipart upload and multiple threads when needed. A transfer configuration can adjust thresholds and concurrency; tune only after measuring your own document sizes and runtime environment. More concurrency can increase resource use.
  • Reliability: Keep the input stream open until the upload call returns. Do not reuse or close it from another thread during transfer. Preserve the bucket and key as a job result only after success, and record enough identifiers to retry or investigate a failed job.
  • Validation: A non-empty byte string alone does not prove that the content is a valid PDF. Ensure the generation library has finalized the document, and where correctness matters validate the output before upload or verify it after retrieval in a separate workflow.
  • Cost: This implementation has no universal price figure: actual AWS charges depend on the account’s current S3 pricing, region, storage class, request volume, data transfer, and retention choices. Consult AWS pricing for the bucket’s region and expected usage. Avoid inventing a per-upload estimate without those inputs.

8. Troubleshooting common failures

Symptom Likely cause What to check or change
AccessDenied The caller or bucket policy does not allow writing this object, or required encryption settings are missing. Check the active identity, bucket policy, object ARN scope, and bucket encryption requirements. Grant only the required permissions.
NoCredentialsError or credential lookup failure Boto3 cannot find usable credentials. Configure the runtime’s IAM role, profile, or supported environment credentials; confirm the process uses the expected profile or role.
NoSuchBucket or bucket-not-found response The bucket name is wrong, unavailable to this account, or the request targets the wrong region/account. Verify the exact bucket name, AWS account, and client region configuration.
The uploaded object is empty or truncated The stream was already at EOF, the PDF generator had not finalized output, or the stream was closed too soon. Finalize the PDF, check that bytes exist, call seek(0), and keep the stream open until completion.
The object downloads as generic binary data Content type was not set. Pass ExtraArgs={"ContentType": "application/pdf"}.
ValueError for unsupported extra arguments An argument is not accepted by the managed transfer helper, or its spelling/capitalization is wrong. Compare the argument with the current Boto3 reference for upload_fileobj; only supported values belong in ExtraArgs.
Upload stalls or fails on large output Network interruption, worker resource limits, transfer configuration, or a policy timeout. Inspect logs and network path, keep retries bounded, review transfer concurrency and timeouts, and consider file-based generation for memory pressure.
PDF opens as corrupt The generator output was read before finalization, bytes were altered, or the wrong buffer was uploaded. Call the library’s finalization method before extracting bytes; validate the local/in-memory PDF independently of S3.

9. When the PDF starts as a website screenshot

If the PDF you need is a capture of a web page, you can build a browser capture pipeline yourself and then upload the resulting PDF bytes using the same Boto3 pattern above. That approach gives control over browser setup, page readiness, and rendering, but you need to manage those pieces in your application.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. For a page-to-PDF capture, request the target URL from its API and save the returned PDF response. See the ScreenshotNeo API documentation for request options. The example below is the documented one-call screenshot request; select PDF output using the API’s supported format option when adapting it to a PDF capture.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server gives AI agents tools named take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.

10. FAQ

Can I upload a PDF from memory without a temporary file?

Yes. Pass the generated bytes through BytesIO to upload_fileobj. This is the standard in-memory handoff; account for the document’s memory footprint when handling large files or many concurrent jobs.

Should I use upload_file or upload_fileobj?

Use upload_file when your input is a filesystem path. Use upload_fileobj when your input is a readable binary stream, including BytesIO or an open file. The choice follows the form of your input.

Do I need to set ContentType?

Set it when downstream clients should recognize the object as a PDF. Use application/pdf; it is object metadata, not a transformation of the uploaded bytes.

Can I use a different PDF generator?

Yes. The generator only needs to produce a finished PDF byte sequence or a binary file-like object. The S3 upload step does not depend on ReportLab.

Does upload_fileobj return a public URL?

No public URL is produced by the upload helper itself. The upload identifies the object by bucket and key. Access depends on your bucket and application policy; keep private documents private and generate access links through your application’s chosen mechanism.