How to Save a Generated PDF to Amazon S3 with Ruby
Upload a generated PDF to Amazon S3 with Ruby using the AWS SDK, Tempfile-safe patterns, metadata, encryption, retries, and troubleshooting.
Use the AWS SDK for Ruby v3 and aws-sdk-s3. Once your application has written the PDF, upload it with Aws::S3::Object#upload_file when you have a path, or Object#put when you already have an open file or IO object. Set content_type: "application/pdf", use credentials supplied by the runtime, and keep the object private unless your access design explicitly requires otherwise.
This guide starts after PDF generation. Your PDF library or framework can produce the bytes; the S3 workflow is the same whether the source is a normal file, a Tempfile, or an in-memory stream.
Prerequisites
- An AWS account and an S3 bucket.
- Ruby and the AWS SDK for Ruby v3.
- Credentials available through the AWS SDK credential chain, such as an IAM role, environment variables, or your shared AWS configuration. Never hard-code access keys.
- Permission for the running identity to write to the target bucket and key.
gem install aws-sdk-s3
In a Bundler project, add gem "aws-sdk-s3" to your Gemfile and run bundle install. The official AWS documentation covers the Ruby S3 examples and API details: S3 examples for the AWS SDK for Ruby and the Aws::S3::Object API.
Upload a generated PDF from a file path
This is the clearest option when your PDF generator has already written a completed file to disk.
require "aws-sdk-s3"
bucket = ENV.fetch("S3_BUCKET")
key = "reports/generated-#{Time.now.utc.strftime('%Y%m%dT%H%M%SZ')}.pdf"
path = "/tmp/generated-report.pdf"
object = Aws::S3::Object.new(bucket, key)
object.upload_file(path, content_type: "application/pdf")
puts "Uploaded s3://#{bucket}/#{key}"
upload_file accepts a String path, Pathname, File, or Tempfile. The object key is the name under which S3 stores the PDF, so choose a collision-safe convention when several reports can have the same title.
Upload with Object#put
Use put when you want to make the file lifetime explicit or when the source is already open.
require "aws-sdk-s3"
bucket = ENV.fetch("S3_BUCKET")
key = "reports/generated.pdf"
path = "/tmp/generated-report.pdf"
object = Aws::S3::Object.new(bucket, key)
File.open(path, "rb") do |file|
object.put(
body: file,
content_type: "application/pdf"
)
end
puts "Uploaded s3://#{bucket}/#{key}"
Open the source in binary mode. The block closes the file after the request completes. The content_type option stores the conventional PDF media type so downloads and browser responses are handled correctly by clients.
Upload a Ruby Tempfile
PDF generators often return a Tempfile. Rewind it before uploading if the generator has written to it or you have already read from it.
require "aws-sdk-s3"
require "tempfile"
bucket = ENV.fetch("S3_BUCKET")
key = "reports/#{SecureRandom.uuid}.pdf"
object = Aws::S3::Object.new(bucket, key)
tempfile = Tempfile.new(["report-", ".pdf"])
begin
# Replace this with your PDF generator.
tempfile.binmode
tempfile.write(pdf_bytes)
tempfile.flush
tempfile.rewind
object.upload_file(tempfile, content_type: "application/pdf")
ensure
tempfile.close!
end
If the Tempfile is already complete, passing it to upload_file is sufficient. If you pass an open Tempfile, your code remains responsible for closing it. A completed path can also be supplied while the temporary file still exists.
Set encryption and other upload options
S3 supports server-side encryption options. Choose the option required by your bucket and account configuration; do not add an encryption value that your IAM policy or bucket settings do not permit.
object.put(
body: File.open("/tmp/generated-report.pdf", "rb"),
content_type: "application/pdf",
server_side_encryption: "AES256"
)
Prefer a block so the file is always closed:
File.open("/tmp/generated-report.pdf", "rb") do |file|
object.put(
body: file,
content_type: "application/pdf",
server_side_encryption: "AES256"
)
end
The AWS Ruby APIs expose additional request parameters. Add only values your application needs, such as metadata or a configured customer-managed encryption key. Keep authorization decisions in IAM, bucket policy, and your application rather than making the object public for convenience. See the S3 bucket API parameters.
Choose a safe object key
An S3 key is a string, not a real directory. Prefixes make listings readable, but uniqueness and authorization are application responsibilities.
user_id = Integer(ENV.fetch("USER_ID"))
report_id = ENV.fetch("REPORT_ID")
key = "private/users/#{user_id}/reports/#{report_id}.pdf"
- Include a report or request ID when overwriting would be dangerous.
- Use a deterministic key only when replacement is intentional.
- Do not put secrets, session tokens, or personal data in a key unless your data policy allows it.
- Simultaneous writes to one key do not preserve every version automatically. Enable bucket versioning or use unique keys when that matters.
Handle errors and confirm success
SDK calls raise AWS service errors for rejected requests and service-side failures. Rescue the error, log the bucket and key without credentials, and decide whether your job should retry or fail.
require "aws-sdk-s3"
begin
object.upload_file(path, content_type: "application/pdf")
rescue Aws::S3::Errors::ServiceError => e
warn "S3 upload failed for s3://#{bucket}/#{key}: #{e.class}: #{e.message}"
raise
end
A successful upload call means the SDK completed the request. For workflows that must verify metadata after upload, issue a head_object request and check the returned content type, size, or checksum according to your requirements.
Large PDFs and multipart uploads
The current AWS SDK for Ruby v3 Aws::S3::Object#upload_file documentation lists a default multipart threshold of 104,857,600 bytes (100 MiB). Files at or above that threshold use multipart upload behavior in that API. The threshold is configurable and can differ across abstractions or SDK versions, so verify the version and configuration used by your application. The TransferManager API documents multipart transfer and parallel part uploads.
For ordinary report PDFs, the path-based API is usually simplest. For very large files, tune multipart settings only after measuring memory, network, and retry behavior in your deployment.
Credentials and IAM
Let the SDK resolve credentials from the environment, shared AWS configuration, or the compute role attached to your workload. Grant the narrowest write permission needed for the bucket and prefix. A typical design permits writing to a specific prefix and denies public access, while a separate application endpoint authorizes downloads.
Do not paste access keys into source code, examples, logs, or PDF metadata. If an upload receives AccessDenied, check the identity actually used by the process, the bucket policy, object ownership settings, encryption-key permissions, and whether the requested key is within the allowed prefix.
Performance, reliability, and cost notes
- Network path: Upload time depends on PDF size, application location, and the route to the S3 region. Keep compute near the bucket when latency matters.
- Retries: Treat transient network and service failures as retryable with bounded backoff. Make the key idempotent when retrying must not create duplicate reports.
- File lifetime: Keep a path or open IO available until the SDK call returns. Delete a Tempfile only after the upload has completed.
- Memory: A file path or file body avoids loading the entire PDF into a Ruby String. Choose the source form that matches your generator and memory limits.
- Multipart: Large uploads can retry individual parts, but they also create multipart cleanup considerations if a process dies. Follow your SDK version’s multipart configuration and lifecycle guidance.
- S3 charges: Storage, requests, and data transfer are billed by AWS according to your account and region. This article does not estimate a price; check the current S3 pricing page for your workload.
Common errors and fixes
| Error or symptom | Likely cause | Fix |
|---|---|---|
LoadError: cannot load such file -- aws-sdk-s3 |
The gem is not installed or not in the bundle. | Add aws-sdk-s3 to the Gemfile, run Bundler, and require it. |
AccessDenied |
The runtime identity or bucket policy does not permit the requested write. | Inspect the resolved IAM identity, bucket policy, prefix, and encryption-key permissions. |
NoSuchBucket |
The bucket name is wrong or the request targets the wrong account or region. | Verify the exact bucket name, account, and configured region. |
| Uploaded PDF is empty | A Tempfile or IO cursor is at the end, or generation had not finished. | Flush the writer, call rewind, and upload only after generation completes. |
| PDF downloads as generic binary data | No content type was supplied. | Set content_type: "application/pdf". |
| Temporary-file error after upload starts | The file was closed or deleted before the SDK finished reading it. | Keep the file open and present through the call; clean it up afterward. |
| Retry creates duplicate objects | The key contains a new random value on every attempt. | Use a stable request or report ID for retries, or record the generated key before retrying. |
| Encryption request is rejected | The selected encryption mode conflicts with bucket or IAM configuration. | Use the encryption setting required by the bucket and grant the necessary key permissions. |
Or skip the browser setup
If your generated PDF starts as a webpage, ScreenshotNeo can return a clean capture through one GET request, including PDF output. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports its verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Use the API example from the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture, PDF controls, custom CSS and JavaScript, waits, blocking rules, headers, cookies, geolocation, caching, signed links, asynchronous jobs, bulk capture, and a usage API. Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Do I need a PDF-specific S3 method?
No. S3 stores bytes. Generate the PDF, then upload the resulting path or IO with the PDF content type.
Should I use upload_file or put?
Use upload_file for a completed path or Tempfile. Use put when you already manage an open IO and want to pass it as the request body.
Can I upload directly from memory?
Yes. Pass a string or IO body to put, while considering the memory cost of holding the complete PDF.
Will S3 make my PDF public?
No. Access follows your bucket policy, IAM configuration, and any application download endpoint. Keep objects private unless public delivery is an explicit requirement.
What happens if the same key is uploaded twice?
The later write can replace the object. Use unique keys or deliberate bucket versioning when retaining previous PDFs is required.


