How to Send Website Data Directly to Amazon S3 With Browse AI
Connect Browse AI to an existing Amazon S3 bucket, export robot Table data as CSV or JSON, and understand file layout, scheduled exports, and common setup errors.
To send website data directly to Amazon S3 with Browse AI, first connect Browse AI to an existing S3 bucket using an AWS CloudFormation stack and an IAM role. Then open the robot’s Tables, choose the data view you want, and export it as CSV or JSON to S3. Browse AI first extracts website data into a robot Table; an export sends that table data to your bucket. It does not automatically send every robot run to S3 by default.
The setup has four parts: prepare AWS and Browse AI, create the AWS integration resources, connect the role in Browse AI, and export a Table. The labels in the AWS and Browse AI interfaces can change, so check the current interface as you follow the steps. See Browse AI’s AWS S3 Integration Guide and S3 export instructions.
Prerequisites
- An AWS account with permission to create CloudFormation stacks and IAM roles.
- An existing S3 bucket. Note its exact name and AWS region.
- Optionally, a bucket folder path (prefix) if you want the integration limited to a particular location.
- A Browse AI robot with extracted data in Tables for the export itself.
- A random 20-character external ID made from letters and numbers. Use the same value in CloudFormation and Browse AI.
The Browse AI guide describes the CloudFormation resource as creating a role with the minimum permissions needed for export. Deleting the stack revokes those permissions. Review the template and the permissions presented by AWS before submitting the stack.
Connect Browse AI to your S3 bucket
- Open CloudFormation. In the AWS console, create a stack using Browse AI’s template:
https://browse-ai-integration.s3.us-east-1.amazonaws.com/s3-integration.yaml. - Enter the stack details. Choose a descriptive stack name, provide the exact S3 bucket name, enter your 20-character external ID, and add the optional bucket path if needed.
- Review and create the stack. Confirm the details and acknowledge the IAM resource creation prompt if AWS presents it. Wait for stack creation to complete.
- Copy the role ARN. Open the stack’s Resources tab or Stack info and find
S3AccessRole. Copy its IAM Role ARN. - Add the integration in Browse AI. Open the relevant robot’s Integrate tab, choose AWS, and add an S3 bucket connection. Provide the bucket region, role ARN, the same external ID, bucket name, and optional path, then save.
Keep the bucket name, region, external ID, and path consistent between the stack and Browse AI connection. The role ARN comes from the completed stack; it is not the bucket ARN.
Export a Table as CSV or JSON
- Open the robot’s Tables and select the tab or view you want to export.
- Set the filters and visible columns to match the intended dataset. Decide whether the export should contain the latest data or historical data where that option is available.
- Choose Export, then select CSV to S3 or JSON to S3.
- Start the export and follow its progress in Tables.
This route sends the data directly to the configured bucket without a local download. Browse AI documents both formats and the Table view settings that affect what is included. See How to export data from Tables.
Choose the format and scope
| Choice | What to consider |
|---|---|
| CSV or JSON | Choose the format your downstream tools expect. The documentation offers both and does not claim one is universally preferable. |
| Table tab and filters | The selected tab and active filters determine which records are in the export. |
| Visible columns | Set the visible columns to the fields you want included. |
| Latest or historical data | Check the historical-data setting if you need records beyond the latest dataset. |
| Combined or per-record output | Use the separate-record option when you need individual JSON files for records, if that option is enabled for the export. |
Understand the S3 file layout
Each export is placed in its own directory, named with a timestamp and unique export ID. Browse AI documents a pattern like export_{ISO 8601 timestamp}_{unique export ID}. Files are organized under that directory by Table tab. If you selected a bucket path, the export directory is beneath that prefix.
- A smaller export may appear as one CSV or JSON file per tab.
- Large exports can be split into numbered parts. The guide uses a file smaller than 100 MB as an example of a small file; treat that as an illustration, not a universal size guarantee.
- When separate-record export is configured, JSON filenames use record IDs.
If an export seems missing, look for the timestamped export directory under the configured bucket and prefix, then inspect its tab folders and any numbered chunks.
Set up recurring exports
Scheduled S3 exports are a separate feature from a one-time Table export. Browse AI’s guide describes scheduled exports as beta and says to contact support to unlock the feature. The documented prerequisites are an approved robot with extracted data and a configured S3 integration. Availability and cadence can depend on what is enabled for your account; confirm them in the account interface or current documentation before designing a production schedule. See the scheduled S3 export guide.
Troubleshooting
| Symptom | Likely cause | What to check |
|---|---|---|
| Browse AI cannot connect to the bucket | The bucket name or region is incorrect, or the bucket does not exist. | Confirm the exact bucket name and region in both AWS and the Browse AI integration. Verify that the bucket already exists. |
| Role access is denied | The AWS account lacks permission to create the required CloudFormation or IAM resources, or the generated role was changed. | Confirm your AWS permissions and inspect whether the stack-created role has been manually modified. |
| External ID verification fails | The external ID entered in Browse AI differs from the value used to create the stack. | Compare the values and enter the same 20-character string in both places. |
| Cannot find the role ARN | The stack may have completed but the role is not obvious in the console. | Open the CloudFormation stack’s Resources tab or Stack info and locate S3AccessRole. |
| Exported fields or records are missing | The active Table view does not match the intended dataset. | Review the selected tab, filters, visible columns, and historical-data setting before exporting again. |
| Export appears incomplete or split | Large exports may be chunked into multiple numbered files. | Inspect the complete timestamped export directory and all part files for each tab. |
| Scheduled export controls are unavailable | The feature is documented as beta and may require account access. | Check the scheduled-export guide and contact Browse AI support to ask whether it can be unlocked for your account. |
Performance, reliability, and cost considerations
- Export only the needed view. Filters and visible columns define the data being sent. Narrowing the dataset can make downstream processing simpler.
- Plan for multiple objects. A single export can create directories, per-tab files, and numbered chunks. Have downstream jobs enumerate the export directory instead of assuming one fixed filename.
- Use export IDs to distinguish runs. Timestamped directories with unique IDs keep exports separate and make it easier to identify a particular run.
- Account for AWS storage and requests. Browse AI’s cited guides explain the export workflow but do not specify AWS storage or request charges. Check your AWS pricing and retention requirements for your bucket and region.
- Do not assume automatic delivery on every robot run. A manual export is an explicit action; recurring delivery requires the separate scheduled-export feature and its beta access.
Or skip the browser setup
If your immediate job is to capture a website as an image or PDF rather than export a Browse AI Table, ScreenshotNeo is a website screenshot API and MCP server. It does not replace Browse AI’s Table-to-S3 workflow; it is an alternative for screenshot capture.
One GET request captures a URL. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.
FAQ
Does Browse AI send each robot run to S3 automatically?
No. The documented workflow exports data from Tables. Scheduled exports are a separate beta feature with an access requirement.
Can I export without downloading a file first?
Yes. After configuring the integration, choose CSV to S3 or JSON to S3 from the Table export flow.
Do I need to create the bucket first?
Yes. The integration setup expects an existing bucket.
Where should I look for an export in S3?
Look under the configured bucket and optional prefix for a timestamped directory with a unique export ID, then inspect its tab folders.


