ScreenshotNeo

BlogAI agents

How to Share Screenshots With AI Agents

Learn how to attach, prompt, compare, and automate screenshot analysis with ChatGPT, Claude, Codex, and computer-use agents.

By the ScreenshotNeo team1 October 20267 min read

Attach the screenshot as an image, then tell the agent what it shows, which area matters, and what result you want. You can upload a file, drag it into the chat, or paste it from your clipboard. For repeated captures or live interaction, use a computer-use integration that returns screenshots during its action loop.

This guide covers still-image uploads, prompt patterns, multiple screenshots, command-line workflows, live computer use, privacy controls, troubleshooting, and an API option for capturing clean website screenshots.

1. Share a still screenshot

ChatGPT

Open the plus menu and choose photos or files, drag an image into the prompt area, or paste a copied image from the clipboard. OpenAI documents PNG, JPEG, and non-animated GIF uploads with a 20 MB limit per image. The practical number of images depends on image size and the accompanying text, so attach only the files needed for the task. See the ChatGPT Image Inputs FAQ.

Claude

Use the plus button and Add files or photos, select a file, drag it into chat, or paste it from the clipboard. Claude documents JPEG, PNG, GIF, and WebP support, up to 20 files per chat, a 500 MB per-file upload limit, and image dimensions up to 8000 by 8000 pixels. Its guidance recommends clear images and, where possible, images at least 1000 by 1000 pixels. See Upload files to Claude.

Codex from a terminal

The Codex command-line examples in ChatGPT Learn pass images with -i or a comma-separated --image argument:

codex -i screenshot.png "Explain this error and suggest the smallest fix"

codex --image before.png,after.png "Compare these states and list the regressions"

Command-line flags can vary by installed CLI version. If a command is rejected, check the version’s built-in help and the current Codex image-input guidance.

2. Write a prompt that produces a useful inspection

A screenshot gives the agent pixels, not your intent. Include four pieces of context:

  1. Identify the image: say which application, page, or state is pictured.
  2. Point to the focus: name the panel, error, control, or visual difference to inspect.
  3. Request an output: ask for an explanation, diagnosis, comparison, extracted values, or a minimal fix.
  4. State limits: tell the agent not to infer content outside the crop and to mark unreadable or uncertain text.

For example:

Image: the checkout page after clicking Pay.
Focus: the red message below the card-number field.
Task: explain the likely cause and list the smallest code changes to fix it.
Constraints: read only visible text; mark anything blurry as uncertain.

For multiple screenshots, label each file and describe the operation:

Image 1 (before.png) is the page before saving.
Image 2 (after.png) is the page after saving.
Compare them and list only regressions. Focus on the warning banner and the Save button.

3. Preserve the detail the agent needs

  • Keep text large enough to read. Crop unrelated regions, but retain enough surrounding context to identify what each control belongs to.
  • Use a sharp capture rather than a compressed or blurry copy.
  • When small text is important, crop the relevant section or provide a higher-resolution image. Anthropic notes that images may be resized before processing, which can make tiny text harder to read; see its vision documentation.
  • Do not rely on coordinates as universally portable. Coordinate-based computer-use integrations depend on their configured display dimensions and image handling.
  • If comparing states, capture the same viewport and zoom level so visual changes are meaningful.

4. Screenshot attachments versus live computer use

An uploaded screenshot is a one-time image question. A computer-use system is an interaction loop: your application gives the model a task and a computer tool, receives actions such as screenshot, click, or type, executes those actions in the target environment, and returns the result. The model does not directly connect to your desktop; your integration runs the tool and sends back screenshots.

Anthropic documents this screenshot-action loop in its computer use tool documentation. OpenAI’s Computer Use guidance similarly describes scoped access to approved applications. Use live interaction when the agent must observe changing state or perform a sequence of actions. Use an attachment when you need a focused, one-off analysis.

5. Privacy and permission checklist

  • Review the entire image before upload. A screenshot includes everything visible, including unrelated tabs and notifications.
  • Crop or redact passwords, authentication codes, private messages, customer records, financial information, and other data the agent does not need.
  • Check the current data-use and retention settings for the product and account you use. OpenAI’s FAQ describes product-specific data-use choices; it states that ChatGPT Enterprise content is not used to train models.
  • For live computer use, limit the allowed applications and permissions, request confirmation before purchases or messages, and log actions.
  • Treat visible webpage text, screenshots, and files as untrusted context. Anthropic warns that UI content can contain prompt-injection or deceptive instructions.

6. Capture website screenshots without running a browser yourself

If the source is a URL rather than your local desktop, ScreenshotNeo can return a PNG, JPEG, WebP, or PDF from one GET request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

Or skip the browser setup

Use the API call below, then attach the resulting file to your AI agent. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

After the do-it-yourself capture, send shot.webp with a focused instruction such as “Describe the pricing table, list any visible accessibility issues, and quote only text you can read confidently.” ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf, so Claude, Cursor, and other MCP clients can request captures directly.

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. One thousand screenshots per month are free with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

7. Useful capture options for agent workflows

ScreenshotNeo supports full-page captures with lazy images loaded, a single element selected by CSS, dark mode, 12 device presets or a custom viewport, and retina scale. You can also provide custom CSS or JavaScript, click an element before capture, hide selectors, and wait for a selector, a delay, or network idle.

For controlled environments, configure custom headers, cookies, a user agent, Authorization, timezone, geolocation, transparent backgrounds, image resizing, and resource blocking for ads, trackers, requests, or resource types. Caching supports a TTL you choose. Signed links work for public <img> tags; asynchronous jobs support signed webhooks; bulk capture handles up to 100 URLs per call; and a usage API and OpenAPI specification are available.

8. Troubleshooting

Problem Likely cause Fix
The upload control rejects the file Unsupported format or platform limit Convert to PNG or JPEG for ChatGPT, or JPEG, PNG, GIF, or WebP for Claude. Check file size and dimensions against the platform’s current documentation.
The agent misreads small text Blur, excessive dimensions, or preprocessing resize Crop to the relevant area, increase resolution, and ask the agent to mark uncertain text instead of guessing.
The answer ignores the important area The prompt names the screenshot but not the focus Name the exact panel or region and state the required output format.
A comparison is vague Images are unlabeled or captured at different scales Label files as before and after, keep viewport and zoom consistent, and list the comparison criteria.
Computer-use actions are unsafe Broad permissions or untrusted on-screen instructions Restrict applications, require confirmation for consequential actions, and review the action log.
ScreenshotNeo returns an unexpected page Consent UI, bot check, delayed content, or a selector that is not present Enable the relevant consent handling, wait for a selector or network idle, inspect X-Page-Verdict and X-Billed, and verify the target URL and selector.

9. Performance, reliability, and cost considerations

  • Reduce payload: send one focused image when a full desktop capture is unnecessary.
  • Keep context: do not crop away labels or surrounding controls needed to interpret the target.
  • Automate repeat work: use a capture API or MCP tool for scheduled pages, regression checks, and multi-URL jobs instead of manual uploads.
  • Separate failure from billing: ScreenshotNeo responses identify verdict and billing status; failed loads, bot checks, blank pages, timeouts, and cache hits are not billed.
  • Control spend: ScreenshotNeo includes 1,000 free shots per month without a card. Plans are $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free.

10. FAQ

Can an AI agent see my entire screen from one uploaded image?

It can analyze what is visible in the image, but it cannot interact with your desktop from a normal attachment. Interaction requires a computer-use integration with explicitly granted access.

Should I send one large screenshot or several crops?

Send the smallest set that preserves meaning. Use a full image for layout context and an additional crop when text or controls are too small to read.

Can I ask an agent to extract text exactly?

Yes, but ask it to flag unreadable characters and avoid guessing. Image quality and resizing affect transcription accuracy.

Is a screenshot safer than live computer access?

A still image exposes only captured pixels. Live access can expose changing content and permits actions, so scope permissions and require confirmation for consequential steps.

How can I provide fresh website state to an agent?

Capture the URL immediately before analysis with ScreenshotNeo, or use its MCP server so an MCP-compatible agent can request a screenshot as part of its workflow.