ScreenshotNeo

BlogAI agents

Computer Use: Build a Safe AI Agent That Operates a Browser or Desktop

Learn how computer use works, how to build the runtime loop around it, and how to keep browser or desktop agents reliable and safe.

By the ScreenshotNeo team29 September 20269 min read

Computer Use: Build a Safe AI Agent That Operates a Browser or Desktop

Computer use lets an AI model operate browser and desktop interfaces by looking at screenshots and choosing actions. To build it, your application must provide the browser or desktop runtime, send observations to the model, execute the model’s requested actions, preserve state between calls, and verify the final result. The model does not independently open a window on your server: your code controls the environment and decides which actions are allowed.

Computer use is useful when a person-facing interface is the only practical way to work with an application, including some legacy software. If a supported API or structured integration can do the job, compare it first; GUI automation is not a universal replacement for APIs. OpenAI describes the capability as a model operating browser and desktop interfaces, with examples such as filling forms and testing user flows (Computer use guide, API documentation).

1. What computer use means

A computer-use agent is a loop between a model and an environment. Your application sends a task and current observations. The model responds with either code to run in a controlled environment or structured input actions. Your runtime carries those actions out and returns a fresh observation. This repeats until the task is complete, stopped, or needs a human.

The environment may be a browser or a desktop session. It is yours to provision and constrain. A model response is a proposal for what should happen next; it is not proof that the page changed or the task succeeded.

Part What it does
Model Interprets the task and observations, then proposes the next step.
Runtime Owns the browser or desktop session and carries out permitted actions.
Observation Screenshot and other results returned to the model so it can choose what to do next.
Policy layer Limits websites, actions, run duration, and actions requiring confirmation.
Verifier Checks the application’s actual end state against the requested outcome.

2. Choose an integration pattern

Code execution

The model writes code that uses a browser or desktop library such as Playwright or PyAutoGUI. Your application executes that code inside an isolated runtime and returns the resulting observations. The current OpenAI guide recommends this approach for GPT-6 Astra. It can be a good fit when a task involves a sequence of higher-level operations in an environment you control.

Structured computer actions

The model returns structured mouse and keyboard actions. Your application translates them into input in the browser or desktop environment. This approach makes the action boundary explicit: inspect each returned action, enforce your policy, execute it, and report the result. OpenAI’s guide continues to support this as an alternative.

Choose based on the runtime, task, and control requirements. Compare browser-only versus desktop scope, how you will isolate sessions, how observations are handled, and whether you can require confirmation before consequential actions. Do not choose based on one headline benchmark: task complexity and the evaluation environment affect reported performance.

3. Build the runtime loop

  1. Start an isolated session. Create a browser or desktop environment for the task. Apply a narrow site allowlist and action policy before the model receives control.
  2. Send the task and current state. Provide an initial screenshot or other useful observation. State the goal and any constraints clearly.
  3. Get the next action. The model may request code execution or return structured input actions. Your application decides whether the request is permitted.
  4. Execute and observe. Run the approved action in the same session, capture a new observation, and send it back. Short action groups make it easier to detect a wrong turn.
  5. Check the actual result. Verify the application state independently. Stop on success, a policy boundary, a timeout, a cancellation request, or a need for human confirmation.

Keep the browser or desktop session alive across model calls. Preserve the tool calls and outputs in the conversation as required by the integration. Continuing a model response does not restore a lost browser login, page state, or runtime variables. When state is uncertain, send a current screenshot rather than asking the model to infer what is on screen.

A computer-use agent depends on a controlled runtime that executes actions and returns observations.
A computer-use agent depends on a controlled runtime that executes actions and returns observations.

Coordinate mapping after resizing

If you resize screenshots before sending them, keep both the displayed dimensions and the environment dimensions. Convert returned screenshot coordinates back to runtime coordinates before clicking. For a simple proportional resize, if the original is W × H and the sent image is w × h, use x_runtime = x_image × W / w and y_runtime = y_image × H / h. Account for any crop, letterboxing, or device scaling too; a proportional formula alone will not correct those transformations.

4. Minimal loop pseudocode

The exact request fields and action schema depend on the selected model and computer-use interface. Keep the orchestration boundary explicit. The following language-neutral pseudocode shows the runtime responsibilities without pretending to define a provider request format:

session = isolated_runtime.start(allowed_sites=["app.example.test"])
conversation = [task_message("Find the latest invoice total")]

for step in range(MAX_STEPS):
    observation = session.screenshot()
    result = model.respond(conversation, observation)
    conversation.append(result.model_message)

    if result.is_final:
        break

    for action in result.requested_actions:
        if action_is_sensitive(action):
            require_user_confirmation(action)
        if not policy_allows(action, session.current_url):
            stop("Action blocked by policy")
        session.execute(action)
        conversation.append(tool_result(session.screenshot()))

verify_application_state(session, expected_outcome)
session.close()

Use the provider’s current computer-use guide for the request format, supported actions, model requirements, and continuation behavior. Do not copy an old example’s tool schema into a new integration without checking it against current documentation.

5. Safety and access controls

  • Isolate the runtime. Use a dedicated browser profile or VM rather than an employee’s general-purpose session.
  • Limit destinations and actions. Allow only the sites and operations needed for the task. Restrict downloads, navigation, and access to unrelated services.
  • Treat screen content as untrusted. A web page, document, or tool result can contain instructions that attempt to redirect the agent. Treat them as data, not as policy overrides.
  • Require confirmation for consequential steps. Purchases, external data transmission, destructive changes, and other hard-to-reverse actions need explicit user control. Typing sensitive information into a form is also data transmission.
  • Bound execution. Set a step limit, wall-clock deadline, and cancellation path. Stop when the task leaves its permitted scope.
  • Verify completion. Check the resulting record or page state instead of relying only on the model’s final message.

These controls follow the safety guidance in the OpenAI computer-use guide. They are part of the product design, not optional polish: the model sees untrusted content and can take actions with real effects.

6. Screenshots: capture and inspect

A screenshot can establish what the agent can currently see, help diagnose a failed interaction, or serve as evidence for a visual check. For your own application, capture directly from the controlled browser or desktop runtime when possible. Decide whether full-page capture is useful: it may include content outside the visible viewport, while a viewport capture better reflects what a user can interact with at that moment.

A screenshot service can remove common overlays before returning a capture.
A screenshot service can remove common overlays before returning a capture.

If you need a screenshot of a public web page outside the agent’s live session, ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. Its API accepts familiar parameter names used by other screenshot APIs, which can make migration simpler. See the ScreenshotNeo API documentation for request options.

Or skip the browser setup

For a one-off screenshot of a public URL, a capture API avoids provisioning a browser just to render an image. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo has full-page capture with lazy images loaded, element capture by CSS selector, dark mode, device presets and custom viewports, retina scale, PDF options, HTML/CSS capture, custom CSS and JavaScript, click-before-capture, hide selectors, wait conditions, request and resource blocking, custom headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable cache TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification. Each response identifies whether the page was clean and whether it was billed through X-Page-Verdict and X-Billed headers; individual cleaning steps can be turned off.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

7. Performance, reliability, and cost

Computer-use tasks can require repeated model calls, screenshots, and browser actions. Keep the loop efficient by grouping a few predictable actions, then returning an observation so the model can check progress. Avoid long blind sequences: they may save a round trip but make it harder to recover after a changed layout or unexpected dialog. Capture a fresh screenshot after navigation, a modal, a failed action, or any point where the state may have changed.

Reliability depends on page stability, session persistence, coordinate mapping, policy boundaries, and verification. Add deadlines and cancellation because pages can hang or the agent can repeat a step. For critical flows, verify with an application state check, such as whether the expected record exists, rather than treating a visually plausible page as proof.

Cost depends on the model and the work required by the task; the dossier does not establish a universal price or expected number of calls. Measure representative tasks in your own environment: record model usage, number of screenshots, retries, runtime duration, and completion rate. The 2025 launch post reported 38.1% success on OSWorld, 58.1% on WebArena, and 87.0% on WebVoyager for that system at that time. Those are task-specific historical benchmark results, not present-day guarantees or a general reliability score (OpenAI’s 2025 launch post).

8. Troubleshooting

Symptom Likely cause Fix
Clicks consistently miss Screenshot was resized, cropped, or displayed at a different scale. Map coordinates to the runtime’s dimensions and include crop offsets or scaling factors.
The next call starts at the wrong page The browser session was recreated, or its state was not preserved. Keep one runtime alive across calls and confirm its URL and screenshot before acting.
The agent repeats an action It did not receive a useful observation after the action, or the page did not change. Return a fresh screenshot and relevant result after each short action group; add a step limit.
Task reports success but nothing changed The final message was accepted as evidence without checking the application. Verify the expected record or state independently before reporting completion.
Unexpected instruction appears on screen Page content is attempting to influence the agent. Treat it as untrusted data, enforce the original policy, and stop if the request exceeds scope.
Automation submits or deletes unexpectedly Consequential actions did not pass through a confirmation gate. Require a human check before purchases, sensitive data transmission, destructive edits, or irreversible actions.
Page appears blank or incomplete Navigation or rendering may still be in progress, or the page failed. Wait for an observable condition, capture a new screenshot, and use a bounded retry before stopping.

9. FAQ

Does computer use require a browser?

No. The runtime can control a browser or a desktop environment, depending on the integration and the task.

Can I use it to automate a legacy application?

Potentially. GUI interaction may help when the application has no suitable API, but first check whether a structured integration is available and more reliable.

Are benchmark scores a forecast for my agent?

No. Published results reflect specific systems, tasks, and evaluation setups. Evaluate your own representative workflows and verify their outcomes.

Does Codex Computer Use have the same data controls as an API integration?

No general equivalence should be assumed. Codex availability and data controls are product-specific. The Help Center describes separate local and cloud contexts and plan-specific controls; check its current wording for your account (OpenAI Help Center).

Implementation checklist

  • Pick code execution or structured actions based on your runtime and task.
  • Persist the session and conversation state across calls.
  • Send fresh observations and map resized screenshot coordinates correctly.
  • Constrain the environment, distrust screen instructions, and gate consequential actions.
  • Bound runs with a deadline, cancellation, and step limit.
  • Verify the real application state before declaring success.