How to capture a website screenshot with an AI agent in OpenAI Agents SDK
Use the OpenAI Agents SDK ComputerTool with a browser harness you control to capture website screenshots. See the setup, code, compatibility notes, and a no-browser alternative.
To capture a website screenshot with an AI agent in the OpenAI Agents SDK, connect the SDK’s ComputerTool to a browser harness that your application runs and controls. The harness implements the SDK’s Computer or AsyncComputer interface, including screenshot(), which returns a base64-encoded PNG of the current display. Add the tool to an Agent, then run it with Runner. The SDK does not supply or host the local browser runtime.
This approach fits tasks where an agent needs to inspect a page and interact with it through computer-style actions. If your application only needs a one-off screenshot, consider whether an agent-driven interaction loop is necessary; the official SDK sources cited here document the ComputerTool path.
How the pieces fit together
- Browser runtime: Your application starts and owns the browser or computer environment.
- Computer implementation: You implement the interface for that runtime. It supplies the screenshot and the interaction methods the agent can use, such as clicking, scrolling, typing, waiting, and keyboard input.
- ComputerTool: The SDK adapts your implementation to the computer-use tool surface.
- Agent and Runner: The agent receives the tool, and
Runnerexecutes the agent’s request and tool cycle. - Screenshot result: The harness returns the current display as a base64-encoded PNG according to the interface contract.
The SDK guide points to its Playwright-based computer-use example as a reference for browser setup and the full harness implementation. Follow that example for the exact methods required by the SDK version you install; the short integration sketch below is not a complete Playwright harness.
Prerequisites and implementation choices
- Choose a browser runtime and run it in your application environment. The SDK documentation does not establish that OpenAI hosts this runtime.
- Choose
Computerfor a synchronous driver orAsyncComputerfor an asynchronous driver. Match the interface to the execution model of your browser driver. - Implement the complete interface expected by the installed SDK version. The screenshot method alone is not enough for an agent that must navigate or interact with a page.
- Choose the model based on the current computer-use support documented for your deployment, and verify the model actually sent on the Responses request.
The names and defaults for supported models and tool formats can change. In the documented behavior described by the current guide, the effective model determines the computer tool request format: the GA path uses a computer payload and can return batched actions[], while the older computer-use-preview path uses computer_use_preview and a single action per call. A model override in run configuration or prompt templates can change which path applies. Check the current [OpenAI Agents SDK computer-use guide](https://openai.github.io/openai-agents-python/computer/) before publishing or deploying against a particular model.
Python setup with the Agents SDK
The following is an integration outline showing where your own browser harness connects to the SDK. It intentionally does not fabricate a driver implementation: use the official Playwright example for the concrete browser setup and all required interface methods. Confirm imports and method signatures against the version you install.
from agents import Agent, ComputerTool, Runner
# Implement this with your browser runtime. It must satisfy the SDK's
# Computer interface, including screenshot() and the supported actions.
computer = YourComputerImplementation()
agent = Agent(
name="Website screenshot agent",
instructions=(
"Use the computer tool to open the requested website, wait for it "
"to render, and inspect the visible page."
),
tools=[ComputerTool(computer)],
)
result = Runner.run_sync(
agent,
"Open https://example.com and capture the current page.",
)
print(result.final_output)
YourComputerImplementation is a placeholder, not an SDK class. Replace it with your implementation of the interface. If your browser driver is asynchronous, use the SDK’s AsyncComputer interface and the corresponding asynchronous run pattern from the installed SDK version. See the [Agents SDK computer-use guide](https://openai.github.io/openai-agents-python/computer/) and [Computer API reference](https://openai.github.io/openai-agents-python/ref/computer/) for the interface contract and example link.
Implementing the screenshot contract
The reference requires screenshot() to return a base64-encoded PNG of the current display. In practical terms, the harness should obtain the current display image from its browser or computer runtime, encode the PNG bytes as base64, and return the encoded value in the form expected by the SDK interface. Implement the remaining action methods required by the interface as well; the agent needs them to operate the page.
def screenshot(self):
"""Return a base64-encoded PNG of the current display."""
png_bytes = self.capture_display_as_png()
return base64.b64encode(png_bytes).decode("ascii")
This is a method sketch, not a standalone implementation: capture_display_as_png() represents the corresponding operation in your chosen runtime. Refer to the official Playwright example for a runnable harness and the exact interface methods.
Run the agent and obtain the screenshot
- Start the browser runtime in the same application environment as the harness.
- Have the harness navigate to the requested URL, either as part of its setup or through the action methods exposed to the agent.
- Register the harness with
ComputerTooland add the tool to the agent. - Run the agent. The SDK and model can request computer actions, and the harness carries them out against the local runtime.
- When the tool requests a screenshot, return the current display as a base64-encoded PNG.
- Use the SDK run result and tool cycle to handle the task outcome. If you need a PNG file outside the tool cycle, decode the screenshot data in your application at the boundary where you own the harness and output.
The precise orchestration and action payload depend on the effective model and SDK version. Do not assume that every model uses the same action shape or number of actions per call.
Operational details and edge cases
Page readiness
A navigation completing does not necessarily mean the page is visually ready. Decide how your harness waits for the content needed by the task: the interface can expose a wait action, and the agent can be instructed to wait before inspecting. Dynamic pages, lazy-loaded content, animations, consent dialogs, and slow resources can all affect what appears in a screenshot. Set a bounded wait strategy in your harness so a stalled page does not hold a run indefinitely.
Authentication and sensitive pages
If the target requires login, provide the browser runtime with the appropriate session state through your application’s normal secure mechanisms. Avoid putting credentials or session tokens in agent instructions or logs. Test with a page and account appropriate for the task, and consider what page content the screenshot and agent context may expose.
Viewport and capture scope
The screenshot represents the current display, so the browser window or viewport configured by your harness determines the visible area. If the task needs content below the fold, the agent may need to scroll and capture again. This computer-use path is based on display screenshots; do not assume it automatically produces a full-page image.
Browser lifecycle and cleanup
Your application owns the browser runtime. Ensure it starts before a run, remains available throughout the tool cycle, and is closed or returned to a pool after the task. Handle browser crashes and navigation failures in the harness so the agent run can terminate with a useful error rather than waiting forever.
Reliability, performance, and cost
- Reliability: The browser harness is a separate application component. Failures can come from browser startup, navigation, page scripts, the computer interface, or the model request. Add bounded waits and cleanup paths at the runtime boundary, and report which stage failed.
- Performance: An agent-driven screenshot can involve multiple model and tool steps, especially when the agent must navigate, wait, scroll, or inspect. For a fixed capture task, reduce unnecessary actions and avoid waiting longer than the page requires. No performance benchmark is established by the cited SDK documentation.
- Cost: This design involves your chosen model usage and the compute needed to run your browser runtime. The cited SDK documentation does not establish a fixed price for this end-to-end workflow. Check current model pricing and your infrastructure costs separately.
- Scaling: Each concurrent task needs an available browser context or runtime that can safely serve that task. Apply your own limits for concurrency, timeouts, and resource cleanup.
Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| The tool is unavailable to the agent | ComputerTool was not added to the agent, or the run is not using the expected agent configuration. |
Confirm the tool is registered in the Agent passed to Runner, and inspect the effective run configuration. |
| The screenshot method fails validation | The method returns raw image bytes or a different encoding instead of the documented base64-encoded PNG. | Return base64 text for PNG bytes in the exact form required by the installed interface. |
| The page is blank or incomplete | The screenshot was requested before navigation or rendering finished, or the page requires scrolling or interaction. | Check navigation errors, add an appropriate bounded wait, and expose the needed interaction methods to the agent. |
| Computer actions are rejected or have the wrong shape | The effective model and computer tool format do not match expectations, or an override selected a different model path. | Verify the model on the actual Responses request and compare its expected GA or preview format with the current guide. |
| The agent cannot interact with the browser | The harness implements screenshot() but omits an action method or uses an incompatible interface version. |
Implement the complete Computer or AsyncComputer contract for your SDK version and compare with the official example. |
| The run hangs on a slow site | Navigation, page readiness, or an action has no effective timeout. | Set bounded timeouts in the browser harness and ensure failed operations return control to the run. |
| Async calls fail or block unexpectedly | The asynchronous browser driver is paired with the synchronous interface or run flow. | Use AsyncComputer for an asynchronous driver and follow the corresponding patterns in the installed SDK version. |
Or skip the browser setup
If you need a website screenshot without building and operating a browser harness, [ScreenshotNeo](https://screenshotneo.com) provides a screenshot API and MCP server. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/) for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await Bun.write('shot.webp', bytes);
Replace the example URL with the page you need. The Node.js example uses Bun.write to save the response; in a Node.js application, write the returned bytes with your preferred filesystem API. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. [Create a free account](https://screenshotneo.com/account/sign-up/).
FAQ
Does the Agents SDK provide a hosted browser?
The documented ComputerTool path uses a computer implementation supplied by your application. You run and manage the browser runtime.
What format does the computer screenshot method return?
The interface contract specifies a base64-encoded PNG of the current display.
Should I choose the synchronous or asynchronous interface?
Use the one that matches your browser driver: Computer for synchronous execution and AsyncComputer for asynchronous execution.
Can I use this for a full-page screenshot?
The documented contract describes a screenshot of the current display. To capture content outside the viewport, use scrolling and additional captures or choose a capture method that explicitly supports full-page output.
Sources
- OpenAI Agents SDK computer-use guide, including the linked Playwright example and current model compatibility guidance.
- OpenAI Agents SDK Computer API reference, including the screenshot return contract.


