Logging Browser Automation Actions for AI Agents
Use Playwright traces, screenshots, DOM snapshots, and OpenTelemetry to reconstruct every browser-agent action and diagnose failures safely.

Direct answer: start a Playwright browser context trace before each agent task, enable screenshots and DOM snapshots, stop the trace on success or failure, and open the resulting archive in Trace Viewer. Add structured logs and OpenTelemetry spans around agent decisions, tool calls, and outcomes so browser evidence can be correlated with your backend traces. Redact credentials, payment data, tokens, and unnecessary page content before exporting or retaining traces.
Playwright tracing is the most direct action-level record of what an AI browser agent did. It captures browser operations and network activity, with optional screenshots and DOM snapshots. Trace Viewer presents a timeline for each action, including the before and after DOM state, screenshots, locator details, timing, console messages, and source locations. Context tracing does not record test assertions; if assertions are part of your failure record, use the Playwright test runner’s tracing features as well.
What to log for an AI browser agent
An agent run usually spans four layers: the model’s plan, the automation tool call, the browser’s observable state, and services called by your application. A useful audit trail preserves enough information to answer five questions:
- What did the agent intend to do?
- Which locator, URL, or element did it target?
- What did the browser show before and after the action?
- What network, console, or navigation events occurred?
- How did the task finish, and can the evidence be correlated with server logs?
Use a stable run ID and step number everywhere. Recommended attributes are run_id, step_number, action_type, target_locator, url, result, error_class, started_at, and finished_at. Keep the agent’s natural-language reasoning out of long-term logs unless you have a clear privacy and retention policy; record the selected action and its result instead.
Playwright tracing: the action-level record
Tracing is attached to a browser context. Start it before the agent begins work and stop it in a finally block so failures produce a trace too. Set screenshots: true and snapshots: true when reviewers need to reconstruct visual and DOM state. Set sources: true when source locations will help diagnose the automation code.

Complete Node.js example
import { chromium } from 'playwright';
import { randomUUID } from 'node:crypto';
import fs from 'node:fs/promises';
const runId = randomUUID();
const tracePath = `traces/${runId}.zip`;
await fs.mkdir('traces', { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
viewport: { width: 1440, height: 1000 },
serviceWorkers: 'block'
});
await context.tracing.start({
screenshots: true,
snapshots: true,
sources: true,
title: `agent-run-${runId}`
});
const page = await context.newPage();
const startedAt = new Date().toISOString();
try {
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.getByRole('link', { name: 'More information...' }).click();
console.log(JSON.stringify({
run_id: runId,
step_number: 1,
action_type: 'click',
target_locator: 'role=link[name="More information..."]',
url: page.url(),
result: 'ok',
started_at: startedAt,
finished_at: new Date().toISOString()
}));
} catch (error) {
console.error(JSON.stringify({
run_id: runId,
action_type: 'task',
result: 'error',
error_class: error?.name || 'Error',
message: error?.message
}));
throw error;
} finally {
await context.tracing.stop({ path: tracePath });
await context.close();
await browser.close();
}
console.log(`Trace written to ${tracePath}`);
Run it with node agent-trace.mjs. Inspect the archive with the Playwright Trace Viewer:
npx playwright show-trace traces/<run-id>.zip
Do not delete the trace before it has been uploaded or reviewed. If your agent can execute several independent tasks in one context, use a separate trace group or context per task so the timeline remains understandable.
Python example
from pathlib import Path
from uuid import uuid4
from playwright.sync_api import sync_playwright
run_id = str(uuid4())
Path("traces").mkdir(exist_ok=True)
trace_path = Path("traces") / f"{run_id}.zip"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(viewport={"width": 1440, "height": 1000})
context.tracing.start(
screenshots=True,
snapshots=True,
sources=True,
title=f"agent-run-{run_id}"
)
page = context.new_page()
try:
page.goto("https://example.com", wait_until="domcontentloaded")
page.get_by_role("link", name="More information...").click()
print({"run_id": run_id, "action_type": "click", "result": "ok"})
finally:
context.tracing.stop(path=str(trace_path))
context.close()
browser.close()
print(f"Trace written to {trace_path}")
What Trace Viewer reveals
Trace Viewer turns a ZIP archive into an action timeline. Select an action to inspect its locator, duration, source location, console output, and network activity. The before and after snapshots help distinguish a bad locator from a page that never reached the expected state. Screenshots show visual changes such as cookie dialogs, overlays, or a page rendered in the wrong responsive layout.
Network records are useful for diagnosing redirects, failed API calls, blocked resources, and unexpectedly slow requests. Console messages often expose JavaScript exceptions that do not surface as a failed click. Store the trace path or object-storage key with the run ID in your job database so an incident reviewer can find it without searching by timestamp.
Tracing limitations and assertions
Playwright’s context.tracing API captures browser operations and network activity, but it does not record test assertions such as expect calls. This matters when an agent reaches a page but your task-level check fails. Add an explicit event around every assertion or outcome:
function recordOutcome(runId, step, result, details = {}) {
console.log(JSON.stringify({
run_id: runId,
step_number: step,
action_type: 'assertion',
result,
...details,
timestamp: new Date().toISOString()
}));
}
recordOutcome(runId, 2, 'error', {
error_class: 'MissingExpectedText',
expected: 'Order confirmed'
});
For suites built with Playwright Test, enable the test runner’s trace modes when you need assertion context and fixture information. Keep the context trace as the browser evidence and your structured event as the task evidence.
Correlating browser evidence with OpenTelemetry
OpenTelemetry is a vendor-neutral framework for generating, collecting, and exporting traces, metrics, and logs. Create one server-side span for the agent run, child spans for model decisions and browser actions, and events for trace creation, upload, and retention. Put the same run_id in span attributes, logs, and the trace filename.
import { trace } from '@opentelemetry/api';
const tracer = trace.getTracer('browser-agent');
async function runStep(runId, stepNumber, actionType, fn) {
return tracer.startActiveSpan(`browser.${actionType}`, async (span) => {
span.setAttributes({
'agent.run_id': runId,
'agent.step_number': stepNumber,
'agent.action_type': actionType
});
const started = Date.now();
try {
const value = await fn();
span.setAttributes({
'agent.result': 'ok',
'agent.duration_ms': Date.now() - started
});
return value;
} catch (error) {
span.recordException(error);
span.setAttributes({
'agent.result': 'error',
'agent.error_class': error?.name || 'Error'
});
throw error;
} finally {
span.end();
}
});
}
Use OTel to connect the browser step to your queue, API, database, and model-provider spans. Browser client instrumentation itself is experimental and mostly unspecified, so prefer stable server-side instrumentation around your Playwright worker. Export traces to the backend already used by your team, then attach a link to the Playwright archive in the run record.
Privacy, redaction, and retention
Screenshots, DOM snapshots, request URLs, headers, response bodies, and console messages can contain credentials, personal data, payment details, access tokens, or private page content. Treat trace archives as sensitive production data.
- Use synthetic accounts and test data where possible.
- Never put API keys or cookies in locator labels or event messages.
- Redact authorization headers and query parameters before exporting network records.
- Keep screenshots and DOM snapshots only as long as the incident or quality workflow requires.
- Encrypt archives in transit and at rest, and restrict Trace Viewer access.
- Record a deletion timestamp or retention class with every run.
There is no universal retention period. Measure how often reviewers need old traces, how large archives become, and what your regulatory obligations require. Sampling successful runs while retaining all failures can reduce storage, but document the sampling rule so an audit trail remains predictable.
Capturing screenshots for a reviewable audit trail
Playwright traces are best when the agent itself must interact with the page. If your workflow only needs a clean visual record of a URL, a screenshot API can remove browser orchestration from the worker.

Or skip the browser setup
ScreenshotNeo returns a PNG, JPEG, WebP, or PDF from one GET request. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether it was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options. The basic call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For agent evidence, save the image beside the run record and persist X-Page-Verdict and X-Billed. ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-element capture, dark mode, device presets, arbitrary viewports, retina scale, custom CSS and JavaScript, click and wait actions, request blocking, headers and cookies, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. These options let you capture a deterministic visual checkpoint without maintaining a Playwright worker.
There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to capture your first run evidence.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| No trace file after a crash | tracing.stop() was never reached |
Stop tracing in finally; write to a durable local or object-storage path. |
| Trace opens but has little visual context | Screenshots or snapshots were disabled | Enable screenshots: true and snapshots: true for diagnostic runs. |
| Assertion failure is absent | Context tracing does not record assertions | Emit explicit assertion events or use Playwright Test tracing. |
| Clicks appear to target the wrong element | Ambiguous or unstable locator | Prefer role, label, or test ID locators; inspect before and after snapshots. |
| Trace is unexpectedly huge | Long sessions, screenshots, or verbose network bodies | Trace per task, sample successful runs, and apply retention and redaction rules. |
| Navigation hangs | Network idle never occurs or a third-party request remains open | Use explicit load conditions and bounded timeouts; record the timeout class. |
| Screenshot contains a consent dialog | The page requires an interaction before content is visible | Click the consent control in Playwright, or use ScreenshotNeo’s consent handling. |
| Screenshot request is not billed as expected | The page verdict was a bot check, blank page, timeout, failed load, or cache hit | Inspect X-Page-Verdict and X-Billed headers and correct the target or wait settings. |
Performance, reliability, and cost
Tracing adds work because screenshots, DOM snapshots, and network records must be collected and compressed. Measure overhead on your pages rather than relying on a general percentage: page size, animation, request volume, and trace duration all change the result. A practical pattern is full tracing for failures and sampled successful runs, with a lightweight structured event for every run.
Reliability improves when each action has a bounded timeout, a stable locator, and an explicit result event. Persist the trace only after the archive is closed, retry uploads separately from browser actions, and make upload keys idempotent using the run ID. For long workflows, split traces at task boundaries so one corrupt archive does not hide every preceding step.
Storage cost is driven by archive size and retention, not just the number of runs. Track bytes per trace, traces per day, review rate, and export cost. ScreenshotNeo charges only for clean shots; failed loads, blank pages, bot checks, timeouts, and cache hits are not billed, which makes repeated evidence collection easier to budget.
Operational checklist
- Create a unique run ID before launching the browser.
- Start tracing before the first agent action.
- Enable screenshots and DOM snapshots for visual reconstruction.
- Log every tool call with step number, locator, URL, result, and timestamps.
- Emit an explicit event for assertions and task outcomes.
- Correlate browser events with OpenTelemetry spans.
- Redact secrets and sensitive page data before export.
- Stop tracing in a
finallyblock on success and failure. - Store the archive key beside the run record.
- Define sampling, retention, access, and deletion rules.
FAQ
Can a Playwright trace replay an AI agent’s reasoning?
No. It records browser operations and state. Store the selected action and tool arguments as structured events if you need a decision history.
Should every successful run keep a full trace?
Not necessarily. Keep all failures, sample successes, and retain lightweight outcome events for every run according to your review and compliance needs.
Is OpenTelemetry required?
No. Playwright tracing is sufficient for browser-focused debugging. Add OpenTelemetry when the run crosses your agent service, queues, APIs, and other backends.
Where should trace archives live?
Use encrypted object storage or another durable store with restricted access. Keep only a reference and metadata in your run database.
Can I use ScreenshotNeo with an AI agent?
Yes. Its MCP server exposes screenshot, page-info, and PDF tools to Claude, Cursor, and other MCP clients, while its HTTP API works from any language that can make a GET request.