Reviewing Visual Changes in AI Agent Builds
Use a repeatable workflow to inspect an AI agent’s code diff, check the rendered UI, compare screenshots, and decide whether a change is ready to merge.
Review an AI agent’s UI change by checking both the code diff and the rendered application, then comparing a controlled screenshot of the new state with an accepted baseline. Judge every visible difference against the requested change, verify the important interaction and viewport, and keep approval with a responsible reviewer. A screenshot or AI summary is evidence for review, not proof that the build is correct.
This guide covers a repeatable manual workflow, browser checks, screenshot comparisons, CI options, common failure modes, and ways to capture pages without maintaining a local browser setup.
1. Start with the request and code diff
Before opening the app, restate the intended outcome in concrete terms: which route changes, what users should see or do, and what must remain working. Then inspect the agent’s changed files and diff. A diff can reveal an accidental broad refactor, an unrelated asset change, or a missing test that a screenshot alone would not explain. VS Code documents reviewing agent edits through the diff view, Source Control, or a pull request workflow. VS Code: review AI-generated changes.
- Read the issue or prompt and write down the expected visible state and behavior.
- Review every changed file, including styles, assets, routing, and tests.
- Look for changes outside the request and ask the agent to explain unclear edits.
- Do not commit, merge, or apply a worktree’s changes until the review is complete.
For UI work, source review is necessary but incomplete. It tells you what changed in the repository; it cannot establish how the app renders or behaves in a browser.
2. Run the app and inspect the rendered result
Start the application using the repository’s documented development command, or use the project’s preview environment. Open the affected route at the intended viewport and inspect the actual page. Check the main interaction as well as its initial appearance: open menus, submit forms, dismiss dialogs, and navigate through the changed flow where relevant.
VS Code’s browser feedback loop describes having an agent start or locate an app, open it, interact with it, inspect page content and screenshots, check console errors, and repeat after fixes. The available tools depend on session setup and settings; use the current documentation for your editor version. VS Code: use browser tools with agents.
- Load the exact changed route, not just the home page.
- Inspect the target viewport and at least one nearby responsive width if layout is affected.
- Exercise the changed interaction from a clean starting state.
- Check console output and network failures for errors related to the change.
- Capture a screenshot when the visual state matters or a failure needs to be explained.
- Give the agent specific evidence and a desired result, then repeat the review after its fix.
A useful feedback request names the route, viewport, observed issue, and expected outcome. For example: “At 390 × 844 on /settings, the Save button wraps below the card. Keep it on one line by reducing the horizontal gap; preserve the desktop layout.”
3. Make the visual check reproducible
A screenshot comparison is useful only when you know what state it represents. Record the route, viewport dimensions, device scale factor, browser or rendering environment, relevant user state, and any setup steps. Use the same conditions for baseline and current captures.
- Route and state: Include query parameters, logged-in or logged-out state, selected tab, and any interaction required to reach the screen.
- Viewport: Fix width and height. If the change targets mobile, include the specific mobile viewport rather than relying on a desktop capture.
- Timing: Wait for the relevant content or selector. Avoid capturing during animation or before fonts and images load.
- Data: Use stable fixtures where possible. Live timestamps, rotating promotions, randomized content, and personalized data can produce noisy diffs.
- Environment: Keep browser, operating system, fonts, locale, timezone, and device scale consistent when comparing pixel output.
- Baseline: Update the accepted baseline only after a reviewer decides the new appearance is intended.
Do not treat every changed pixel as a defect. Antialiasing, dynamic content, and rendering differences can create noise; conversely, a small but meaningful shift in a button or error message can matter. Inspect the diff in context and keep ignored regions reviewable.
4. Compare screenshots and interpret the difference
For a focused review, capture the same route and state before and after the agent’s change. A visual regression system can automate this comparison and present differences for review. Argos describes deterministic pixel comparison, adjustments intended to reduce rendering noise, pull request review, and comments pinned to pixels; these are vendor-described capabilities, not a guarantee that every reported difference is meaningful. Argos.
When a diff appears, ask:
- Is the changed region part of the requested work?
- Does the new appearance support the requested behavior and hierarchy?
- Did a neighboring component, breakpoint, or shared layout change unintentionally?
- Can the difference be explained by content, timing, fonts, or environment rather than code?
- Does the relevant interaction still work, including keyboard use where applicable?
BrowserStack Percy documents AI-assisted visual summaries that can be compared with pull request context. Its classification is advisory: the vendor cautions that AI may miss or misinterpret changes and asks reviewers to inspect before approval. Standard visual diffs remain useful when the classification is unavailable. Percy: visual review agent.
5. Use agent-authored browser checks carefully
An agent can help write or run browser checks, but generated tests need the same scrutiny as generated application code. Give it current framework documentation and repository conventions. Inspect selectors and assertions, run a focused test, repeat it to look for instability, and review the generated diff. Selenium’s guidance recommends reviewing locators and repeating tests; it also notes that a screenshot captured at failure can reveal an overlay or other context absent from a stack trace. Selenium documentation.
Prefer checks that assert user-visible outcomes rather than implementation details. A test that confirms a dialog opens and its primary action is available is usually more resilient than one tied to incidental DOM structure. Keep a failure screenshot and relevant console output with the report so the agent and reviewer can see the state that failed.
6. Choose a review approach that fits the team
| Approach | Useful evidence | Good fit | Limit to account for |
|---|---|---|---|
| Manual editor and browser review | Diff, rendered page, interaction, console | Small changes, exploratory work, one-off routes | Review conditions can vary unless recorded |
| Browser tests with screenshots | Assertions, failure state, screenshot, logs | Important repeatable user flows | Tests and selectors can be flaky or too coupled to markup |
| Visual regression in CI | Baseline comparison, changed pixels, PR context | Shared components and recurring UI changes | Dynamic content and environment differences create noise |
| Combined review | Code, screenshots, behavior, and PR intent | Changes with meaningful visual or interaction risk | Needs clear ownership of baseline updates and final approval |
Tools such as Percy and Argos document screenshot comparison and pull request review workflows. Choose based on supported stack and browser, CI integration, collaboration needs, data handling, noise controls, and current pricing. The research for this article did not verify current prices or make independent product comparisons, so check vendor documentation before selecting a service.
7. Troubleshooting visual review
| Symptom | Likely cause | What to do |
|---|---|---|
| Screenshot is blank or incomplete | Capture happened before navigation, app rendering, or required content finished | Wait for a stable page condition or target selector; check the route and application logs. |
| Large diff on every run | Unstable data, animation, timestamps, ads, or mismatched capture conditions | Use stable fixtures, wait for animation to end, align viewport and environment, and isolate genuinely dynamic regions. |
| Click is intercepted or a control cannot be reached | Cookie banner, modal, overlay, or sticky element covers the target | Capture the failure state, identify the covering element, then decide whether the test should dismiss it or the page behavior needs fixing. |
| Diff misses a visible regression | Wrong route, viewport, or interaction state; baseline already contains the issue | Verify capture metadata, inspect related breakpoints and states, and confirm the baseline was previously accepted. |
| Agent says the UI is correct but evidence is weak | Summary is being treated as proof, or only source changes were reviewed | Ask for the route, viewport, screenshot, interaction result, and console status. Make the approval decision from the evidence. |
| Browser test passes once but fails later | Timing or selector instability, asynchronous content, or shared state | Repeat the focused test, wait on meaningful conditions, remove order-dependent state, and inspect the failure screenshot and logs. |
8. Keep the approval decision with a reviewer
AI summaries and visual classifiers can prioritize what to inspect, but they do not establish that a change matches intent. Review the relevant screenshots, confirm the important interaction and viewport, and validate the integrated result before keeping or merging the agent’s work. VS Code explicitly recommends testing the integrated result before archiving or deleting an agent session. VS Code agent review guidance.
- Ready: The requested behavior is present, the important views look right, relevant checks pass, and no unexplained regression remains.
- Revise: A specific mismatch or regression is visible; send the agent a focused correction with evidence.
- Investigate: The result is ambiguous because a baseline, state, environment, or test is unreliable; stabilize the check before approving.
Or skip the browser setup
If you need a page capture for review without setting up a local browser, ScreenshotNeo is a website screenshot API and MCP server for developers. Its API returns a screenshot or PDF from one GET request. Cookie banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. The MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Install the ScreenshotNeo API documentation for request options and integration details. This runnable cURL example saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Use a capture API as evidence for the page it actually loaded; still verify route, state, viewport, and intended behavior before approving a change. ScreenshotNeo includes full-page and element captures, device presets and custom viewports, dark mode, custom CSS and JavaScript, selector and network-idle waits, request blocking, cookies and headers, caching, async jobs, bulk capture, and PDF options. Its parameter names also work with those used by other screenshot APIs. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, no card required.
FAQ
Should every agent UI change get a visual regression test?
No. Use repeatable screenshot checks where the route or component is important enough to justify maintaining a stable baseline. For small or exploratory changes, a focused manual browser review may be sufficient.
Can an AI visual summary approve a pull request?
Treat it as triage. Review the screenshots and context yourself; vendor documentation cautions that AI classifications may miss or misinterpret changes.
What should I attach when asking an agent to fix a visual issue?
Provide the route, viewport, screenshot or diff, exact observed mismatch, expected appearance or behavior, and any relevant console or test output.
When is the review complete?
After the requested result and key interaction are verified, relevant regressions are resolved, and the integrated build has been checked.


