ScreenshotNeo

BlogHow-to

How to Add Visual AI to Automated Tests

Add AI-assisted visual checks to browser tests with repeatable states, focused checkpoints, baseline review, and a Playwright example.

By the ScreenshotNeo team4 October 20268 min read

To add visual AI to automated tests, keep your existing browser test responsible for reaching a known application state, then add a visual checkpoint that captures that state, compares it with an approved baseline, and sends differences for review. A screenshot by itself is only an image: for example, Cypress’s screenshot command does not compare images. Comparison and baseline review come from an integration or service. AI-assisted comparison is specific to some tools, not a property of every visual testing system.

This guide uses Applitools Eyes with Playwright for the concrete AI-assisted example. The fixture and eyes.check() call are Applitools-specific; check the current integration documentation for package setup and configuration details, which can change.

1. Understand the visual testing workflow

A useful visual test has four parts:

  1. Reach a meaningful state. Use the existing test to open the page, sign in, populate stable test data, or open the modal you want to protect.
  2. Capture a page or element. Place the checkpoint after the interface has rendered in the intended state.
  3. Compare against an approved baseline. The tool evaluates the new rendering against the reference image or representation.
  4. Review the difference. Accept a baseline update when the change is intentional; investigate and reject an unexpected difference.

Functional assertions and visual checks answer different questions. An assertion such as “the save button is enabled” checks a specific behavior or property. A visual comparison can catch a shifted button, changed typography, missing icon, unexpected wrapping, or layout change that those assertions do not cover. Keep functional tests: visual checks complement them rather than replacing them.

2. Choose checkpoints that protect important UI

Start with a small set of valuable screens: a key landing page, a shared component, a critical form, or a meaningful state such as an open dialog. A checkpoint in every test can create a review burden without adding useful coverage. Use an element-level check when you want to localize ownership and isolate a component. Use a full-page check when the overall layout is what you need to protect.

Checkpoint choice Useful when Watch for
Element A component has a clear owner and can be checked in a stable context. Changes outside the selected element are not covered by that checkpoint.
Full page Page composition, spacing, and overall layout matter. Unrelated dynamic content can create noisy differences.
Named state A page has important states such as validation errors, expanded menus, or a modal. Make the state explicit and repeatable before capture.

Give checkpoints names that identify both the screen and state, such as Account settings — validation error. Clear names make results easier to find during review.

3. Add an Applitools Eyes checkpoint to Playwright

The example follows the documented Applitools Playwright fixture style. It navigates to a page and asks Eyes to check a full-page rendering with a strict match level:

import { test } from '@applitools/eyes-playwright/fixture';

test('homepage visual checkpoint', async ({ page, eyes }) => {
  await page.goto('https://example.com');
  await eyes.check('Homepage', {
    fully: true,
    matchLevel: 'Strict',
  });
});

The fixture supplies the eyes object to the test. Call eyes.check() at the point where the intended UI has rendered. The checkpoint name should describe the state. Here, fully: true requests a full-page check and matchLevel: 'Strict' selects a comparison setting shown in the integration example. Choose settings for your interface and consult the current vendor docs before implementation.

This snippet is the checkpoint portion of an integration, not a complete project bootstrap: it assumes a Playwright project and the Applitools package and account configuration are already set up as described in the vendor’s documentation. Do not assume a bare Playwright screenshot automatically gains comparison or baseline management.

4. Make captures repeatable

Visual comparison is useful only when the same test state produces a sufficiently consistent rendering. Before adding more checkpoints:

  • Wait for the particular interface state you need instead of capturing immediately after navigation.
  • Use controlled test data and avoid time-dependent content where possible.
  • Run comparisons in a consistent rendering environment when practical.
  • For unavoidable dynamic areas such as third-party widgets or ads, mask or ignore the smallest relevant region if your integration supports it.
  • Prefer a focused checkpoint over relaxing comparison for an entire page just to hide one unstable area.

Applitools’ integration supports options such as full-page capture, match level, and ignored dynamic regions; confirm their current API spelling and behavior in its documentation. Cypress also cautions that incidental screenshots can create review overhead and recommends meaningful visual testing.

5. Establish and review baselines in CI

  1. Add the checkpoint to a test that already reaches the desired state reliably.
  2. Run it and establish or approve the initial baseline through the integration’s workflow.
  3. Run the visual test in the existing pipeline so changes are checked alongside normal browser tests.
  4. Inspect each reported difference. Accept a new baseline only when the UI change is intended; investigate and reject unexpected differences.

A baseline is an approved reference, not proof that the UI is correct. The initial image needs review, and future updates need review when they represent product changes. AI-assisted analysis can help compare renderings, but it does not decide whether a design change is intentional or acceptable for your product.

6. Select an integration that matches your suite

Visual testing integrations differ in how they capture, render, compare, and review changes. Cypress lists integrations including Applitools Eyes, Argos, Chromatic, Happo, LambdaTest SmartUI, Percy, Sauce Labs Visual, SmartBear VisualTest, and Wopee.io. The list is not a claim that every service uses AI. Cypress describes Applitools Eyes as AI-assisted; it describes Percy in terms of DOM snapshots rendered across browsers and responsive widths.

Selection question Why it matters
Does it support your framework and language? The integration should fit the browser tests your team already runs.
Do you need component checks, end-to-end states, or both? Tools vary in the kinds of checkpoints and workflows they support.
What does “AI” mean for this option? Verify the specific comparison capability rather than applying the label to all visual tools.
How are captures produced? Full-page, element, DOM-based, browser, viewport, and device choices affect what differences you can detect.
How does review work? Check baseline approval, diff grouping, permissions, and how CI reports results.
Can you send the page data to the service? Consider your test data and privacy requirements when choosing hosted rendering and review.

For Playwright users seeking AI-assisted visual comparison, Applitools Eyes is a documented option. Compare the current integration and review flow against your team’s needs. Cypress itself captures screenshots but does not provide image comparison through its built-in screenshot command; its documentation describes plugins and services for that step.

7. Troubleshooting common visual test problems

Symptom Likely cause What to do
A screenshot exists, but there is no visual result or diff. Capture and comparison are separate functions; the test may only save an image. Configure a comparison integration and its baseline workflow. Cypress’s built-in screenshot command alone does not compare images.
The same test reports differences on repeated runs. Unstable state, changing data, time-dependent content, inconsistent rendering, or third-party content. Stabilize the state and environment; isolate or mask only the smallest uncontrollable area.
A change appears in a diff, but its cause is unclear. The checkpoint name or scope does not make the UI state obvious. Use descriptive names and capture the specific state that matters; review the changed region in context.
A legitimate UI update keeps failing against the old reference. The baseline still represents the previous design. Review the intended product change, then approve the new baseline through the tool’s review workflow.
Unexpected changes are being accepted as routine updates. Baseline approval is happening without checking whether the change is intentional. Inspect diffs before accepting. Keep baseline updates tied to reviewed UI changes.
The Applitools fixture import or checkpoint options do not match the project. Package setup or API details may have changed, or the project is not configured for the documented fixture. Follow the current Applitools Playwright integration documentation and verify package and setup instructions.

8. Performance, reliability, and cost considerations

Each checkpoint adds capture and comparison work to a test run, and a large number of low-value checkpoints also adds review work. Keep the first rollout small, prioritize important states, and expand when a checkpoint covers a real risk. The cited documentation does not establish neutral cross-vendor speed, accuracy, or price benchmarks, so compare current vendor plans and operational fit directly rather than relying on an invented ranking.

Reliability depends on deterministic test data, a stable rendered state, and a consistent environment. Keep a human approval step for meaningful baseline changes. For hosted integrations, check the provider’s current pricing, retention, and data handling terms for your test content; those details are not established by the workflow documentation summarized here.

9. Or skip the browser setup

If your immediate need is a clean screenshot of a URL rather than a baseline comparison inside an existing browser suite, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. It is a capture tool; use your visual testing integration for baseline comparison and review.

Install no browser harness for a simple capture. This cURL example saves a WebP screenshot of Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed along with known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server lets AI agents use screenshot, page-info, and PDF tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, no card required.

10. FAQ

Does adding AI visual testing remove the need for functional assertions?

No. Functional assertions check behavior and specific conditions; visual comparison checks rendered appearance against a reference. Keep both where they cover different risks.

Can I use a screenshot command as a visual regression test?

Not by itself. A capture must be compared with a reference, and the resulting difference needs a review workflow.

Should every test get a visual checkpoint?

Usually start with a few important screens and states. Add checkpoints when they protect a meaningful UI risk and the result can be reviewed usefully.

Does every visual testing integration use AI?

No. Check the documented capabilities of the specific integration. The cited Cypress documentation describes Applitools Eyes as AI-assisted and Percy as using DOM snapshots for rendering.

Sources