ScreenshotNeo

BlogHow-to

How to Use Applitools Eyes Test Manager

Find the right Eyes run, review checkpoint differences against their baselines, and decide when to accept or reject a change.

By the ScreenshotNeo team4 October 20267 min read

Applitools Eyes Test Manager is where you review visual test results and manage baselines. Open the run for the relevant test, compare each checkpoint with its baseline, investigate differences, then accept and save an intended UI change or reject a change that appears to be a defect. A visual difference is a prompt to investigate, not a reason to accept automatically.

Test execution and screenshot capture happen in your test suite through the Eyes SDK. Test Manager is the review stage of that flow. The workflow below covers how to find a run, interpret a checkpoint, and keep baseline updates deliberate.

1. Run a visual test and open its results

Run the test suite that contains your Eyes checks. Each check captures a UI state, called a checkpoint, and compares it with the stored baseline for that test. The run result or a link from your test output should take you to the corresponding results in Test Manager.

If the run is new, its first captured images establish initial baselines because no earlier references exist. Review these images for correctness before treating them as expected behavior for future runs. On later runs, Eyes compares checkpoints with those references and flags differences for review. See Applitools’ Test Manager overview and its explanation of how baselines are created.

2. Find the relevant run

Start from the Test Manager home screen and locate the run associated with the code change you are reviewing. The home screen groups runs and can show context such as test name, result, date and time, browser or device, operating system, and viewport. Use available status filters and open a test to inspect its details. Interface labels can change, so use the run’s identity and environment details to confirm you have the right result.

  1. Match the test name and application to the change under review.
  2. Check the run’s execution date and result.
  3. Confirm browser, operating system, and viewport before interpreting a difference.
  4. Open the test or checkpoint with the unresolved difference.

Baselines are associated with a particular test and environment. Applitools lists operating system, viewport size, browser, application name, and test name among the identifying factors. A different browser or viewport may therefore use a different baseline. Confirm that you are looking at the expected environment before accepting or rejecting anything. Applitools baseline details.

3. Compare the checkpoint with its baseline

Open the checkpoint to inspect the current image and its baseline. Use the available comparison views and difference highlighting to locate what changed. Look at the changed area in context: a highlighted difference might be an intended design update, a content change, an unexpected layout shift, or a rendering difference tied to the environment.

For each difference, ask:

  • Was this UI change part of the intended code or design change?
  • Does it affect the content, layout, color, typography, or a user action?
  • Is the test running against the intended browser, operating system, viewport, and application?
  • Could dynamic content, an experiment, or test data explain the alternate image?
  • Would accepting this image make an unintended regression the reference for future runs?

If the difference reflects an approved product change, accept it and save the baseline update. If it looks like a bug, reject it so the existing expected image remains the reference. If you do not yet know which explanation is correct, investigate before resolving the difference. Accepting every flagged result can normalize a regression.

4. Accept or reject, then save reviewed changes

  1. Accept an intended change. This updates the baseline to reflect the reviewed checkpoint image.
  2. Reject a suspected defect. This keeps the prior baseline as the expected reference and records the difference as unintended.
  3. Save the reviewed update. Confirm the save action so accepted baseline changes become the reference for later runs.
  4. Re-run when appropriate. A new run can confirm the code and reviewed baseline now produce the expected result.

The exact controls and labels may vary with the current Test Manager interface. The essential decision remains the same: accept changes that are intentional; reject changes that represent defects. Applitools documents this accept/reject workflow in its Test Manager walkthrough.

5. Use baseline variations for expected alternatives

Some applications can legitimately show more than one version of a step, such as when an A/B experiment serves different page variants. In that case, baseline variations let Eyes compare a checkpoint with multiple accepted references. Use a variation when the alternative is known and valid; do not use one to make an unexplained difference disappear.

Applitools documents a maximum of 20 baseline variations for a test step. Create and save a variation for a distinct approved rendering. When the experiment ends and the alternatives are no longer valid, review and remove or merge variations as appropriate. See Applitools’ baseline variations documentation for the current management steps and limits.

6. Understand match sensitivity

Match level affects which visual changes are surfaced. The cited Applitools settings page describes Strict as the default and recommended level in its context: it checks visible details such as text, fonts, color, graphics, and element positions while aiming to ignore platform-dependent pixel differences. Confirm the level used by your test and the guidance for your Eyes configuration before interpreting results. Applitools Visual AI options.

When a difference is noisy, first confirm the environment, checkpoint, and test data. Then review whether the configured match level fits what the test is meant to validate. A setting that ignores a category of change can also hide a meaningful regression, so match sensitivity should reflect the purpose of the check.

7. Troubleshooting common review problems

Symptom Likely cause What to do
The run appears to have no matching baseline This may be the first run for the test and environment, or the test/environment identity differs from the existing baseline. Check application name, test name, browser, operating system, and viewport. Review the initial checkpoint before establishing it as the expected reference.
A difference appears after switching browsers or viewport sizes Those settings can identify a separate baseline. Confirm which environment the run represents and review its matching baseline instead of accepting the image into an unrelated reference.
A checkpoint is flagged but the change may be intended The run shows a visual difference; it cannot determine product intent for you. Check the change request or design decision, inspect the changed area, and accept only after confirming it is expected.
The checkpoint differs between runs for a known experiment The application can render multiple legitimate variants. Consider an explicit baseline variation for each approved alternative. Keep within the documented maximum of 20 per test step.
Many differences appear to be caused by shifted content A changed element, content length, or rendering context may have moved other elements. Inspect the earliest meaningful difference and check the environment and test data. Do not accept the whole checkpoint until you understand the cause.
An accepted change is not used by later tests The baseline update may not have been saved, or later tests may be using a different test/environment identity. Confirm the save completed, then compare the later run’s application, test, browser, operating system, and viewport.
A rejected result is still unresolved The review may not have been fully saved or the run may contain other unresolved checkpoints. Check the status of each differing step and save the review. Resolve each checkpoint based on its own evidence.

Performance, reliability, and review cost

Test Manager is downstream of test execution: it reviews the checkpoints your suite captured. Reliable reviews depend on meaningful checkpoints, consistent test and environment identity, and stable test data. If runs vary across browsers, viewports, or application states, the added variation can make it harder to identify which change matters.

The main operational cost of baseline review is the risk of accepting the wrong image. A mistaken acceptance can make a defect the expected appearance for future runs; a mistaken rejection can leave an intended product update unresolved. Review changed regions, preserve the environment context, and use variations only for known alternatives. The supplied research does not establish performance benchmarks or pricing for this workflow, so this guide makes no claims about either.

Or skip the browser setup

For capturing a website screenshot directly, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns an image or PDF. It is a separate capture workflow from Applitools Eyes visual testing and baseline review.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
  • Cookie banners are accepted like a visitor, and cookie banners, popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, and failed loads are never billed.
  • An MCP server lets AI agents use screenshot tools.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.

FAQ

Does Test Manager run my visual tests?

No. Your test suite and Eyes SDK capture checkpoints; Test Manager is where you review the resulting runs and manage baselines.

Should I accept every unresolved difference to make the run pass?

No. Accept only reviewed changes that are intended. Reject suspected defects so the existing expected image remains the reference.

How many baseline variations can one test step have?

Applitools’ documentation gives a maximum of 20 variations for a test step.

Does a different browser always use the same baseline?

Not necessarily. Browser and other environment details are among the factors Applitools lists for identifying a baseline. Check the run context.