How to Ignore Flaky Changes in Visual Regression Tests
Fix flaky visual tests by stabilizing data, timing, and rendering first, then narrowly mask truly unpredictable content in Playwright or Chromatic.
A visual regression test is flaky when repeated captures differ even though the application code has not changed. The reliable fix is to stabilize the rendered state first: use fixed data, make fonts and assets predictable, wait for the UI state the test needs, and control incidental animation. Mask or ignore only the smallest region whose changing pixels are not part of the behavior being tested. A mask can hide a real layout regression along with volatile text.
This guide shows how to diagnose changing pixels, stabilize Playwright screenshots, use Playwright masks and styles, configure Chromatic ignores, and decide when a baseline update is justified. For tool-specific behavior, follow the linked official documentation because defaults can differ across tools and versions.
1. Find out what is actually changing
Before changing a threshold or adding an ignore, capture the same page several times with unchanged code. Compare the images and classify the differences:
- Content: timestamps, randomized values, rotating recommendations, live counters, or data from a changing API.
- Timing: a screenshot taken before a request, font, image, transition, or client render has finished.
- Motion: a CSS animation, video, animated image, or moving carousel captured at different frames.
- Rendering: fonts or image resources that load inconsistently, or different browser and operating-system rendering.
- Layout: a late-loading asset changes element dimensions, or the viewport or browser environment differs between runs.
If the whole page shifts, investigate viewport, browser/environment consistency, and layout readiness before masking individual elements. Late resources, dynamic inputs, and layout behavior are recognized causes of unstable visual tests. Chromatic’s unstable tests guide recommends making data and resources predictable.
2. Stabilize the page before suppressing differences
- Use fixed inputs. Replace live API data with a fixture or seeded dataset. Fix timestamps, random values, locale, and user state when they are not the subject of the test.
- Make assets predictable. Use local static images or placeholders where appropriate. Ensure fonts are served reliably or preloaded before capture.
- Wait for the intended state. Wait for a meaningful selector or application-ready condition, not a generic delay. A sleep can hide a race on one machine and still fail on another.
- Keep the capture environment consistent. Pin the viewport and browser configuration used by the test. Avoid comparing screenshots made with different device scale factors or rendering environments.
- Control motion. Disable incidental transitions and animation when the assertion concerns the settled page. Keep motion visible in a separate test if animation behavior is what you need to verify.
Chromatic recommends stable data and resources, including static or placeholder assets where useful and reliably served or preloaded web fonts. The correct readiness condition is application-specific. Chromatic: debugging unstable tests.
3. A runnable Playwright example
Playwright screenshot assertions compare a capture with a stored reference. This example fixes the viewport, replaces a changing timestamp with deterministic text, waits for the page’s ready marker, and captures the settled state. Install Playwright with npm install -D @playwright/test, save this as tests/dashboard.spec.js, and run npx playwright test. Replace the URL and selectors with your application’s values.
const { test, expect } = require('@playwright/test');
test('dashboard has a stable visual state', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.goto('http://localhost:3000/dashboard');
// Supply deterministic data before application code reads it, if the app supports it.
await page.addInitScript(() => {
Math.random = () => 0.42;
Date.now = () => new Date('2025-01-15T12:00:00Z').getTime();
});
// Prefer waiting for the state the test needs over an arbitrary sleep.
await page.getByTestId('dashboard-ready').waitFor({ state: 'visible' });
await expect(page).toHaveScreenshot('dashboard.png', {
animations: 'disabled',
fullPage: true
});
});
Important: addInitScript runs after navigation begins only when registered before the page’s application code executes. In the example it is registered after goto, which is too late for code that already read time or randomness. Register it before navigation as shown in this corrected version:
const { test, expect } = require('@playwright/test');
test('dashboard has a stable visual state', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.addInitScript(() => {
Math.random = () => 0.42;
Date.now = () => new Date('2025-01-15T12:00:00Z').getTime();
});
await page.goto('http://localhost:3000/dashboard');
await page.getByTestId('dashboard-ready').waitFor({ state: 'visible' });
await expect(page).toHaveScreenshot('dashboard.png', {
animations: 'disabled',
fullPage: true
});
});
Only freeze time or randomness if that matches the test contract. If the page uses a server-rendered timestamp or remote API data, control those inputs at the data boundary too; overriding browser APIs cannot make server responses deterministic.
4. Mask only content that is intentionally variable
When a changing region cannot reasonably be stabilized and its content is outside the test’s purpose, mask the narrowest element. In Playwright, pass locators to mask. The mask overlays the element’s bounding box, so it can hide a position or size change as well as the volatile pixels. Do not mask a region if its geometry is part of the regression contract. Playwright visual comparisons and the PageAssertions API document screenshot comparison and masking options.
const { test, expect } = require('@playwright/test');
test('dashboard ignores only its live clock', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.goto('http://localhost:3000/dashboard');
await page.getByTestId('dashboard-ready').waitFor({ state: 'visible' });
await expect(page).toHaveScreenshot('dashboard.png', {
animations: 'disabled',
mask: [page.getByTestId('live-clock')]
});
});
You can also apply temporary screenshot CSS to hide or normalize volatile content. This is useful when a selector is easier to target through a stylesheet. Keep the rule local to screenshot capture so it does not change application behavior outside the visual assertion.
const { test, expect } = require('@playwright/test');
test('screenshot uses stable treatment for live content', async ({ page }) => {
await page.goto('http://localhost:3000/dashboard');
await page.getByTestId('dashboard-ready').waitFor({ state: 'visible' });
await expect(page).toHaveScreenshot('dashboard.png', {
stylePath: 'tests/visual-stability.css'
});
});
/* tests/visual-stability.css */
[data-visual-volatile] {
visibility: hidden !important;
}
Playwright’s comparison documentation describes stylePath for filtering volatile elements and screenshot masks. A stylesheet can also affect layout depending on the chosen rule: hiding an element with visibility preserves its space, while display: none changes layout. Pick the behavior that reflects what the test intends to assert. Playwright visual comparisons.
5. Ignore an element in Chromatic
Chromatic supports excluding a specific DOM element from visual diffs with the .chromatic-ignore class or data-chromatic="ignore". For example, mark a live clock:
<time class="chromatic-ignore" data-testid="live-clock">
12:34:56
</time>
Or use the data attribute:
<time data-chromatic="ignore" data-testid="live-clock">
12:34:56
</time>
Chromatic says the ignored pixels include the element’s bounding box and position. Avoid ignoring containers whose size or placement should be checked. Chromatic: ignore elements.
6. Handle animations and delayed resources deliberately
First ask whether motion is part of the behavior being tested. For a settled-state screenshot, disable or complete incidental animation before capture and wait for the resulting UI. For an animation test, keep it observable and assert its behavior separately.
Animation defaults are tool-specific. Chromatic documents that it pauses video and animated GIFs at their first frame; if an animation cannot be disabled, its guidance suggests waiting for completion or ignoring that element. Do not assume another screenshot tool behaves the same way. Chromatic animation handling.
For delayed images, fonts, or data, wait for the condition that proves the needed resource or state is ready. If the intended test checks a loading state, capture that state explicitly rather than waiting for the final page. This keeps the assertion aligned with what the test is meant to protect.
7. Tune comparison thresholds and update baselines carefully
Playwright provides screenshot comparison settings such as maxDiffPixels. A tolerance may account for small, known rendering noise, but a permissive threshold can conceal real visual regressions. Set it based on a known source of harmless variation, and keep it as narrow as practical. Playwright visual comparisons.
When a UI change is intentional, inspect the diff first, confirm the new appearance is correct, and then update the committed reference using the documented snapshot update workflow, for example npx playwright test --update-snapshots. Do not automatically accept every failure: a baseline update is a review decision, not a flake fix.
8. Choosing local assertions, hosted review, or a screenshot API
Playwright’s screenshot assertions keep comparison and reference images in the test workflow. Chromatic describes a hosted workflow that uploads captured archives for cloud comparison and review. Compare tools on the actual needs of the team: where baselines live, how volatile regions are excluded, how animation and delayed resources are handled, browser coverage, and how reviewers approve changes. Confirm current plan limits and supported environments directly before purchasing; those details are not established here. Playwright visual comparisons, Chromatic visual tests.
Percy’s Playwright client documents ignored selector and coordinate regions, as well as animated image options; check the documentation for the package version in use before implementing them. Percy Playwright client library.
For a hosted screenshot API, ScreenshotNeo is an option when you want consent banners, newsletter popups, and chat widgets removed before capture, and billing limited to clean shots. Its response identifies the page verdict and billing status in headers. The API can capture screenshots or PDFs; it is separate from the visual diff and baseline review workflow described above.
Or skip the browser setup
ScreenshotNeo takes a screenshot with one GET request. See the ScreenshotNeo API documentation for parameters and options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. The response includes
X-Page-VerdictandX-Billedheaders. - An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Troubleshooting flaky visual tests
| Symptom | Likely cause | Fix |
|---|---|---|
| Large parts of the page move between runs | Viewport, environment, fonts, or late-loading resources differ | Pin viewport and browser setup; ensure resources and fonts are available before capture; wait for the page state the test needs. |
| A timestamp or counter changes | Live data changes on each run | Use a fixture, seed, or fixed clock where appropriate. If the content itself is irrelevant, mask only that small element. |
| A screenshot sometimes catches a transition | Capture races with animation or a client render | Disable incidental motion or wait for a specific completion condition. Test animation separately if it matters. |
| A mask hides a regression | The masked element’s box includes changed geometry | Reduce the mask to the smallest region, or stabilize the source value and remove the mask. |
| Snapshot update makes the test pass but the cause is unknown | The baseline was accepted without classifying the diff | Review the image difference, confirm the UI change is intended, then update the reference deliberately. |
| Threshold increase hides too many changes | Comparison tolerance is broader than known harmless noise | Lower the tolerance and fix nondeterministic data, timing, or resources at their source. |
| Freezing time did not stabilize the page | Time or data is generated on the server or comes from an external request | Control the server fixture or mock the relevant response; browser-side overrides only affect browser code that reads those APIs. |
FAQ
How do I stop screenshot tests failing because of timestamps or animations?
Fix the timestamp or data source where possible. Disable incidental animation for settled-state captures, or wait for a defined end state. Mask a timestamp only when its value is outside the test’s purpose and its position and dimensions are also irrelevant.
Should I mask a dynamic element or fix the test data?
Fix the data when it can be made deterministic. Mask only unpredictable content that does not belong to the assertion, and remember that a mask can hide the element’s geometry too.
When is it safe to update a visual snapshot?
After reviewing the diff and confirming the new rendering is an intended change. A passing test after an update does not establish that the new baseline is correct.
Can a screenshot API replace visual regression testing?
A screenshot API can capture a page, but visual regression still requires a comparison policy, reference images, and review of meaningful changes. Use the API for capture where it fits your workflow; keep assertions and approvals aligned with the risk you need to catch.


