ScreenshotNeo

BlogHow-to

How to Fix Different Argos CI Screenshots on Local and CI Runs

Find out why Argos screenshots differ between local and CI runs, then stabilize page state and use one consistent Playwright environment.

By the ScreenshotNeo team4 October 202610 min read

Argos screenshots can differ between local and CI runs for two broad reasons: the page was in a different state when it was captured, or the machines rendered the same page differently. Stabilize the page first, then capture and update baselines in one canonical environment with a pinned operating system, Playwright version, and browser. Use a per-screenshot tolerance only for understood residual noise; a broad threshold can hide real regressions.

A pixel diff proves that pixels changed, but not that the product regressed. Fonts, images, animations, live data, timestamps, random ordering, viewport settings, and machine rendering are common causes. Argos compares screenshots captured in the real test browser; using Argos does not remove the need to make the page deterministic.

1. Classify the screenshot difference

Before changing thresholds or updating snapshots, inspect the diff and identify its shape. These patterns are clues rather than guarantees:

What changed Likely cause to investigate
Text wraps differently, or font weight changes across much of the page A web font failed to load, loaded late, or rendered differently on the two systems.
A whole image is blank, different, or appears late The asset request failed, the image had not decoded, or lazy loading had not been triggered.
Only a spinner, animation, timestamp, or rotating widget differs The capture happened at a different point in time or uncontrolled data was included.
Lists show the same records in a different order The data or query lacks a stable order, or random data was not seeded.
Most edges have small pixel changes, with little layout movement Font rasterization, anti-aliasing, GPU paths, or scrollbar behavior may differ by machine.
The whole layout shifts or wraps at different points Viewport dimensions, device scale factor, browser version, or scrollbar width may differ.

Open the Argos diff alongside the actual page and, where useful, a Playwright trace. First establish whether the changed pixels represent unstable page content or an environment mismatch. Do not use a threshold to make an unexplained difference disappear.

2. Make the page deterministic before capture

Use a stable test fixture and wait for the visible state you intend to compare. A fixed delay by itself is not a reliable readiness check: slow CI workers can need longer, while fast runs waste time.

Wait for fonts and images

Wait for web fonts and image decoding before taking a screenshot. Make sure the same font files are available in CI, and check CI network logs for font or image errors. If images are lazy-loaded, scroll them into view or deliberately trigger the loading behavior before capture.

// Wait for browser font loading and images already present in the document.
await page.evaluate(async () => {
  await document.fonts.ready;
  const images = Array.from(document.images);
  await Promise.all(images.map(async (image) => {
    if (!image.complete) {
      await new Promise((resolve) => {
        image.addEventListener('load', resolve, { once: true });
        image.addEventListener('error', resolve, { once: true });
      });
    }
    if (image.decode) {
      try { await image.decode(); } catch { /* A failed image is diagnosed separately. */ }
    }
  }));
});

This waits for images currently represented in the DOM; it does not make off-screen lazy images load. Scroll through the page or target the relevant elements first, then wait for their image requests and decoding.

Stop animation and blinking carets

Disable CSS transitions and animations for visual tests. Also account for JavaScript-driven animation, canvas, video, and blinking input carets; CSS rules cannot freeze application timers or every canvas renderer. Control those in the app or test fixture.

await page.addStyleTag({ content: `
  *, *::before, *::after {
    animation-delay: 0s !important;
    animation-duration: 0s !important;
    animation-iteration-count: 1 !important;
    scroll-behavior: auto !important;
    transition-duration: 0s !important;
    caret-color: transparent !important;
  }
` });

Apply this before capture and verify that the resulting state is representative. If the page depends on animation to reach its final state, wait for that state explicitly before disabling or freezing animation.

Control APIs, clocks, randomness, and ordering

  • Mock live API responses with fixtures so the screenshot uses the same records on every run.
  • Freeze dates and times when relative labels or timestamps appear.
  • Seed random values used for visible content, or replace them with fixed test values.
  • Sort records explicitly and use a stable tie-breaker, such as an ID, when values can be equal.
  • Hide or stub third-party widgets that cannot be made deterministic, while keeping the test focused on the UI you own.

Network idle can help identify a stable point, but it is not a universal guarantee: pages that poll, stream, or keep long-lived connections may never become idle. Prefer an app-specific readiness condition, such as a loaded heading or an absent busy indicator, and use network idle only when it fits the page.

3. Pin the capture environment

macOS and Linux can differ in font rasterization, anti-aliasing, GPU paths, and scrollbar or viewport behavior. A practical rule for committed visual baselines is to capture and update them in the same pinned Playwright image and browser version used in CI. Keep the viewport and device scale factor fixed as well.

Argos’s CI guide recommends using an official Playwright Docker image pinned to the exact Playwright version, and updating snapshots inside that image or in CI. Check the current guide and image tags when setting up a workflow; version numbers in examples age. See Argos’s Playwright visual regression testing in CI guide.

Argos’s quickstart also demonstrates Chromium flags --disable-lcd-text and --font-render-hinting=none to stabilize text rendering across macOS and CI. Treat these as rendering controls to evaluate in your environment, not as a substitute for loading the intended fonts. See the Argos Playwright quickstart.

4. Configure Argos with Playwright

Argos’s documented Playwright integration uses @argos-ci/playwright, its reporter, and the argosScreenshot helper. Install the package and Playwright in your project, then configure the reporter and upload only in CI:

// playwright.config.ts
import { defineConfig } from '@playwright/test';

export default defineConfig({
  reporter: process.env.CI
    ? [['list'], ['@argos-ci/playwright/reporter']]
    : [['list']],
  use: {
    trace: 'on-first-retry',
    screenshot: 'only-on-failure',
  },
});
// tests/homepage.spec.ts
import { test, expect } from '@playwright/test';
import { argosScreenshot } from '@argos-ci/playwright';

test('homepage visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1280, height: 800 });
  await page.goto('http://127.0.0.1:3000', { waitUntil: 'domcontentloaded' });

  // Prefer an app-specific ready condition in addition to document readiness.
  await page.getByRole('heading', { name: 'Welcome' }).waitFor();
  await page.evaluate(() => document.fonts.ready);

  await argosScreenshot(page, 'homepage');
});

Use your actual local test URL and a readiness condition that exists in your app. The Argos helper documents waiting for fonts, decoded images, and network idle; it also waits for [aria-busy] loaders to disappear, hides carets and scrollbars, and pauses CSS animations. Check the behavior in the version you install and do not assume it controls your app’s JavaScript timers, data, or canvas.

In CI, provide ARGOS_TOKEN as a secret, or use the GitHub Actions OIDC/tokenless authentication option described by Argos. The quickstart’s GitHub Actions example uses npm ci and npx playwright install --with-deps chromium; verify current action and image versions against the official docs. A default-branch build must upload at least one build before pull requests can compare against a baseline; otherwise Argos describes those pull requests as orphaned.

Retries and failure screenshots are useful diagnostics. They do not make nondeterministic rendering deterministic. Keep traces and screenshots on failure to investigate functional test errors, while addressing visual instability at its source.

5. Choose where baselines are generated and updated

For a small suite, Playwright’s native toHaveScreenshot with committed PNG baselines can be a straightforward option. Generate and update those files in the same Docker environment as CI. For hosted review, Argos uploads screenshots captured in the real test browser for diffing and review. In either approach, page stabilization and environment parity remain necessary.

Choose based on rendering parity, how baselines are stored and updated, the review workflow your team needs, repository size, and whether local captures or hosted review fit your process. Do not mix local macOS baselines with Linux CI captures and expect pixel-identical output.

6. Tune comparison only after stabilizing

Keep strict comparison for ordinary content. If a specific canvas, map, or remaining anti-aliasing noise is understood, tune tolerance only for that screenshot or region where the tool supports it. Then confirm that meaningful changes to text and layout still fail review.

A global threshold can mask a real UI regression across unrelated screenshots. Argos’s stabilization guide gives vendor-specific threshold details; check its current documentation before relying on a setting. The goal is to account for known residual noise, not unexplained changes.

7. Troubleshooting checklist

Symptom Cause to check Fix
Fonts or line breaks differ Font files are missing, return an error in CI, or have not loaded before capture; OS rasterization also differs. Check network responses and computed fonts, wait for document.fonts.ready, and capture on the same pinned OS/browser. Consider Argos’s documented Chromium text-rendering flags.
Images are blank or inconsistent Requests failed, lazy images were not triggered, or decoding had not completed. Inspect request failures, scroll target images into view, wait for loading and decoding, and use stable fixtures for remote assets.
Only a spinner or moving element differs The screenshot landed at a different animation or loading point. Wait for an explicit ready state, disable CSS animation, and control application timers, canvas, or video separately.
Text, prices, or dates change between runs Live data or the system clock is visible. Mock the API and freeze the clock or supply fixed test data.
Rows move around between runs Records are returned in unstable order or generated randomly. Use deterministic fixtures, seed random values, and sort with a stable tie-breaker.
CI and local screenshots have widespread small edge changes Different OS, browser build, font libraries, GPU path, or scrollbars. Use one pinned capture environment. Do not try to fix a machine mismatch with a global tolerance.
The entire page wraps differently Viewport size or device scale factor is different, or content width changed. Set the same viewport and scale factor; verify browser version and scrollbar behavior.
Argos pull request build has no baseline No default-branch build has uploaded screenshots yet. Run the default branch through the configured reporter first, then retry the pull request comparison.
Increasing tolerance makes tests pass but misses visible changes The threshold is broad enough to mask meaningful differences. Restore stricter comparison, stabilize the page, and tune only the affected screenshot for understood residual noise.

8. Performance, reliability, and cost considerations

Readiness waits improve reliability but can add time. Prefer waiting for the actual fonts, images, and app state needed by a screenshot over adding a long fixed sleep to every test. Reuse deterministic fixtures and avoid waiting for network idle on pages that never settle. Pinning the CI image reduces environment drift, but remember to update the Playwright package, browser image, and baseline together when you intentionally upgrade.

Visual review has a maintenance cost: baseline changes need review, and unrelated environmental noise consumes attention. Keep tests focused on meaningful UI, isolate uncontrolled third-party content, and update baselines only after reviewing the intended change. Hosted diffing changes the review workflow; it does not guarantee stable captures or eliminate the work of approving changes. Argos-specific plans and thresholds can change, so consult its current product pages for current terms.

Or skip the browser setup

If you need a clean page capture without wiring up a browser locally, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its one-call API returns a PNG, JPEG, WebP, or PDF. This example requests a WebP capture; see the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as image:
    image.write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', image));

ScreenshotNeo accepts cookie and consent banners as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month with no card.

FAQ

Should I update the baseline when CI fails?

Only after reviewing the diff and confirming the change is intentional. A failing diff may reveal a real regression or an unstable capture; updating it without diagnosis can make either harder to catch.

Can retries fix a flaky visual test?

Retries can expose intermittent failures and preserve diagnostic artifacts, but they do not make the page deterministic. Fix the unstable input or capture environment.

Does using Argos make local and CI rendering identical?

No. Argos reviews screenshots captured in the test browser. You still need stable page content and a consistent environment if you compare local captures with CI baselines.

What is the most important first change?

For committed baselines, capture and update in the same pinned Playwright environment used in CI, while ensuring the page reaches a deterministic state before the screenshot.

Sources: Argos Playwright Quickstart; Argos, “How to Fix Flaky Visual Tests: Every Root Cause, Solved”; Argos, “Playwright Visual Regression Testing in CI: Complete Guide”.