ScreenshotNeo

BlogHow-to

Loki screenshot tests are flaky: how to stabilize them

Stabilize Loki screenshot tests by making stories deterministic, pinning the capture environment, choosing a suitable diff engine, and reviewing baseline changes.

By the ScreenshotNeo team4 October 20268 min read

A flaky Loki screenshot test is often a determinism problem in the story or rendering environment, not an image-comparison problem. Make asynchronous work finish explicitly, stop changing motion or content, and capture baselines and tests in the same pinned browser environment. Only then tune the diff engine or tolerance.

This guide follows Loki’s documented behavior and options. CLI flags and defaults can vary by installed version, so check your version’s help and configuration before copying version-specific settings.

1. Make story readiness explicit

A story that fetches data, computes state on mount, or re-renders after an event may be captured before its final state is ready. Loki handles most network traffic and image loading, but other asynchronous work may need an explicit completion signal. Use Loki’s asynchronous callback pattern in the story so capture waits for the work you control.

// Illustrative pattern: adapt the callback signature to the Loki version in your project.
export const LoadedState = {
  render: (args, { done }) => {
    loadFixture().then((data) => {
      renderLoadedComponent(data);
      done();
    }).catch((error) => {
      reportStoryError(error);
      done(error);
    });
  },
};

The names and Storybook integration depend on your setup; use Loki’s documented async story callback for your installed version. The key is to call the completion callback only after the final render is ready. Avoid replacing it with a fixed sleep: a delay may be too short on a slow runner and unnecessarily long on a fast one.

Keep fixtures stable as well. Freeze timestamps, random values, generated IDs, and data ordering where they affect visible output. Prefer local fixture data over a live service when the screenshot is meant to verify layout rather than service availability.

2. Remove motion and changing content

Loki disables common CSS transitions and animations and requestAnimationFrame by default. Its documented limits include looping requestAnimationFrame, GIF, animated SVG, Lottie, and React Native Animated content. These can still show different frames from one capture to another.

  • For animated content, provide a test-only static state or disable the animation in the story.
  • Use Loki’s documented context/decorator approach or Loki-targeted CSS where suitable.
  • If a story cannot be represented meaningfully by one still frame, selectively skip that story rather than weakening assertions across the suite.

Do not hide a whole component just because its animation is difficult to stabilize if the component’s static appearance is part of what the test should cover.

3. Pin the capture environment

Browser rendering can vary with the host operating system, browser version, settings, hardware, power source, and headless mode. Generate reference images and run CI captures in the same controlled environment. Pin the browser build, OS image, headless mode, viewport, fonts, and relevant rendering settings.

  1. Build Storybook as a static site for CI and point Loki at that build.
  2. Use a versioned, compatible browser or Docker image for both baseline generation and CI.
  3. Keep viewport dimensions and device scale settings consistent.
  4. Install the same fonts and avoid relying on system fonts that differ between machines.

Loki’s CLI provides a Docker image option, and its configuration supports named targets and viewport dimensions. Its documentation includes an example image, but verify compatibility with your installed Loki and browser versions before using an older sample unchanged.

4. Check load failures and capture boundaries

Loki fails a test by default when a story makes a network request that fails. Investigate failed requests first. Use fetchFailIgnore only for a known, intentional URL pattern whose failure is part of the test setup; do not use it to hide unexplained errors.

Check the capture region when a screenshot appears to include unexpected content. Loki’s chromeSelector crops to the selector’s dimensions; it is not a DOM-only screenshot. Absolutely positioned elements above the selected area may still appear in the captured image.

5. Choose a diff engine using representative changes

Loki documents three diff engines. Their sensitivity, speed, tolerance behavior, and dependencies differ. Compare them with both known intentional changes and known accidental regressions before changing engines or thresholds.

Engine Documented trade-offs Consider it when
pixelmatch More sensitive to changes on large images and less susceptible to anti-alias flakiness. You want pixel-level comparison and have reviewed how its sensitivity behaves on your actual screenshots.
gm Generally faster, requires GraphicsMagick, and bases tolerance on the overall image. The external dependency is acceptable and your visual diffs suit overall-image tolerance.
looks-same JavaScript-only and slower; measures tolerance across neighboring pixels, which may help with differing pixel densities. You need its neighboring-pixel behavior and accept the runtime trade-off.

There is no universal safe tolerance established by these sources. Choose one from real diffs, and verify that representative layout, color, and content regressions still fail.

6. Use retries and tolerance narrowly

Loki’s CLI documentation lists chromeRetries with a default of zero, chromeLoadTimeout with a default of 60,000 milliseconds, and chromeTolerance defaulting to zero. Tolerance behavior depends on the selected diff engine. Confirm the current defaults for your installed version.

First fix state, timing, and environment differences. A retry can help with an occasional capture failure, but persistent pixel differences point to a nondeterministic fixture or rendering setup. Keep tolerance small and document the known noise it is intended to exclude. A larger threshold can allow genuine visual regressions through.

7. Treat baselines as reviewed code

Loki separates testing, updating, and approving reference images, and stores references, current screenshots, and diffs in configurable directories. In CI, build a static Storybook and require references so missing baselines cannot make the result misleading.

  1. Run the test and inspect expected, actual, and difference images.
  2. Decide whether each change is intended by reviewing the source change and the visual result.
  3. Update and approve references only for accepted changes.
  4. Commit reviewed reference updates with the code change that explains them.

Playwright’s visual comparison guidance likewise uses explicit snapshot updates and recommends reviewing and committing screenshot changes. A baseline update is an assertion change: it should receive the same care as changing expected test data.

8. Troubleshoot common failures

Symptom Likely cause What to do
The diff changes between repeated runs Async work, animation, timestamps, randomness, or live data is still changing. Add an explicit readiness signal; freeze visible inputs; disable or statically render motion.
Local screenshots pass but CI differs Different OS, browser, fonts, headless mode, viewport, or rendering settings. Pin and share the capture image and browser configuration for baseline generation and CI.
A story fails with a network error A request from the story failed. Fix the request or use a stable fixture. Ignore only a deliberate, understood URL pattern.
Images or data are missing in the capture Capture occurred before non-network async work completed, or the asset failed to load. Signal completion after the final render and inspect failed requests and asset loading.
A crop still contains content outside the target chromeSelector crops screenshot dimensions; positioned elements can overlap the crop. Adjust the story layout or capture target and inspect positioning and overflow.
CI reports a missing reference or behaves unexpectedly References were not created, included, or required consistently. Build a static Storybook, require reference images in CI, and follow the update/approve workflow.
Changing tolerance makes failures disappear The threshold may be masking a real change. Revert the broad threshold, inspect representative diffs, and tune only for understood rendering noise.
Retrying sometimes passes but does not resolve the issue A transient capture failure may coexist with persistent nondeterminism. Keep retries for occasional failures only; stabilize story state and environment for changing pixels.

9. Performance, reliability, and maintenance

Faster diffing is useful only if it still catches the changes your suite is meant to detect. The documented gm engine is generally faster but adds a GraphicsMagick dependency; looks-same is slower and JavaScript-only. Measure on your representative suite rather than assuming an engine will be faster for your project.

Reliability comes primarily from repeatable inputs and a repeatable renderer. Static Storybook builds, local fixtures, pinned browser images, consistent fonts and viewport settings, and explicit async completion reduce sources of variation. Retries may recover occasional capture failures, but they do not make changing UI state deterministic.

Keep baselines scoped to the environment that generated them. When changing the browser, OS image, fonts, or diff engine, expect that references may need a deliberate review. Record the environment and reason alongside baseline changes so later updates are explainable.

10. Do-it-yourself checklist

  • Does every story with async work signal completion after its final visible render?
  • Are clocks, random values, IDs, and fixture ordering stable?
  • Are animations disabled or represented by a deliberate static state?
  • Do local reference generation and CI use the same browser, OS image, fonts, viewport, and headless settings?
  • Are failed requests investigated, with ignore patterns limited to known cases?
  • Have diff engine and tolerance choices been checked against intentional and accidental changes?
  • Does CI require references, and are all updates reviewed as code?

Or skip the browser setup

If your task is capturing a website rather than asserting Loki story snapshots, ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. See the API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot, and each step can be turned off. Bot checks, blank pages, and failed loads are never billed; cache hits are also free. Claude, Cursor, and other MCP clients can use its screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month with no card.

FAQ

Should every story be included in visual tests?

No. Include stories that have a meaningful stable still state. Selectively skip content that cannot be represented as a useful still image, and keep the rest of the suite deterministic.

Can I fix flakiness by switching to a more tolerant engine?

Not by itself. First establish whether the input state and rendering environment are stable. Then compare engines on real diffs and choose a narrowly justified tolerance.

Does disabling requestAnimationFrame stop every animation?

No. Loki documents limitations for looping frame callbacks and formats or libraries such as GIF, animated SVG, Lottie, and React Native Animated. Handle those explicitly in the story or test styling.

When should I update a reference image?

After confirming the visual change is intended and reviewing the expected, actual, and diff images. Treat the reference change as a reviewed modification to the test’s expected output.