ScreenshotNeo

BlogHow-to

How to Compare Website Screenshots with Loki Visual Monitoring

Use Playwright to compare website screenshots with approved baselines, then use Loki logs to investigate failures. Loki does not compare image pixels.

By the ScreenshotNeo team4 October 202611 min read

Short answer: use a browser test such as Playwright Test to capture a page and compare it with a reviewed screenshot baseline. Use Grafana Loki to query application and test logs from the same run and time window when a comparison fails. Loki is a log aggregation system; it does not compare screenshot pixels. Keep screenshots and visual diffs as test artifacts, and correlate them with Loki using run identifiers, timestamps, and links.

This distinction makes “Loki visual monitoring” a useful description of a workflow, but not a built-in Loki image-diff feature. Playwright documents screenshot baselines and visual assertions; Grafana documents Loki for log aggregation and querying. Playwright visual comparisons · Grafana Loki documentation

1. Decide what each system should answer

System Input Question it answers Evidence
Playwright visual test Browser-rendered screenshot and approved baseline Did the rendered page change beyond the comparison policy? Baseline, actual screenshot, diff, test result
Loki and Grafana Timestamped log streams with labels What errors or events occurred around the test or incident? Log lines, timestamps, labels, queries, dashboards, alerts

Loki indexes a limited set of stream labels and stores log data in chunks. LogQL selects streams and filters or processes their log lines. This is useful for investigating a failed screenshot test, provided relevant application or test events are being sent to Loki. The log query does not inspect the PNG or calculate a visual diff. See Loki’s overview and LogQL query documentation.

2. Add a Playwright screenshot comparison

The example below captures a stable route at a fixed viewport, waits for a page-specific ready condition, masks a volatile timestamp, and compares against a baseline. It is a complete Playwright Test file. Install Playwright Test and its browser in your project, save this as tests/home.visual.spec.ts, and set BASE_URL to the application under test.

import { test, expect } from '@playwright/test';

test('home page matches its visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 900 });
  await page.goto(process.env.BASE_URL ?? 'http://127.0.0.1:3000', {
    waitUntil: 'domcontentloaded',
  });

  // Prefer an application-specific readiness signal over an arbitrary sleep.
  await page.locator('main').waitFor({ state: 'visible' });

  // Mask a known volatile region only if its content is not under visual test.
  await expect(page).toHaveScreenshot('home.png', {
    fullPage: true,
    animations: 'disabled',
    mask: [page.locator('[data-testid="live-clock"]')],
    maskColor: '#999999',
  });
});

Playwright’s first run creates a reference image if one is missing; inspect that generated image before accepting it into version control. On later runs, the assertion compares the current capture to that baseline. Commit the intended snapshots with the test so reviewers can see baseline changes in code review. Read Playwright’s visual comparison guide for snapshot naming, storage, and updates.

Minimal project setup

In a Node.js project, install the test runner and browser, then run the named test. The first run is baseline generation and review; the next run performs the comparison.

npm install --save-dev @playwright/test
npx playwright install chromium
npx playwright test tests/home.visual.spec.ts --project=chromium
# Review the generated snapshot, then commit it.
# On subsequent runs, the same command checks for visual changes.

If the repository does not yet have a Playwright config, a small config can pin the project to Chromium and set the base URL. Keep the browser version and execution environment stable between baseline creation and CI comparisons.

import { defineConfig } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  use: {
    browserName: 'chromium',
    baseURL: process.env.BASE_URL ?? 'http://127.0.0.1:3000',
    viewport: { width: 1440, height: 900 },
  },
  expect: {
    toHaveScreenshot: {
      animations: 'disabled',
      // Set tolerances only after reviewing actual diffs.
      // maxDiffPixels: 20,
    },
  },
});

If you use baseURL, the test can navigate to a relative route such as page.goto('/'). The explicit URL in the first example also works without that config.

3. Make screenshots reproducible

Visual comparisons are meaningful only when the inputs are controlled. Playwright warns that screenshots can vary with the operating system, browser version and settings, hardware, power source, and headless mode. Run comparisons in the same environment used to create the baseline, ideally through the same CI image and browser project. Playwright recommends matching the baseline environment.

  • Route and data: use deterministic fixtures or seeded test data. Avoid pages whose content changes between runs unless that behavior is what you want to test.
  • Viewport and device scale: pin viewport dimensions and browser project. A different viewport changes responsive layout; a different scale can alter rasterization.
  • Authentication and state: create a known logged-in state and consistent cookies, local storage, and feature flags. Keep secrets out of committed snapshots and logs.
  • Fonts and assets: wait for the page’s actual ready condition and ensure required fonts and images have loaded. A missing font can shift wrapping and layout across the entire page.
  • Animation and time: disable animation for captures when animation is irrelevant; mask or freeze clocks and rotating content only when those regions are not under test.
  • Consent and overlays: establish a deliberate consent state for the test. A banner or dialog can obscure the page and create a large diff if its state varies.
  • Lazy content: scroll or otherwise trigger the relevant lazy-loaded regions before a full-page screenshot when the application requires it.

Do not hide every changing region automatically. A rotating advertisement may be noise for a layout test; a changing price, status, or user-facing alert may be exactly the behavior that deserves coverage. Mask only after deciding what the test is intended to protect.

4. Set a diff policy and review baselines

Playwright supports a maximum differing-pixel count, a maximum differing-pixel ratio, and a perceived color threshold. These controls answer different questions: a pixel count gives an absolute allowance, a ratio scales with image size, and the color threshold controls how much color difference a pixel comparison treats as significant. Options may be set on an assertion or in the project’s expect.toHaveScreenshot configuration. Consult the snapshot assertion options and visual comparison docs for their precise semantics.

await expect(page).toHaveScreenshot('dashboard.png', {
  fullPage: true,
  maxDiffPixels: 25,
  // maxDiffPixelRatio: 0.001,
  // threshold: 0.2,
});

The values above illustrate where the settings go; they are not universal recommendations. Start with strict comparison on stable pages. When a test fails, inspect the expected image, actual image, and diff before changing the tolerance. Increasing a threshold to silence unexplained failures can hide real layout or rendering regressions.

  1. Generate a baseline from an intentional, reviewed page state.
  2. Run the same test in the pinned environment.
  3. Inspect the diff and decide whether the change is intended, a regression, or capture noise.
  4. For an intended design change, update the baseline deliberately with npx playwright test --update-snapshots, inspect the changed files, and commit them alongside the UI change.

The update command replaces expected snapshots; it does not decide whether the new design is correct. Baseline review remains a human decision.

5. Correlate a failed screenshot test with Loki

Give each CI run a stable identifier, and include it in test-run logs or structured application context where appropriate. Useful correlation fields include a run identifier, build or commit version, environment, service name, route, and event timestamp. Keep Loki stream labels low-cardinality and stable; put per-run or per-request detail in log fields or line content instead of creating a new label value for every event.

After a screenshot assertion fails:

  1. Open the Playwright report and inspect the expected, actual, and diff images first.
  2. Record the test’s run identifier, build version, route, and capture time.
  3. In Grafana Explore, select the relevant Loki stream labels, then narrow the time range to just before and after the capture.
  4. Filter for the run identifier or relevant error text, and inspect request failures, application exceptions, or test setup events that could explain the rendered state.
  5. Keep the screenshot and diff with the CI/test artifacts. Link to those artifacts from the test report or run context; use Loki for log evidence and correlation.

A conceptual LogQL query looks like this; replace the labels and text filter with the labels your deployment actually sends:

{service_name="web"} |= "visual_run_id=RUN_ID"

LogQL syntax and available labels depend on your log pipeline. The query only finds matching log lines. A screenshot file does not become searchable in Loki because its test emitted a log entry; image artifact storage and retention require a separate workflow. See Loki query documentation.

6. Understand what screenshot alerts do and do not mean

A screenshot test failure means the current rendered page differed from its accepted baseline under that test’s capture conditions and comparison policy. It does not, by itself, prove that the application logged an error: a CSS shift can be a real visual regression with no exception. Conversely, an error in Loki may not change the screenshot if the affected feature was outside the captured route or state.

Grafana alert notification screenshots are a separate capability concerning Grafana panels associated with alert rules. They do not provide website screenshot comparison, and the documentation says notification screenshots are not supported in Loki. See Grafana notification documentation and check the screenshot support statement in the relevant Grafana alert screenshot documentation before designing around it.

7. Troubleshooting

Symptom Likely cause Fix
Snapshot missing on first run No reference image exists yet. Review the generated actual image as the proposed baseline; commit it only after approval.
Large diff on every run Browser, OS, fonts, viewport, device scale, or runtime environment differs. Pin the browser project and run baseline generation and comparison in the same environment.
Intermittent diff in one region Animation, clock, rotating content, async data, or third-party embed is changing. Use deterministic test data, disable irrelevant animation, wait for a meaningful ready signal, or narrowly mask content outside the test’s scope.
Blank or partially loaded screenshot Navigation completion was mistaken for application readiness, or an asset/request failed. Wait for a visible app-specific element; inspect browser errors and correlated application logs. Avoid relying on arbitrary fixed sleeps as the main readiness mechanism.
Diff appears after a legitimate UI change The expected baseline describes the previous design. Review the actual and diff, then update snapshots intentionally and include them in the same reviewed change.
Diff passes despite a visible minor change Configured pixel or color tolerance is too permissive for the page. Inspect the comparison settings and lower the allowance based on the page’s risk; do not copy a threshold from another project blindly.
Loki query returns no related lines Logs may not include the run identifier, labels may differ, or the selected time range may miss the event. Confirm ingestion and actual label names, widen the time range, and add run/build context to future test or application logs.
Loki query is slow or expensive The query selects too many streams or scans a wide time range. Use selective stable labels first, limit the time window to the capture, then apply text filters.
New CI workers fail against old snapshots The image, browser, fonts, or rendering environment changed. Reproduce in the original environment if possible; if the environment change is intentional, review and regenerate affected baselines as a controlled migration.

8. Performance, reliability, and cost considerations

  • Capture cost in CI: each screenshot requires browser rendering and comparison. Prefer targeted routes and states over taking redundant full-page captures of every route at every configuration.
  • Full-page tradeoff: full-page images cover more content but take more capture and review effort, and can expose more dynamic regions. Use element screenshots when the component is the behavior under test.
  • Parallel execution: parallelize independent tests while keeping each test’s data and browser context isolated. Avoid having multiple jobs update the same baselines concurrently.
  • Baseline storage: snapshots grow with pages, states, browsers, and viewport variants. Keep only variants that answer a meaningful support question; review repository and artifact retention policies.
  • Flakiness: retries can help reveal intermittent failures, but they should not turn unexplained visual instability into a passing signal. Track repeated diffs and remove the source of nondeterminism.
  • Loki query efficiency: choose useful stream labels and a narrow time range before applying broad text filters. Logs improve diagnosis only when relevant events are collected and correlated.
  • Reliability boundary: screenshots show a rendered result at a point in time; logs show recorded events. Neither guarantees the other is complete. Keep the test result and screenshot artifact as primary evidence for the visual assertion.

This workflow uses open-source test and logging components, but its infrastructure and operating costs depend on your own CI execution, artifact retention, and Loki deployment. The research sources provide no universal performance or cost figures for a given application.

Or skip the browser setup

If you need a screenshot without installing and maintaining browser capture code, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. It is a capture service, not a replacement for reviewed visual baselines or Loki log investigation. See the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
  writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))
);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. Sign up for 1,000 free screenshots a month, no card required.

FAQ

Can Loki compare two PNG files?

No. Loki aggregates and queries logs. Use a browser visual test or a dedicated image comparison workflow to compare screenshot pixels.

Should screenshot images be stored in Loki?

The documented workflow treats screenshots and diffs as test artifacts. Image storage alongside an observability system would require a separately designed artifact workflow.

Does a visual diff mean the application is broken?

Not necessarily. It may show an intended design change, a rendering-environment difference, volatile content, or a genuine regression. Review the actual image and diff before changing the baseline.

Can Grafana alert screenshots replace browser visual tests?

No. Grafana alert screenshots relate to Grafana panels and alert notifications, not website screenshot baselines. Grafana documentation also notes that this feature is unsupported in Loki.

References