ScreenshotNeo

BlogHow-to

How to Monitor a Landing Page Screenshot Before and After A/B Tests

Use repeatable Playwright screenshots to catch landing-page regressions around A/B tests, and measure experiment outcomes separately with analytics.

By the ScreenshotNeo team4 October 202611 min read

To monitor a landing page before and after an A/B test, capture each intended page state with a repeatable browser setup, compare it with a saved screenshot baseline, and review every difference against the test hypothesis. Use the visual comparison to catch rendering regressions; use your experiment platform and analytics to decide which variant performs better. A screenshot cannot identify a conversion winner.

This guide uses Playwright Test for local, repository-managed visual baselines. The same principles apply if you use a hosted visual review service. Keep the browser, operating system, viewport, page state, and capture timing consistent between baseline creation and later runs.

1. Decide what the screenshot check should protect

Before writing a test, list the shared regions that should remain visually sound and the parts the experiment deliberately changes. For example, the hero heading and button may differ between variants, while navigation, pricing details, a form, and the footer should remain intact.

  • Record the page URL and the state being captured, including variant, viewport, and any required setup.
  • Identify intentional changes in the experiment hypothesis.
  • Choose whether to capture the whole page or a specific element.
  • Decide how to handle changing content such as timestamps, rotating promotions, personalized text, or live counters.
  • Assign someone to review and approve baseline updates.

Do not treat every pixel difference as a defect. A difference tells you the rendered output changed. Review it to determine whether the change is expected, harmless, or a regression outside the experiment’s intended scope.

2. Create a Playwright screenshot baseline

Install Playwright Test in a project that uses Node.js, then install its browser binaries. The following setup uses Chromium. Keep the Playwright version and operating-system image stable in CI so the baseline and later captures use the same rendering environment.

npm init playwright@latest
npx playwright install chromium

Create tests/landing-page.spec.ts. Replace the example URL with the landing page and make sure the page is in the intended test state before the screenshot assertion runs.

import { test, expect } from '@playwright/test';

test('landing page visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 1000 });
  await page.goto('https://example.com/landing', { waitUntil: 'networkidle' });
  await expect(page).toHaveScreenshot('landing-page.png', {
    fullPage: true,
    animations: 'disabled',
  });
});

Run the test once to create the initial reference screenshot, then run it again to compare the current capture with that reference:

npx playwright test tests/landing-page.spec.ts
npx playwright test tests/landing-page.spec.ts

The first run may report that a new snapshot was written. Review and commit the reference image through your normal version-control process. Later runs compare against the committed reference and report differences for review. Playwright’s screenshot assertion waits for two consecutive screenshots with the same result before comparing; screenshot assertions also disable animations by default. Keep the explicit animation setting if you want that behavior to be clear in the test. See the Playwright visual comparisons guide and toHaveScreenshot API reference.

Capture both variants

A visual baseline must correspond to a reproducible state. If your application can expose deterministic variants in a test environment, capture each one separately and give the files meaningful names. For example:

import { test, expect } from '@playwright/test';

test('control variant', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 1000 });
  await page.goto('https://example.com/landing?variant=control', {
    waitUntil: 'networkidle',
  });
  await expect(page).toHaveScreenshot('landing-control.png', {
    fullPage: true,
    animations: 'disabled',
  });
});

test('treatment variant', async ({ page }) => {
  await page.setViewportSize({ width: 1440, height: 1000 });
  await page.goto('https://example.com/landing?variant=treatment', {
    waitUntil: 'networkidle',
  });
  await expect(page).toHaveScreenshot('landing-treatment.png', {
    fullPage: true,
    animations: 'disabled',
  });
});

The query parameter here is only an example. Use the variant-selection mechanism your application or test setup actually supports. Do not expose a production-only or user-assignment parameter unless your application defines it. If the experiment assigns variants using cookies, set the appropriate test cookie or use a deterministic test account before navigation.

Capture one region instead of the whole page

If the change is localized, a locator screenshot can make the check more focused. It can also reduce noise from unrelated content farther down the page.

const form = page.locator('[data-testid="lead-form"]');
await expect(form).toHaveScreenshot('lead-form.png');

Prefer stable selectors such as test IDs. A selector that matches multiple elements or changes with the page markup can make the test ambiguous or brittle.

3. Make captures repeatable

Rendering may vary with the host operating system, browser version, browser settings, hardware, power source, and headless mode. Playwright recommends running tests in the same environment used to generate the reference screenshots. A baseline captured on a developer laptop may differ from one generated in a Linux CI container even when the page code is unchanged. See Playwright’s guidance on consistent screenshots.

Control the browser and viewport

  • Use the same Playwright and browser versions for baseline creation and comparison.
  • Set the viewport explicitly; do not rely on a machine’s current browser window size.
  • Use the same operating-system image and headless setting in the baseline and comparison jobs.
  • Wait for the page state you intend to inspect. A fixed delay alone is often less reliable than waiting for a known element or application state.
  • Keep device scale factor and other browser context settings consistent if your project sets them.

Handle unstable content carefully

Disable animations where possible. For content that changes on every run but is outside the experiment, use a screenshot style to hide it or mask it. For example, this hides a rotating promotion during the screenshot:

await expect(page).toHaveScreenshot('landing-page.png', {
  fullPage: true,
  animations: 'disabled',
  style: `
    .rotating-promotion,
    [data-testid="current-time"] {
      visibility: hidden !important;
    }
  `,
});

Use a selector that targets only the unstable content. Masking a large area can hide the very regression the test is supposed to catch. If a timestamp or rotating banner is part of the feature under test, keep it visible and instead make its value deterministic in the test environment.

Choose the right wait condition

networkidle can be useful for pages that settle after their requests finish, but it may never occur on pages with polling, analytics, or long-lived connections. In that case, wait for the meaningful page element and, if needed, a short deliberate settling interval:

await page.goto('https://example.com/landing', { waitUntil: 'domcontentloaded' });
await page.locator('[data-testid="hero"] h1').waitFor({ state: 'visible' });
await page.waitForTimeout(300);

Use a fixed delay only to allow a known visual transition to settle after the relevant state is ready. It should not replace waiting for the page’s actual content.

4. Review differences as visual QA

When Playwright finds a changed screenshot, inspect the diff alongside the test hypothesis. Ask whether the changed region is expected, whether shared page regions still render correctly, and whether the capture was made in the same state and environment as the baseline.

  • Expected difference: update the relevant variant baseline after review, with the reason recorded in the change.
  • Unexpected difference within the changed region: check whether the implementation matches the intended design.
  • Unexpected difference in a shared region: investigate it as a possible visual regression, such as a shifted layout, missing form, or broken navigation.
  • Broad differences everywhere: first check browser version, operating system, viewport, fonts, and capture state before changing snapshots.

For hosted snapshot review, Chromatic documents an integration with Playwright, and BrowserStack documents a Percy workflow for Playwright. Choose based on your team’s review process, CI setup, browser needs, retention terms, and existing tools. The cited documentation describes their workflows; it does not establish comparative accuracy or current pricing.

5. Measure the A/B test separately

An A/B test randomly shows two or more versions of a page to samples of users and evaluates an outcome such as click-through, a key event, or engagement. Use your experiment platform and analytics setup for assignment, exposure tracking, metric definitions, and outcome analysis. Google Analytics describes A/B testing as a randomized experiment with two or more page variants and explains that Analytics can be used to interpret results when integrated with a third-party A/B testing tool: Google Analytics A/B test guidance.

A screenshot diff answers “what changed in the rendered page?” It does not answer “which version performed better?” Do not declare a winner from visual comparisons. Keep visual checks and behavioral analysis linked to the same variant names, but use the experiment’s outcome data to make the performance decision.

6. End the experiment cleanly

Test duration depends on traffic and conversion rates. Google Search Central advises running a test only as long as needed for a reliable decision, then updating the site to the desired content and removing test-only elements such as alternate URLs, scripts, and markup. See Google’s A/B testing guidance for Search.

  • Record the selected variant and the visual baseline that represents the desired final page.
  • Remove temporary variant-selection code, alternate URLs, test scripts, and markup when no longer needed.
  • Keep useful visual regression checks for the final page and shared components.
  • Verify that redirects and canonical behavior are correct if the test used alternate URLs.

Google Optimize and Optimize 360 are no longer available; Google states they were discontinued on September 30, 2023. Do not follow legacy instructions that depend on those products for a new workflow. See Google Ads Help’s notice.

Screenshot capture options and services

For this workflow, the main choice is whether to keep the visual baseline in your repository, use hosted snapshot review, or capture a page on demand through an API. Match the tool to the job: Playwright assertions can compare repeatable test captures; hosted services can provide a review workflow; a screenshot API can return a rendered image without requiring you to maintain browser setup for that capture.

ScreenshotNeo is a website screenshot API and MCP server. Its API supports full-page screenshots, element capture, device and viewport options, custom CSS and JavaScript, waits, caching, and PDF output. It can help when you need screenshots on demand or want an AI agent using an MCP client to capture a page. Its API does not replace your experiment platform or analytics, and an API screenshot by itself is not a saved-baseline comparison workflow. See the ScreenshotNeo documentation for request options.

Or skip the browser setup

To request a screenshot with one GET call, replace the sample URL and use your API key. This saves the returned response body as a WebP image. See the ScreenshotNeo API documentation for options and response headers.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month with no card.

Performance, reliability, and cost

  • Local Playwright: screenshot assertions run as part of browser tests, so account for browser installation, capture time, and snapshot review in CI. Reusing a stable CI image reduces environment-driven diffs.
  • Snapshot maintenance: every intentional visual change may require a reviewed baseline update. Keep snapshots focused and remove obsolete variant baselines when the experiment ends.
  • Page readiness: waiting too little can capture incomplete content; waiting for network idle on a page that never settles can stall the test. Prefer a meaningful selector or application-ready signal.
  • Hosted review: evaluate the service’s account terms, retention, CI behavior, and review process for your needs. The sources linked above document integrations, not current prices or comparative performance.
  • Screenshot API: ScreenshotNeo lists a free tier of 1,000 shots per month and paid tiers from $5 for 3,000; higher plans are Starter $5/3,000, Growth $15/15,000, Pro $39/60,000, Scale $99/250,000, and Business $249/1,000,000. Yearly billing gives two months free. Every feature is on every plan. Only clean shots are billed; inspect response headers for the verdict and billing status.

Troubleshooting

The test fails with a new or unexpected screenshot

Cause: the page changed, the baseline is stale, or the environment differs. Fix: inspect the image and diff; verify browser and operating-system versions, viewport, fonts, and variant state. Update the reference only after confirming the change is intended.

The entire page differs between local and CI runs

Cause: host rendering differences, browser version drift, headless settings, device scale factor, or missing fonts. Fix: create and compare baselines in the same CI image with pinned browser dependencies and explicit viewport settings.

The screenshot captures a loading state or misses lazy content

Cause: the assertion runs before the relevant content is ready or before lazy-loaded content enters the viewport. Fix: wait for a page-specific ready selector; for long pages, scroll through the content before capturing if the application only loads it on scroll. Check the resulting image rather than assuming navigation completion means every visual asset is ready.

The test hangs waiting for network idle

Cause: polling, analytics, streaming, or another persistent request prevents the network from becoming idle. Fix: navigate with domcontentloaded or another suitable lifecycle event, then wait for the page element that indicates the tested state is ready.

The screenshot changes on every run

Cause: animations, clocks, random content, rotating banners, live data, or personalization. Fix: disable animations, inject deterministic test data, or hide/mask only the unstable area with screenshot styling. Keep meaningful experiment content visible.

The diff flags an intentional A/B change

Cause: the two variants have distinct expected visuals, but the test compares one with the other’s baseline. Fix: keep separate named baselines for each reproducible variant and review each against its own expected state.

The screenshot API returns an error or an unexpected page

Cause: the target may require authentication, may block automation, may be blank, or may fail to load. Fix: check the target URL and access requirements, inspect the response status and ScreenshotNeo’s X-Page-Verdict and X-Billed headers, then consult the API documentation for supported request options. A failed or blank capture is not a visual baseline to approve.

FAQ

Can a screenshot comparison tell me which variant won?

No. It can show visual changes, but the winner depends on the experiment’s behavioral outcome metrics and analysis.

Should I compare the treatment directly with the control?

Compare each variant with its own expected baseline when checking for regressions. Review control-versus-treatment differences separately to confirm the intended change.

Can I use screenshots from different machines?

You can, but environment differences may create noise. Keep baseline generation and comparison on the same operating system, browser version, settings, and viewport.

Is Google Optimize still available?

No. Google Optimize and Optimize 360 were discontinued on September 30, 2023.