How to Run Visual Regression Tests Across Multiple Branches
Set branch-aware visual baselines, compare pull requests with the right reference, and keep screenshot results stable across CI.
Run visual regression checks on both pull requests and the shared integration branch, and decide which baseline each check uses. A branch regression test asks whether the UI changed since its approved visual state. A pull request review against the merge base asks what the branch introduces. Those are different questions, so choose the comparison mode deliberately, review intentional changes before accepting them, and keep rendering inputs consistent.
This guide covers Playwright’s committed screenshot snapshots, Chromatic’s branch baselines and PR review, and Percy’s Git and Visual Git baseline strategies. It also shows how to use ScreenshotNeo when you need clean captures of live pages as part of a separate visual-check workflow.
1. Choose the comparison you need
Before configuring CI, specify what “changed” means for your check. A green result only answers the question asked by that comparison.
| Method | What is compared | Where the expected state lives | Useful when |
|---|---|---|---|
| Playwright screenshot assertions | The current test screenshot against a golden image | Snapshot files in the repository, committed with the test | You want Git-reviewed image files and control over snapshot updates |
| Chromatic UI Tests | A build against accepted snapshots for its branch | Hosted accepted snapshots associated with branch/build history | You want branch-scoped regression checks and hosted review |
| Chromatic UI Review | The pull request head against its merge base | A generated changeset; this is not a UI Test baseline | You want to review the changes a PR would introduce to its base |
| Percy Git | A build against a base-branch build selected through Git history | Git build history; approvals apply to the build | You approve or reject snapshots as a group |
| Percy Visual Git | Snapshots against the latest approved snapshots on each branch | Branch snapshots; approvals can be individual | You need per-snapshot approval granularity |
For hosted branch-aware Storybook or Playwright review, Chromatic is a documented fit. BrowserStack Percy is another relevant option when you need to choose between whole-build and individual-snapshot approval. Select based on baseline ownership and approval behavior, not just on whether the tool displays a diff.
Sources: Chromatic branch and baseline documentation and Percy baseline management.
2. Make a branch baseline policy
Write down these rules before enabling automatic baseline acceptance:
- Baseline owner: Does the team review image files in Git, approve hosted snapshots, or both?
- Comparison target: Is the check against the branch’s accepted state, or against the PR’s merge base?
- Approval authority: Who can accept a visual change, and should approval apply to a full build or individual snapshots?
- Main branch behavior: Does main need a clean, tested build so new branches inherit a known state?
- Branch synchronization: How often should long-lived branches merge or rebase from main?
- Rendering contract: Which browser, operating system or container, fonts, viewport, and test data generate the images?
In Chromatic, each branch has its own accepted baseline. A new branch inherits from its branch point, but later accepted changes on main do not automatically update that feature branch’s baseline. This is intentional branch isolation; a feature branch that falls behind main may therefore show changes already accepted there. Merge or rebase main into it, then rerun and review the resulting diffs.
Keep the distinction between regression and PR review explicit. A PR-to-merge-base comparison answers “what does this branch introduce relative to its base?” A regression comparison answers “what changed since the approved visual state?” One passing does not prove the other baseline is current.
Source: Chromatic: branches, baselines, and Git history.
3. Build stable Playwright visual tests
For repository-managed golden images, Playwright’s toHaveScreenshot() assertion captures an image and compares it with the expected file in the test snapshot directory. On the first run, a missing snapshot is created. Inspect it and commit it with the test so later runs have an explicit reference.
import { test, expect } from '@playwright/test';
test('account page visual state', async ({ page }) => {
await page.setViewportSize({ width: 1280, height: 800 });
await page.goto('http://127.0.0.1:3000/account');
await page.getByRole('heading', { name: 'Account' }).waitFor();
await expect(page).toHaveScreenshot('account-page.png', {
fullPage: true,
animations: 'disabled',
});
});
This example assumes the app is running at the given local address and has stable test data. Replace the URL and page readiness condition with your own application. Playwright snapshot names include browser and platform context, and browsers/platforms can render differently; generate and compare snapshots in a consistent environment.
Update snapshots intentionally
- Run the visual test and inspect the actual, expected, and diff images in the report.
- If the UI change is intended, update the expected image with
npx playwright test --update-snapshots. - Review the changed image files alongside the code diff.
- Commit the new snapshots with the change that caused them.
Do not make snapshot updates an unconditional CI step. That turns detection into silent acceptance and removes the review decision the test is meant to support.
Source: Playwright visual comparisons.
4. Run checks on pull requests and main
Run visual checks for changes proposed by a pull request and for commits to the shared integration branch. Testing main keeps its visual state visible to the team and helps preserve a clean reference as branches are created and merged.
A minimal GitHub Actions example for native Playwright snapshots follows. Adjust Node.js version, package manager, install command, and app startup command to match the repository. The workflow expects a lockfile and a test:e2e script that starts or targets the test application.
name: Visual tests
on:
pull_request:
push:
branches: [main]
jobs:
visual:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
cache: npm
- run: npm ci
- run: npx playwright install --with-deps chromium
- run: npm run test:e2e
Use the same OS/container and browser version for baseline generation and CI comparisons where practical. Retain Playwright’s HTML report and failure artifacts long enough for reviewers to diagnose a diff. For larger suites, Playwright documents splitting work across CI jobs (sharding); keep shards on the same rendering image and configuration.
For Chromatic, follow its GitHub Actions and Playwright integration guidance, and ensure Git metadata is available in CI. History depth and checkout details matter when the service associates commits, pull requests, and baselines. Review how your CI provider constructs pull-request events: some workflows test a synthetic merge commit, which can affect the apparent base and the changes included in review.
Sources: Playwright CI, Chromatic GitHub Actions, and Chromatic for Playwright.
5. Handle merges and baseline updates
- Keep main green and tested. Treat the shared branch as an actively checked branch, not merely the place feature work eventually lands.
- Sync long-lived branches. Merge or rebase main into a feature branch when main’s UI changes affect its expected state.
- Review the right diff. Use branch regression checks to find departures from accepted state and PR-to-merge-base review to see what the PR adds.
- Accept only reviewed changes. Update native snapshot files in Git or approve hosted snapshots after confirming the UI change is expected.
- Configure automation narrowly. Chromatic documents
autoAcceptChangesfor accepting incoming changes on main in certain squash/rebase workflows, andignoreLastBuildOnBranchfor ignoring a target branch’s latest build. Confirm the behavior matches your merge policy before enabling either setting.
These controls affect baseline selection or acceptance behavior. A misconfigured automatic acceptance rule can make an unintended visual change the new reference, so keep human review in the process unless the branch’s changes are deliberately trusted.
Source: Chromatic GitHub Actions guidance.
6. Keep captures reproducible
Visual diffs are only useful when changes in the image correspond to changes you care about. Pin and standardize the rendering inputs as much as your CI allows.
- Browser and operating system: Use consistent browser binaries and runner/container images. Playwright notes that host OS, version, settings, hardware, power source, and headless mode can affect rendering.
- Viewport and device scale: Fix dimensions and scale for each snapshot. Test additional viewports as separate, named cases when responsive behavior matters.
- Fonts and assets: Ensure fonts are installed and loaded before capture. Wait for meaningful readiness signals rather than relying only on elapsed time.
- Animation and time: Disable animations where motion is not under test. Freeze clocks or provide deterministic dates when timestamps affect the view.
- Network and data: Use stable fixtures or controlled test data. Avoid live services whose content changes independently of the commit.
- Dynamic regions: Mask or hide volatile content only when it is outside the behavior under test. A mask that covers too much can hide real regressions.
- Diff tolerance: If using thresholds, choose them deliberately and keep them narrow enough to expose meaningful changes. Thresholds cannot repair unstable rendering.
Playwright’s documentation summarizes the key constraint: “For consistent screenshots, run tests in the same environment where the baseline screenshots were generated.” Source: Playwright visual comparisons.
7. Capture live pages with ScreenshotNeo
Browser-based test runners are suited to app states you control, but sometimes a workflow needs a clean screenshot of a deployed URL—for example, an external page, a staging build reachable over the web, or a set of URLs to review. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. It returns PNG, JPEG, WebP, or PDF from one GET request. For repository-owned visual regression, keep your chosen baseline and approval system; a URL capture does not replace that policy.
Or skip the browser setup
Use the API call below to capture a page. Create an API key and replace the target URL with the page you need. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await (await import('node:fs/promises')).writeFile('shot.webp', bytes);
- Cookie/consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
- Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include
X-Page-VerdictandX-Billedheaders. - An MCP server offers
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - Every plan includes every feature. Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewport, retina scale, PDF settings, HTML/CSS capture, custom CSS and JavaScript, click/wait behavior, hiding selectors, request/resource blocking, headers, cookies, user agent and Authorization, timezone, geolocation, transparent background, resizing, cache TTL, signed image links, async jobs with signed webhooks, bulk capture up to 100 URLs per call, usage API, and OpenAPI spec. Common parameter names used by other screenshot APIs also work.
Pricing: Free includes 1,000 screenshots per month with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. See ScreenshotNeo for the product and plan details. Sign up for 1,000 free screenshots a month, with no card.
8. Troubleshoot branch visual test failures
| Symptom | Likely cause | Fix |
|---|---|---|
| A feature branch reports changes that were already accepted on main | Its branch baseline did not automatically inherit later main approvals | Merge or rebase main into the feature branch, rerun, and review the diff |
| Nearly every screenshot changes in CI | Rendering environment differs: browser, OS, fonts, viewport, headless mode, or other inputs | Use the same runner/container and browser version as baseline generation; verify fonts, dimensions, and settings |
| Hosted service selects an unexpected baseline or misses commits | Git metadata or relevant commit history is absent in CI | Check checkout depth and ensure Git and the history needed by the integration are present |
| PR diff contains surprising base-branch changes | The CI event may test a synthetic merge commit, or the comparison base may not match expectations | Inspect the event checkout and tool’s base configuration; compare the PR head with the intended merge base |
| A changed image becomes expected without clear review | Snapshot update or hosted auto-acceptance ran without an approval gate | Remove unconditional update steps; require image review and limit auto-acceptance to a policy-approved branch/workflow |
| Images differ intermittently between runs | Dynamic data, animation, delayed fonts/assets, or unstable network content | Use deterministic fixtures, wait for a meaningful ready condition, disable irrelevant motion, and mask only truly volatile regions |
| Playwright cannot find a snapshot or creates a new one | The golden file is missing, named differently, or generated for another browser/platform context | Check the test name and snapshot path/context; inspect and commit the intended baseline from the matching environment |
Playwright supports sharding CI runs across jobs. If using shards, ensure the same test inputs and rendering environment are used across jobs, and collect reports/artifacts in a way reviewers can inspect.
9. Performance, reliability, and cost
Performance
Keep the suite focused on representative states instead of taking redundant full-page captures for every test. Run independent tests in parallel or shard them where your CI setup supports it, while preserving the same browser and OS image. Waiting for a selector or application-ready condition is usually more reliable than adding a long fixed delay, and deterministic data prevents repeated work caused by flaky retries.
Reliability
Baseline drift is a process issue as much as a capture issue. Test main, preserve Git history for hosted baseline resolution, sync active branches, and require review of expected-image changes. Keep the PR merge-base comparison separate from branch regression approval so reviewers can tell whether they are seeing new PR work or divergence from an older accepted state.
Cost
Playwright’s native screenshot assertions use repository snapshots and run in your own CI environment; account for the CI runtime and artifact storage your provider charges under its own terms. Hosted services have their own plan limits and pricing, which can change; consult their current product pages before choosing. ScreenshotNeo’s stated plans are listed above, and only clean shots are billed; failed/bot/blank/cache outcomes are free. No unsupported defect-reduction or performance benchmark is needed to choose a baseline policy.
10. Frequently asked questions
Should pull requests compare with main or their own branch baseline?
It depends on the question. Use a merge-base comparison to see what the PR adds relative to its base. Use a branch regression baseline to catch changes since that branch’s accepted visual state. Teams may run both.
Do feature branches inherit every later baseline update from main?
Not in Chromatic’s branch model. A branch starts from its branch point; sync main into it to bring later main changes into the branch’s code and visual state.
Should visual snapshots be committed to Git?
With Playwright’s native assertions, expected screenshots are repository files and should be reviewed and committed. Hosted services store and associate approvals according to their own baseline model.
Can a passing PR review prove that regression baselines are current?
No. A merge-base review and an approved-baseline regression test compare different references.


