How to Run Playwright Visual Regression Tests in Bitbucket Pipelines
Set up Playwright screenshot comparisons in Bitbucket Pipelines with reproducible baselines, useful failure artifacts, and fixes for CI-only visual diffs.
Run Playwright visual regression tests in Bitbucket Pipelines by using a Linux pipeline step with browsers available, installing the project dependencies, and running npx playwright test. Use Playwright’s Docker image with a tag aligned to the Playwright package in your project, start with one CI worker, and commit reviewed screenshot baselines generated in the same environment. Preserve test results and reports as pipeline artifacts so reviewers can inspect failures.
1. How Playwright visual regression tests work
Playwright Test’s toHaveScreenshot() assertion captures a screenshot and compares it with a reference image. On the first run, Playwright creates reference snapshots; review those files and commit the intended images. Subsequent runs compare new captures against the committed references and fail when the difference exceeds the configured tolerance.
These comparisons are pixel-sensitive. Operating system, browser version, browser settings, hardware, power source, and headless mode can all affect rendering. A baseline made on a developer laptop may differ from an otherwise identical page rendered by a Linux CI worker. Generate and compare baselines in the same controlled environment wherever possible.
2. Add the Bitbucket Pipelines configuration
Enable Bitbucket Pipelines for the repository, then add a bitbucket-pipelines.yml file at the repository root. This minimal example uses the documented Playwright image tag as a versioned example. Keep the image’s Playwright version aligned with the version installed by the project; update both deliberately when upgrading.
image: mcr.microsoft.com/playwright:v1.63.0-noble
pipelines:
default:
- step:
name: Playwright visual tests
caches:
- node
script:
- npm ci
- npx playwright test
artifacts:
- test-results/**
- playwright-report/**
The image supplies a Linux environment prepared for Playwright browsers. The node cache is an example dependency cache; retain it only if it suits the repository and package manager. Artifact paths are relative to the clone directory, so make sure they match the directories your Playwright configuration actually writes.
Configure Playwright for CI
Set CI-friendly output paths and a single worker in the project’s Playwright configuration. A representative playwright.config.ts is:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: Boolean(process.env.CI),
retries: process.env.CI ? 1 : 0,
workers: process.env.CI ? 1 : undefined,
reporter: process.env.CI
? [['list'], ['html', { outputFolder: 'playwright-report', open: 'never' }]]
: 'list',
outputDir: 'test-results',
use: {
baseURL: process.env.BASE_URL ?? 'http://127.0.0.1:3000',
trace: 'retain-on-failure',
screenshot: 'only-on-failure',
},
});
In this example, workers is set to one in CI for stability and reproducibility, while local runs can use Playwright’s default worker behavior. Retry count is a policy choice: retries can help collect diagnostic evidence for intermittent failures, but should not be used to make unstable visual tests appear reliable. If the application needs to be started by the test run, configure a webServer entry in Playwright or add the appropriate startup and readiness commands to the pipeline before running tests.
The configuration above is an example, not a required set of options. If the repository already defines a base URL, reporter, output directory, or CI policy, preserve that setup and make the artifact paths agree with it. The Playwright image tag shown above is an example from the official CI guidance, not a promise that it will remain the newest tag.
3. Write a visual test and create its baseline
Use a stable route and explicitly control state that affects the rendered page. For example, a test can visit the landing page and compare its content:
import { test, expect } from '@playwright/test';
test('landing page visual appearance', async ({ page }) => {
await page.goto('/');
await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible();
await expect(page).toHaveScreenshot('landing.png');
});
With a configured baseURL, page.goto('/') resolves against that origin. Replace the heading and route with elements from your application. On the initial baseline-generation run, review the generated snapshot files before committing them. The repository’s snapshot directory is created and managed by Playwright; keep its committed reference images under version control so pull requests compare against an intentional baseline.
To intentionally update references after a visual change, run:
npx playwright test --update-snapshots
Review snapshot changes like code changes. An update command accepts the current rendering as the new reference; it does not determine whether the change is correct.
4. Keep screenshot comparisons reproducible
- Align versions: keep the Docker image and the project’s Playwright package on matching versions. Upgrade them together and review any resulting browser-rendering changes.
- Use a consistent environment: generate baselines and compare them with the same operating system, browser build, settings, and headless behavior whenever practical.
- Control page state: use predictable test data, wait for meaningful page readiness, and avoid capturing during animations or asynchronous content updates.
- Handle volatile content narrowly: Playwright supports a stylesheet option for hiding or neutralizing content that is genuinely irrelevant to the comparison. Filter only sources of irrelevant variability; do not hide elements whose visual changes matter.
- Keep concurrency conservative: begin with one worker in CI. Increase it only after checking runner capacity and confirming concurrent tests remain stable.
- Review every baseline update: a changed reference can represent a real regression or a deliberate design update. Keep the decision visible in the code review.
For a page with known dynamic regions, Playwright’s screenshot assertion can take options. One example is a stylesheet that masks a changing timestamp while leaving the rest of the page testable:
await expect(page).toHaveScreenshot('dashboard.png', {
stylePath: './tests/visual-stability.css',
});
/* tests/visual-stability.css */
.test-only-clock {
visibility: hidden !important;
}
Use selectors or test fixtures that target only content whose variation is outside the purpose of the test. Keep important status labels, totals, and layout regions visible so visual regressions remain detectable.
5. Preserve failure evidence as Bitbucket artifacts
Artifacts let reviewers download reports and test output after the pipeline step completes. The sample pipeline retains test-results/** and playwright-report/**; verify that these folders are where your configuration emits screenshots, traces, and reports. Artifact paths must stay inside BITBUCKET_CLONE_DIR.
Atlassian’s documentation currently specifies a maximum artifact retention of 14 days and a 1 GB size limit. These limits can change, so check the current Bitbucket documentation when setting retention expectations. Avoid uploading bulky outputs that are not useful to debug a failed comparison.
6. Caching, runtime, reliability, and cost
What to cache
A dependency cache can reduce repeated package installation work, provided it matches the project and package manager. Browser-binary caching is optional and generally not recommended by Playwright: restoring the cache can take about as long as downloading browser binaries. If you choose to cache browser binaries anyway, key the cache against the Playwright version so a browser update cannot silently reuse an incompatible cache.
How to keep the job dependable
One worker is the recommended starting point for CI stability and reproducibility. More workers may shorten a suite, but can also compete for runner resources and expose timing or shared-state problems. Measure stability and available capacity before increasing concurrency. Keep the browser image, package version, baseline environment, and artifact configuration explicit so failures are easier to reproduce.
Cost and duration considerations
The supplied research does not establish a universal runner size or a fixed cost per run. Pipeline execution time depends on your suite, dependency installation, browser startup, capture count, and runner allocation. Reduce avoidable work by running only the needed test selection for a given pipeline where appropriate, using a suitable dependency cache, and keeping artifact output focused. Do not assume that browser caching saves time without comparing its restore time with a normal install.
7. Troubleshooting visual tests in Bitbucket
| Symptom | Likely cause | What to do |
|---|---|---|
| Tests pass locally but screenshots fail in CI | Different OS, browser version, settings, hardware, or headless behavior changes rendering. | Generate and compare baselines in the same controlled Linux image used by CI. Align the image tag with the project’s Playwright version, then review and commit updated references only when the difference is expected. |
| Browser executable or shared-library errors | The runner image does not include the browser dependencies expected by the installed Playwright version, or versions are misaligned. | Use the Playwright Docker image recommended for CI and align its version with the package. If using a different environment, follow Playwright’s CI guidance for installing browser binaries and required dependencies. |
| Snapshots are missing or every run looks like a first run | Reference screenshots were not committed, or the test is looking for a different project or snapshot path. | Inspect the generated snapshot files, commit the intended baselines, and check the test name, project configuration, and repository status. |
| Pipeline reports no artifacts | Artifact paths do not match actual output directories, or outputs are outside the clone directory. | Check outputDir, reporter output, and Bitbucket artifact globs. Keep artifact paths relative to the repository clone directory. |
| Visual diffs change from run to run | Volatile page content, animations, incomplete loading, or resource contention makes captures inconsistent. | Control page state and readiness, neutralize only irrelevant dynamic content with a screenshot stylesheet, and begin with one CI worker. |
| Pipeline is slow after adding browser caching | Cache restore and extraction cost may rival downloading browser binaries. | Remove browser caching or key it to the Playwright version and compare the actual pipeline behavior. Keep dependency caching separate from browser-binary caching. |
| Updating snapshots removes evidence of a regression | The current output was accepted without reviewing the diff. | Run the update locally or in the controlled environment, inspect every changed image, and commit only intentional visual changes. |
8. Or skip the browser setup
If your task is to capture a page image rather than compare a committed Playwright baseline, ScreenshotNeo provides a one-call screenshot API. It can return PNG, JPEG, WebP, or PDF and also offers an MCP server for AI agents. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, and cache hits are not billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. ScreenshotNeo is made by Yorker Media. Sign up for 1,000 free screenshots a month with no card.
9. FAQ
Should visual baselines be committed to the repository?
Yes. Commit reviewed reference screenshots so CI has a stable comparison target and code review can show intentional visual changes.
Can I use a different Playwright image tag?
Yes. Choose a tag appropriate for the project, and keep it aligned with the installed Playwright version. The documented tag in this guide is a versioned example.
Should I set retries to zero or one?
That is a team policy decision. Retries can preserve diagnostic evidence for intermittent failures, but fix the underlying instability rather than relying on retries to make a flaky visual test acceptable.
Where can I find the official guidance?
See the Playwright CI guide, Playwright visual comparisons, and Atlassian’s documentation for Docker images in build environments and pipeline artifacts.


