How to Run Playwright Screenshot Tests in Azure DevOps
Configure Azure Pipelines to run Playwright visual tests, compare stable screenshot baselines, and publish reports and traces—even when tests fail.
Run Playwright screenshot tests in Azure DevOps by installing the repository’s locked Node dependencies and matching Playwright browsers, running npx playwright test in Azure Pipelines, and publishing the JUnit results and HTML report even when a test fails. Use expect(page).toHaveScreenshot() for visual assertions, commit reviewed baseline images, and generate or update baselines in the same browser and operating-system environment used in CI.
1. Add Playwright screenshot assertions
Use the Playwright Test runner and the toHaveScreenshot assertion. A focused test might look like this:
import { test, expect } from '@playwright/test';
test('landing page visual appearance', async ({ page }) => {
await page.goto('https://example.com');
await expect(page).toHaveScreenshot('landing.png');
});
On the first run, Playwright creates a reference screenshot. Review the generated image and commit it with the test. Later runs compare a new capture against that baseline. A changed screenshot is a signal to inspect: it may reflect an intentional design change, a regression, or environmental variation.
Keep baseline generation and CI comparison consistent. Operating system, browser version, browser settings, fonts, hardware, power source, and headless mode can all affect pixels. Prefer creating and reviewing baselines in the same environment as the pipeline. Do not accept unexplained differences as harmless noise; change comparison tolerances only when you can justify the tradeoff. Playwright documents screenshot assertions, baselines, and capture options in its visual comparisons guide.
Reduce noise from genuinely volatile regions
If a timestamp, rotating banner, or other unstable region makes a meaningful page comparison unreliable, apply a narrowly scoped stylesheet to the capture. Keep the rest of the page visible so the test continues to cover the UI that matters. Playwright’s screenshot assertion supports stylePath:
import { test, expect } from '@playwright/test';
test('landing page without volatile timestamp', async ({ page }) => {
await page.goto('https://example.com');
await expect(page).toHaveScreenshot('landing.png', {
stylePath: './screenshot-test.css',
});
});
/* screenshot-test.css */
.test-only-timestamp {
visibility: hidden !important;
}
Use a selector that targets only the volatile content. Hiding broad page regions can make a test pass while concealing a real rendering problem.
2. Configure Playwright for CI
Start with one worker in CI. This favors stable and reproducible runs; if runtime becomes a problem, shard the suite across jobs before increasing concurrency on a single agent. A configuration can also retain a trace on the first retry:
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
reporter: [
['list'],
['junit', { outputFile: 'test-results/e2e-junit-results.xml' }],
['html', { outputFolder: 'playwright-report', open: 'never' }],
],
retries: process.env.CI ? 1 : 0,
workers: process.env.CI ? 1 : undefined,
use: {
baseURL: process.env.BASE_URL || 'https://example.com',
trace: 'on-first-retry',
},
});
The HTML reporter writes its output when the test run completes, including a run with failed tests. The JUnit reporter creates XML that Azure DevOps can surface as test cases. Traces on retry help investigate intermittent failures without recording a trace for every test; Playwright notes that tracing every test is performance-heavy. See the official Trace Viewer guide.
3. Run the suite in Azure Pipelines
Azure DevOps runs YAML CI workflows through Azure Pipelines. This Linux-hosted example installs Node.js, restores exactly the lockfile dependencies with npm ci, installs Playwright browsers and their Linux system dependencies, then runs the tests. Update the Node version to one supported by your repository and keep the lockfile’s Playwright package version aligned with the browser install.
trigger:
- main
pool:
vmImage: ubuntu-latest
steps:
- task: UseNode@1
inputs:
version: '22'
displayName: 'Install Node.js'
- script: npm ci
displayName: 'Install locked dependencies'
- script: npx playwright install --with-deps
displayName: 'Install Playwright browsers and Linux dependencies'
- script: npx playwright test
displayName: 'Run Playwright tests'
env:
CI: 'true'
BASE_URL: 'https://example.com'
- task: PublishTestResults@2
displayName: 'Publish JUnit test results'
condition: succeededOrFailed()
inputs:
testResultsFormat: 'JUnit'
testResultsFiles: 'test-results/e2e-junit-results.xml'
failTaskOnFailedTests: false
testRunTitle: 'Playwright visual tests'
- task: PublishPipelineArtifact@1
displayName: 'Publish Playwright HTML report'
condition: succeededOrFailed()
inputs:
targetPath: 'playwright-report'
artifact: 'playwright-report'
publishLocation: 'pipeline'
The reporting tasks use succeededOrFailed() so their results and report can be published after the test command fails. The test task itself still fails the job when assertions fail. Confirm that each configured reporter creates its output in the expected location; publishing a missing report directory will itself fail.
Playwright’s official CI guide shows this installation flow for Linux agents. Windows and macOS agents need no additional browser configuration beyond installing Playwright and running the tests. If you switch operating systems, regenerate and review the visual baselines for that environment.
4. Keep browser versions and baselines aligned
For a hosted Linux agent, npx playwright install --with-deps installs browser binaries and system dependencies. Another option is a Playwright container image. If you use one, match its Playwright version to the package version in the repository and update the image when you update that dependency. The official CI documentation shows a versioned container example; its particular tag is an example, not a recommendation to pin that release indefinitely.
| Approach | Best fit | Keep in mind |
|---|---|---|
Hosted Linux agent plus install --with-deps |
A straightforward pipeline using the agent’s Node setup | Install browser and system dependencies during the job; keep the package lock and browser install consistent. |
| Version-matched Playwright container | A team that wants an explicitly selected browser environment | Update the container tag together with the Playwright dependency. |
Do not compare screenshots from a local machine with CI baselines without checking the OS and browser differences first. If a visual diff appears after an agent image, font, browser, or dependency update, determine whether the rendering environment changed before updating snapshot files.
5. Speed up a stable suite with sharding
Begin with one CI worker to reduce machine-level concurrency and improve reproducibility. If the test duration is too long, divide the suite among Azure jobs using Playwright’s --shard option. Sharding spreads test files across jobs; it does not remove the need for each job to install the matching browsers and dependencies.
# Example commands for a two-job Azure Pipelines matrix
npx playwright test --shard=1/2
npx playwright test --shard=2/2
Pair this with an Azure Pipelines matrix strategy or separate jobs, and give each job a distinct result/report location if both publish artifacts. Add shards gradually and check that the longer setup and artifact collection steps do not erase the time saved by parallel execution. Playwright describes CI worker and sharding choices in its CI documentation.
6. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable or shared library missing | Browser binaries or Linux system dependencies were not installed on the agent. | On Linux run npx playwright install --with-deps after npm ci, or use a matching Playwright container. |
| Screenshot assertion fails only in CI | The baseline and pipeline differ in OS, browser version, fonts, settings, hardware, or headless rendering. | Compare the environments and generate/review the baseline in the CI environment. Inspect the actual and expected images before accepting changes. |
| Unexpected snapshot changes after a dependency update | The update may have changed the browser build or rendering behavior. | Check the Playwright package, browser version, and agent/container image together; review resulting pixel changes deliberately. |
| Test-results XML is not published | The JUnit reporter path and Azure task glob do not match, or the reporter did not write a file. | Set the reporter outputFile and testResultsFiles to the same path, then inspect pipeline logs. |
| HTML artifact task fails after tests fail | The report folder was not created or the task uses the wrong path. | Confirm the HTML reporter is configured with that output folder and verify it exists before publishing. |
| Intermittent CI failure has little diagnostic detail | No trace was captured for the failing retry. | Set trace: 'on-first-retry', reproduce the failure, and open the trace from the Playwright report or with Trace Viewer. |
| Tests pass locally but differ against a remote browser service | The runner host OS can determine expected screenshot paths, while the remote browser may render on another OS. | For Playwright Workspaces, follow Microsoft’s service-specific advice to run comparisons in the service or use ignoreSnapshots for that run if appropriate. Do not apply this workaround automatically to ordinary Azure-hosted browser runs. |
For remote Playwright Workspaces, see Microsoft’s visual comparisons guidance; its advice is specific to that service and should not be generalized to a normal hosted Azure Pipelines agent.
Or skip the browser setup
For a one-off screenshot or a workflow that should not install and maintain a browser on the pipeline agent, ScreenshotNeo provides a screenshot API and MCP server. One GET request returns an image or PDF. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the page verdict and billing status in response headers. Its MCP server gives AI agents such as Claude and Cursor the take_screenshot, get_page_info, and capture_pdf tools. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Learn more at ScreenshotNeo, then sign up for 1,000 free screenshots a month with no card.
Frequently asked questions
Should I commit Playwright screenshot baselines?
Yes. Review the generated images and commit the intended references with the tests so CI has stable expectations to compare.
Can I use Azure’s Windows or macOS agents?
Yes. Playwright’s CI guidance says those agents need no additional browser configuration beyond installing Playwright and running tests. Keep baseline environment differences in mind.
Should I retry every visual test?
A retry can help capture diagnostic traces for intermittent failures, but it does not explain or resolve a genuine visual change. Inspect the screenshot diff and trace before accepting a result.
What if I need a screenshot artifact but not a regression assertion?
Use a screenshot capture workflow and publish its output as an artifact. A captured image by itself is not a visual regression test; comparison requires an expected baseline and an assertion.


