BlogScreenshots on your device
How to Capture Screenshots in Windows UI Test Automation
Capture reliable Windows UI evidence with winapp, Appium, Selenium, or Playwright, then diagnose overlays, flaky rendering, and CI failures.
Direct answer: capture screenshots at the layer that owns the UI. For native Windows applications, use the Microsoft winapp CLI or Appium’s Windows driver. Use a stable window handle (HWND) when several windows exist, and capture the smallest useful scope: an element for a control, a window for a dialog or page, or the whole screen when menus, flyouts, tooltips, or other overlays are visible. For browser tests, use Selenium’s page or element screenshot APIs, or Playwright’s screenshots and visual assertions.
Drive the application to the required state, wait until it settles, save a PNG named with the test and failure ID, and publish it as a CI artifact. Keep visual-regression baselines on a consistent browser, operating system, font set, rendering configuration, and headless mode.
Choose the capture layer
| Test target | Recommended tool | Best scope | Important detail |
|---|---|---|---|
| Win32, WPF, WinForms, or WinUI | winapp CLI or Appium Windows Driver | Element or window | Give controls stable AutomationId values where supported. |
| Desktop WebDriver tests | Appium with the Windows Application Driver plugin | Window or element | Microsoft says WinAppDriver is no longer under active development and recommends Appium with the Windows Application Driver plugin. |
| Browser UI | Selenium | Page or element | WebDriver returns screenshot data encoded in Base64. |
| Browser visual regression | Playwright Test | Page, element, or baseline | Use expect(page).toHaveScreenshot() with a controlled environment. |
| Menus, flyouts, tooltips, overlays | winapp screen capture | Full screen | Foreground the target and capture with --capture-screen. |
Capture native Windows apps with winapp
The Microsoft winapp CLI uses Windows UI Automation and can capture a whole window, an element crop, or the screen. Its normal Windows Graphics Capture path captures the DWM-composited surface and can work while the window is occluded; the documentation describes a PrintWindow fallback when WGC is unavailable.
1. Inspect the UI
winapp ui inspect -a notepad
Use inspection to discover the application identity and element identifiers. If multiple instances are open, identify the process, title, PID, or HWND and prefer the most stable identifier available.
2. Capture a window
winapp ui screenshot -a notepad
winapp ui screenshot -a notepad --output smoke-test.png
winapp ui screenshot -a notepad --json
The command captures the application window as PNG. --json is useful in automation because the response can be logged and the output path collected as a CI artifact.
3. Capture by HWND
winapp ui screenshot -w 131906
An HWND is often safer than a title when a test opens multiple windows with similar names.
4. Capture one element
winapp ui screenshot txt-searchbox-e5f6 -a myapp
Element screenshots keep failure evidence focused on the control under test. Prefer a stable AutomationId over a localized label or a position-based selector.
5. Capture overlays and popups
winapp ui screenshot -a myapp --capture-screen
Use screen capture for popup menus, dropdowns, flyouts, and tooltip overlays. The command brings the target window to the foreground, so account for that side effect in parallel or interactive test suites.
Build a reliable screenshot step
- Identify the target: resolve the app by process or app name, title, PID, or HWND. Store the resolved identifier in the test log.
- Reach the state: perform clicks, keyboard input, navigation, and data setup before capturing.
- Wait for stability: wait for a known element, an application-ready signal, or a short settling interval. Avoid arbitrary long sleeps when a state condition is available.
- Select scope: element for a control, window for a dialog or page, screen for overlays.
- Name the artifact: include test name, scenario, and failure identifier, for example
checkout-invalid-card-test-42.png. - Publish it: upload the file or JSON result to the CI artifact store even when the test fails.
PowerShell failure-hook example
param(
[string]$App = "myapp",
[string]$Output = "artifacts\ui-failure.png"
)
New-Item -ItemType Directory -Force (Split-Path $Output) | Out-Null
winapp ui screenshot -a $App --output $Output --json
if ($LASTEXITCODE -ne 0) {
Write-Error "Screenshot capture failed with exit code $LASTEXITCODE"
exit $LASTEXITCODE
}
Write-Host "Screenshot saved to $Output"
Appium for Windows desktop tests
For WebDriver-style desktop automation, use Appium with the Windows Application Driver plugin. Microsoft’s current guidance identifies WinAppDriver as the original tool and says it is no longer under active development.
Give controls stable AutomationProperties.AutomationId values in frameworks that support them. In a test, locate the application window, perform the action, wait for the expected state, and call the WebDriver screenshot method. Keep the driver session and application version fixed in CI so artifacts remain comparable.
Selenium browser screenshots
Selenium drivers expose page and element screenshot operations. The WebDriver endpoint returns image data encoded in Base64; your test runner should decode it and write a PNG artifact.
// C# example
using OpenQA.Selenium;
using OpenQA.Selenium.Chrome;
using var driver = new ChromeDriver();
driver.Navigate().GoToUrl("https://example.test");
var checkout = driver.FindElement(By.CssSelector("[data-testid='checkout']"));
checkout.GetScreenshot().SaveAsFile("artifacts/checkout.png");
Playwright screenshots and visual regression
import { test, expect } from '@playwright/test';
test('checkout error state', async ({ page }) => {
await page.goto('https://example.test/checkout');
await page.getByRole('button', { name: 'Pay' }).click();
await expect(page).toHaveScreenshot('checkout-error.png', {
maxDiffPixels: 50,
stylePath: 'tests/visual-mask.css'
});
});
Playwright stores PNG by default and supports WebP. Use stylePath to mask dynamic content and maxDiffPixels to set an explicit tolerance. Its documentation warns that rendering changes with browser and operating-system versions, fonts, settings, hardware, power source, and headless mode. Generate and compare baselines in the same project environment.
Element, window, or screen?
| Scope | Use it for | Trade-off |
|---|---|---|
| Element | Button, field, grid, or component assertion | Small, focused artifact; can miss an overlapping popup. |
| Window | Dialog, page, or native application state | Shows context while excluding unrelated monitors and windows. |
| Screen | Menus, flyouts, tooltips, drag indicators, and cross-window overlays | Captures pixels outside the app and may foreground the target window. |
CI and reliability checklist
- Run with a deterministic screen resolution, DPI scale, theme, locale, and timezone.
- Install the exact fonts used by visual baselines.
- Use a dedicated interactive desktop session when the tool requires a visible desktop.
- Wait for application state instead of relying only on fixed delays.
- Record app version, OS build, browser/driver version, and capture scope with each artifact.
- Retry only transient startup or transport failures; do not retry a real visual mismatch until its cause is understood.
- Keep screenshots on failure and delete passing-run artifacts according to your retention policy.
Troubleshooting
“Window not found”
Cause: the title changed, multiple instances exist, or the app has not finished starting. Fix: inspect the UI, target the process/PID or HWND, and wait for the window-ready condition.
Blank or black image
Cause: the surface is not ready, the app is minimized, or the capture path cannot access the rendered surface. Fix: wait for a visible state, restore the window, try the normal WGC path, and use the documented fallback behavior when WGC is unavailable.
Popup missing
Cause: a window capture excludes overlays owned by another surface. Fix: use --capture-screen, which is intended for popup menus, dropdowns, flyouts, and tooltips.
Element selector fails intermittently
Cause: a generated or localized identifier. Fix: add a stable AutomationId or test ID and wait for the element to be available and enabled.
Playwright diffs change between machines
Cause: different browsers, OS versions, fonts, hardware, power settings, or headless mode. Fix: pin the environment, regenerate baselines there, and use a narrowly scoped diff threshold.
Screenshot file is missing in CI
Cause: the test exits before the artifact upload or writes to a workspace that is not collected. Fix: create the artifact directory first, log the absolute path, and configure upload steps to run after failures.
Performance, reliability, and cost
Element captures are usually smaller and faster to transfer than full-screen images. Capture only the scope needed for the assertion, and avoid capturing on every passing step. Full-screen captures are appropriate when the evidence includes overlays or multiple windows.
For visual regression, rendering determinism matters more than raw capture speed. Pin browser and OS images, fonts, viewport, scale factor, and power mode. For failure diagnostics, retain the screenshot alongside logs and the UI state that produced it.
Self-hosted desktop capture consumes the time and resources of the test machine. If you only need website screenshots rather than a native desktop surface, an API can remove browser installation and display-session maintenance.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. It is useful when the target is a web page rather than a native Windows window: cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; and each response reports its page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP tools let Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options. This one-call example captures a WebP image:
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, custom CSS and JavaScript, clicks, selector/delay/network-idle waits, request and resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, PDF output, and an OpenAPI specification. The parameter names used by other screenshot APIs also work, which can simplify migration.
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Should I capture the screen or the application window?
Capture the window for normal app evidence. Capture the screen when the evidence includes menus, flyouts, tooltips, or overlays outside the window surface.
Is WinAppDriver still the preferred Microsoft tool?
No. Microsoft’s current guidance says WinAppDriver is no longer under active development and recommends Appium with the Windows Application Driver plugin.
What format should test screenshots use?
Use PNG for lossless diagnostics and visual comparisons. Keep the same format and rendering environment for every baseline.
How do I prevent dynamic content from breaking visual tests?
Mask or hide dynamic regions, wait for the page to settle, and set a narrowly scoped pixel-difference threshold. Playwright supports stylePath and maxDiffPixels.


