ScreenshotNeo

BlogHow-to

How to Add Visual Testing to BDD Tests

Add screenshot checks to BDD tests at stable, meaningful UI states. Learn how to capture, compare, review baselines, and troubleshoot visual regressions.

By the ScreenshotNeo team4 October 20269 min read

Add visual testing to BDD tests by capturing a screenshot after a scenario reaches a meaningful, stable UI state, then comparing it with an approved baseline. Keep the Gherkin scenario focused on behavior; put screenshot capture and comparison in the UI automation layer. Review every difference before accepting a new baseline.

This guide uses Playwright Test with Applitools Eyes for a concrete integration. The same placement principle applies to Cucumber with Playwright and other BDD runners, but their package setup and hooks differ. Verify those details against documentation for the versions in your suite.

1. Choose a meaningful visual checkpoint

BDD examples describe behavior that a team can discuss and automate. Cucumber describes BDD as collaborative work that closes the gap between business and technical teams and produces shared understanding and system documentation checked against behavior. A visual assertion complements that purpose by checking the rendered interface at a behavior outcome; it does not replace the behavioral assertions.

Choose a point where the page communicates an important outcome, such as a completed sign-in, a validation error, or a submitted form. Avoid capturing after every step: each checkpoint adds maintenance and review work, so prioritize states where a layout or rendering regression would matter.

For example, a readable scenario can describe successful sign-in, while the step definition or test lifecycle captures the resulting account page:

Feature: Sign in
  Scenario: A valid user reaches their account
    Given I am on the sign-in page
    When I sign in with valid credentials
    Then I should see my account page

Keep the scenario understandable to its readers. Place vendor-specific screenshot calls in the automation code behind the scenario rather than turning the Gherkin steps into SDK instructions.

2. Stabilize the UI before capture

A screenshot comparison is only useful when the current run and baseline represent the same intended state. Control the inputs that can change pixels:

  • Test data: use known accounts and deterministic content. Reset state when one scenario can affect another.
  • Viewport and browser: keep the browser, viewport, device scale, and environment consistent with the baseline.
  • Loading: wait for the relevant content, fonts, and images. Prefer waiting for a meaningful element over a fixed delay when possible.
  • Motion and transient UI: settle animations and remove or control timestamps, rotating content, cursors, and other unstable elements.
  • Dynamic regions: mask or ignore only content that is intentionally variable. Broad masks can hide real defects.

Do not assume that a page’s initial load event means the interface is visually ready. Wait for the state your scenario is asserting, then capture.

3. Add a visual assertion in Playwright

Applitools documents a Playwright Test fixture that provides both page and eyes. The example below follows that fixture pattern and checks the account screen after the scenario’s behavioral steps.

import { test, expect } from '@applitools/eyes-playwright/fixture';

test('a valid user reaches their account', async ({ page, eyes }) => {
  await page.goto('http://localhost:3000/sign-in');
  await page.getByLabel('Email').fill('test@example.com');
  await page.getByLabel('Password').fill('correct-test-password');
  await page.getByRole('button', { name: 'Sign in' }).click();

  await expect(page.getByRole('heading', { name: 'Your account' }))
    .toBeVisible();

  await eyes.check('Account after sign-in', {
    fully: true,
    matchLevel: 'Strict'
  });
});

This is an integration example, not a universal setup recipe. Install and configure the current Applitools Playwright package for your project and versions, including its required credentials and runner setup. The documented pattern imports test from @applitools/eyes-playwright/fixture and calls eyes.check().

The checkpoint name should identify the screen or state. A name such as Account after sign-in is more useful in review than a generic name such as Screenshot 1. Use full-page capture when the entire page matters; use a viewport or targeted region when that is the intended contract.

4. Keep behavioral assertions alongside visual checks

A visual comparison can show that something changed, but it does not explain whether a business rule is correct. Retain behavioral assertions for outcomes that matter, such as navigation, validation, and dynamic values. For example, the test above checks that the account heading is visible before it captures the screen.

Use textual assertions where they are useful for dynamic aspects, rather than expecting every changing value to remain pixel-identical. A visual test should add coverage for rendered presentation, not become the only check that the scenario succeeded.

5. Compare and review baselines deliberately

Visual testing compares the current capture with a previously stored baseline. Treat a baseline as the approved reference for a specific application, environment, viewport, browser, and state. A difference is a prompt for review, not automatic proof of a defect.

  1. Run the scenario against the intended environment and inspect the comparison.
  2. If the UI change is intentional, approve the new screenshot as the baseline.
  3. If it is unexpected, reject the change, retain the previous baseline, and investigate the regression.
  4. When updating a baseline, confirm that the change reflects the intended product behavior and was captured under the expected conditions.

Keep the checkpoint name and scenario context visible in the test report so a reviewer can connect a pixel difference to the behavior under test.

6. Choose the comparison and storage approach

Decide how visual checks fit your team’s review process. These are implementation choices, not a universal ranking:

Decision Questions to answer
Comparison method Do you need pixel-level comparison, or a semantic or AI-assisted way to judge visual changes?
Baseline location Will references live with the project or in a hosted review service? Who can approve updates?
Coverage Is one browser and viewport enough, or do important users need cross-browser or device coverage?
Dynamic content Which values should be controlled, asserted separately, or narrowly ignored?
CI behavior How will failures show the scenario, named checkpoint, and comparison for review?
Baseline changes How will the team distinguish an intentional design update from a regression?

The right setup depends on your runner, app, and review workflow. Applitools documents its Playwright integration and eyes.check() options, including full-page capture, match level, and ignored regions. Its eyesConfig also includes settings such as appName and whether visual differences fail the test. Check the current vendor documentation for exact option names and behavior.

7. Integrate with a Cucumber or other BDD runner

For Cucumber with Playwright, keep the feature and step definitions readable, and call the visual SDK once the step or hook has reached the rendered state you want to verify. A shared test lifecycle or page-object layer can be a suitable integration point, depending on how the suite manages browser sessions and scenario state.

For Ruby/Cucumber, Java, or another runner, do not copy the Playwright fixture example. Find the visual SDK’s current integration instructions for that language and runner, then connect its capture call to the scenario lifecycle or relevant step. The available Applitools Cucumber article about Ruby setup dates to 2018; use it only as historical architectural context, not as current package instructions.

Before adopting an integration, verify that each scenario gets the intended browser and visual-test session, that failures are reported to the runner, and that baseline approval fits your team’s permissions and CI process.

8. Run visual checks in local development and CI

Run the visual check in the same feedback loop as the UI scenario. A useful failure should tell the reviewer which scenario and checkpoint changed and provide access to the comparison. The precise CI configuration depends on your runner and visual service; avoid hiding visual failures behind a separate, unreviewed job.

  • Use stable test data and a consistent rendering environment for baseline and comparison runs.
  • Keep credentials and service configuration in the approved environment mechanism for your CI platform.
  • Make visual differences visible as actionable review items.
  • Require deliberate approval before baseline updates become the new reference.

9. Troubleshoot common visual test failures

Symptom Likely cause What to do
Differences vary from run to run Uncontrolled data, animation, delayed assets, or transient content Stabilize the data and viewport, wait for the meaningful UI state, and narrowly mask unavoidable variable regions.
The screenshot shows a loading or incomplete page Capture happens before the scenario’s UI is ready Wait for a scenario-relevant element or content condition before calling the visual check.
Large differences after a browser or environment update The rendering environment changed from the baseline environment Confirm browser, viewport, device scale, fonts, and app environment. Decide whether the change is intentional before updating references.
A meaningful defect passes unnoticed An ignored region is too broad or the checkpoint misses the affected state Narrow ignored areas and add a checkpoint at the state where the defect would appear.
The visual check runs but does not fail the test as expected Runner or SDK difference-failure configuration is not set as intended Review the current integration’s failure settings, including the documented eyesConfig behavior, and confirm how the runner reports differences.
Step definitions cannot access the visual fixture The suite uses a different lifecycle or fixture model than Playwright Test Use the integration documented for the actual BDD runner and language; do not assume the Playwright Test fixture is available in Cucumber hooks.
Reviewers cannot tell what changed Generic checkpoint names or missing scenario context Use descriptive names and surface the scenario and checkpoint in reports.

10. Performance, reliability, and cost considerations

Each visual checkpoint adds capture and comparison work to the test run, and each baseline adds something the team must maintain. Keep checks focused on valuable states, and avoid duplicating nearly identical screenshots without a reason. Full-page captures can be useful, but capture only the area that represents the intended visual contract.

Reliability comes primarily from repeatable inputs and deliberate baseline governance: control test data, wait for stable content, keep rendering conditions consistent, and review changes before accepting them. The research sources do not establish neutral prices, service limits, or performance benchmarks for visual-testing vendors, so compare those against current vendor terms for your expected test volume.

Or skip the browser setup

If you need a screenshot for a visual check without maintaining a browser capture setup, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Or call it from Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Or use Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
  • Cookie and consent banners are accepted or removed before capture, along with known newsletter popups and chat widgets; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers say which page verdict occurred and whether it was billed.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents, including Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan.

For repeatable visual regression tests, keep your capture URL, viewport, and page state controlled, and review baseline changes just as you would with browser-based captures. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

How do I add visual regression testing to Cucumber tests?

Keep the scenario focused on behavior, then call the visual SDK at a meaningful rendered state through the step definition or test lifecycle. Use the SDK’s current instructions for your language and runner.

Can I add screenshot testing to existing BDD tests?

Yes. Add a named checkpoint to selected existing scenarios after their meaningful UI outcome, then establish and review the baseline. You do not need to rewrite the behavior specification.

Where should visual assertions go in a Gherkin scenario?

The assertion belongs in the automation behind the scenario, after the page reaches the state worth checking. Keep vendor calls out of the Gherkin text unless your team’s readers specifically benefit from expressing that detail there.

Should every scenario have a screenshot?

No. Select scenarios whose rendered result is valuable to review and whose UI state can be made repeatable.