ScreenshotNeo

BlogHow-to

How to Implement BDD Testing for Test Automation

Learn how to implement BDD through team discovery, Gherkin examples, step definitions, and a maintainable automation feedback loop.

By the ScreenshotNeo team4 October 202610 min read

To implement behavior-driven development (BDD) for test automation, start by agreeing on concrete examples of the behavior a user needs. Write those examples as readable scenarios, connect their steps to automation code, and use the results to guide implementation. The tool is part of the workflow; BDD is the collaboration and feedback loop around the examples.

This guide uses Cucumber and Gherkin to show the mechanics. The same workflow applies with another runner that can express readable examples and connect them to your application.

1. Start with behavior and examples

Choose a small upcoming story or behavior. Bring together product or business, testing, and development perspectives to discuss what the user is trying to do, what should happen, and which cases might change the answer. Cucumber describes this collaborative work as Discovery, Formulation, and Automation. Example Mapping and Event Storming are two techniques it names for surfacing examples and questions. Cucumber’s BDD guide

Do not begin by translating a ticket into a sequence of clicks. First clarify the rule. For account access, the useful question might be: “What should happen when a registered customer provides valid credentials?” Discuss invalid credentials, locked accounts, and any unanswered product rules too. If the group cannot agree on an expected result, record the question and resolve it before encoding an assumption in a test.

Run a focused discovery discussion

  1. Pick one behavior that is small enough to understand and automate.
  2. Ask for concrete examples: a typical success case, relevant alternatives, and important failure cases.
  3. Have product, testing, and development identify rules, boundaries, dependencies, and open questions.
  4. Agree on observable outcomes that a person can recognize and automation can check.
  5. Keep unresolved questions visible and return to discovery when an example reveals uncertainty.

Cucumber calls its cross-functional discussion a Three Amigos conversation, while noting that the group need not consist of exactly three people or meet only once. Initially, have the whole team shape the language. Later, a developer or automation owner and tester can draft examples together if product or business representatives actively review them. Cucumber guidance on team roles

2. Formulate examples as Gherkin scenarios

Gherkin is a structured language for writing examples that people can read and Cucumber can execute. A feature groups related scenarios. Within a scenario, Given describes initial context, When describes an event, and Then describes the expected outcome. And and But continue a sequence. Cucumber matches each step to a step definition in code. Gherkin reference · Step definitions

Feature: Account access

  Scenario: A registered customer signs in
    Given a registered customer
    When the customer signs in with valid credentials
    Then the account overview is available

Save the example in a .feature file in source control alongside the software. Keep it tied to a meaningful business rule, and make the expected result observable. For example, “the account overview is available” can be implemented by checking a page heading, an authenticated API response, or another reliable product outcome, depending on the system under test.

Cucumber recommends aiming for three to five steps per example, while allowing as many as needed. Treat that as a readability guide, not a hard limit: if a scenario grows long, check whether it combines multiple behaviors or has become a script of implementation details. Cucumber’s BDD guide

Keep scenarios behavior-focused

Prefer “When the customer signs in with valid credentials” to “When Bob opens /login, types into the email field, types into the password field, and clicks the blue button.” The step definition can manage URLs, selectors, and interaction details. This keeps the shared specification focused on behavior and less dependent on a particular interface. Cucumber guidance on better Gherkin

Each scenario should explain one behavior clearly enough that a failure has a useful meaning. Avoid checks of implementation details that can change without changing the behavior users care about. A scenario can use arguments or a data table when a step needs structured values; keep those values understandable in the example rather than hiding the rule in helper code. Gherkin reference

3. Connect the steps to automation

A step definition is code associated with a Gherkin step. It performs setup or an action against the system under test, then checks the outcome. The exact runner setup depends on your language and application; Cucumber documents implementations and APIs for its supported language ecosystems. Use the official documentation for the matching implementation and your chosen browser, API, or application test integration. Cucumber installation guides · Cucumber API reference

Illustrative JavaScript step definitions, assuming your project has installed and configured the JavaScript Cucumber package and provides the named application helpers:

const { Given, When, Then } = require('@cucumber/cucumber');
const assert = require('node:assert/strict');

Given('a registered customer', async function () {
  this.customer = await createRegisteredCustomer();
});

When('the customer signs in with valid credentials', async function () {
  this.response = await signIn(this.customer.email, this.customer.password);
});

Then('the account overview is available', async function () {
  assert.equal(this.response.status, 200);
  assert.equal(this.response.body.page, 'account-overview');
});

The helpers in this example (createRegisteredCustomer and signIn) are application-specific; implement them using the actual test environment. This sketch shows the mapping, not a drop-in test for an unspecified product. See the step-definition documentation for syntax and the relevant Cucumber runner setup.

Make the test state reliable

  • Arrange known data for each scenario, or use isolated fixtures, so one run does not depend on another.
  • Keep setup in appropriate hooks or helper functions when it is technical plumbing rather than part of the business example.
  • Use a stable test environment and wait for the specific outcome the scenario needs; avoid arbitrary sleeps when a condition can be observed.
  • Ensure failures show which behavior or expected outcome was wrong. Avoid swallowing errors in shared steps.
  • Keep step definitions reusable where that improves clarity, but do not make business language so generic that its meaning becomes opaque.

4. Run one example, implement, and refine

Automate one useful example at a time. Run it, inspect the result, and implement or adjust the behavior. When a failing example exposes a product ambiguity, return to discovery and agree on the rule. When understanding changes, update both the shared specification and the implementation so they stay aligned. Cucumber describes BDD as iterative collaboration with automatically checked documentation. Cucumber’s BDD guide

  1. Add one agreed scenario to a feature file.
  2. Run the Cucumber runner and identify any undefined or failing steps.
  3. Implement the step definitions and required test setup.
  4. Run the scenario against the system under test and use its feedback to complete the behavior.
  5. Review the example with the people who own the rule; refine wording, expected outcomes, or scope as needed.
  6. Repeat for the next useful example, then run the relevant feature or suite in the team’s normal build workflow.

BDD is not a claim that a particular tool or syntax guarantees better software. The practical aim is to make important behavior explicit, shared, and checkable. Cucumber reproduces Fred Brooks’s sentence, “The hardest single part of building a software system is deciding precisely what to build.” Cucumber’s BDD guide

5. Choose a tool by fit

Cucumber and Gherkin are one common route, but tool choice should follow the team’s needs. Compare candidates on the programming-language ecosystem, whether people can express and execute readable examples, how well the runner integrates with the application under test, and whether the mapping from steps to code stays understandable. The cited Cucumber sources explain its workflow and mechanics; they do not establish a comparative ranking of competing tools.

Before adopting a runner, build a small vertical slice: one feature, one scenario, one real integration with the system under test, and a useful failure report. Confirm that the team can run it locally and in its normal automation environment, and that the scenario still reads clearly to someone outside the automation implementation.

6. Keep the suite maintainable

  • Write examples for rules, not screens. Keep interface mechanics in step definitions or lower-level helpers.
  • Keep one behavior per scenario. Split unrelated outcomes so failures point to a clear issue.
  • Review the language as the product evolves. Stale examples are misleading documentation; update or remove them when behavior changes.
  • Use shared helpers with restraint. Reuse setup and interactions, but make each step’s action and consequence easy to trace.
  • Make test data explicit and isolated. Avoid hidden dependencies between scenarios and unpredictable shared state.
  • Keep the automation proportionate. Automate examples that clarify and protect meaningful behavior; do not turn every technical detail into business-facing Gherkin.

7. Troubleshooting BDD automation

Symptom Likely cause What to do
A step is reported as undefined No step definition matches its text, or the definition file is not loaded by the runner. Check the runner’s configured paths, the expression syntax, and the step wording. Add or correct the definition, then rerun the single scenario.
A scenario passes alone but fails in a suite Shared state, order dependence, or test data collisions. Give scenarios isolated setup and cleanup, use unique or resettable data, and remove assumptions about execution order.
A browser step times out intermittently The test waits for a fixed delay, a network or application condition is variable, or the environment is overloaded. Wait for a meaningful state or response with a bounded timeout, inspect environment health, and capture diagnostic output on failure.
The scenario reads like a click script Interaction details have leaked into the shared specification. Rewrite steps in terms of user intent and outcomes; move selectors, URLs, and low-level actions into definitions or helpers.
A step definition is reused but confusing Over-generalized wording or hidden conditional behavior obscures what the step does. Make the business meaning explicit, split distinct actions, or pass a clear argument or table value.
A test asserts the wrong result The example encoded an assumption before the rule was agreed, or product behavior changed. Return to the relevant product, testing, and development collaborators, agree on the example, then update the scenario and implementation together.
The report is difficult to interpret Failures are swallowed, assertions are too broad, or several behaviors share one scenario. Assert the expected outcome directly, preserve useful errors, and divide scenarios when each behavior needs a distinct result.

8. Performance, reliability, and cost

BDD itself does not determine suite speed or financial cost. Those depend on the runner, the system under test, the environments, and how many examples you execute. Keep feedback useful by starting with the scenario or feature being changed, then run the broader set in the team’s established workflow. Avoid arbitrary waits and unnecessary setup work; isolate data and dependencies so retries do not conceal unstable behavior.

There is no benchmark in the cited material that supports a promised defect reduction, time saving, or return on investment. Evaluate your own suite by whether scenarios clarify important rules, provide actionable failures, and remain aligned with the product. Consider the maintenance cost of each example and its automation when deciding what belongs in the suite.

9. Capture visual evidence for browser scenarios

For browser-based scenarios, a screenshot can help a developer inspect the page state behind a visual failure. Capture at the point where the scenario’s expected state should be visible, and retain the scenario name and useful diagnostics with the image. Screenshot capture complements assertions; it does not establish that the behavior is correct on its own.

Capture a page with your own browser setup

For a local Cucumber browser test, use the browser automation library already integrated with your runner. The following Playwright example is a runnable standalone Node.js script after installing Playwright in the project. It opens a page, waits for a meaningful page state, and saves a screenshot.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  try {
    await page.goto('https://example.com', { waitUntil: 'domcontentloaded', timeout: 30000 });
    await page.locator('body').waitFor({ state: 'visible', timeout: 10000 });
    await page.screenshot({ path: 'scenario.png', fullPage: true });
  } finally {
    await browser.close();
  }
})();

Install the dependency with npm install --save-dev playwright, then install its browser with npx playwright install chromium. Replace the example URL and readiness condition with the application and observable state for your scenario. Playwright documents navigation and screenshot options in its screenshot guide and Page API.

ScreenshotNeo for capture without browser setup

Or skip the browser setup:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the request options. ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. Cookie banners and consent popups are accepted or removed before the shot, along with newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the shot was billed. Its MCP server lets AI agents using Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does BDD require Cucumber?

No. BDD is the collaborative practice of clarifying behavior through examples and checking those examples. Cucumber and Gherkin provide one way to express and automate them.

Should every test be written in Gherkin?

No. Use shared scenarios where they help express meaningful behavior. Keep lower-level technical checks in the testing layers and formats that make them clear.

Who should write the scenarios?

People who understand the product rule and the automation should collaborate. Product or business representatives should review the wording and outcomes, even when developers and testers draft the examples.

What should happen when an example is ambiguous?

Pause automation of that assumption and return to the relevant collaborators to agree on the behavior. Then update the example and its implementation.