ScreenshotNeo

BlogGuides

Common BDD Pitfalls and How to Avoid Them

Avoid common BDD pitfalls by collaborating on concrete examples, writing behavior-focused scenarios, and keeping automation clear and reusable.

By the ScreenshotNeo team4 October 20269 min read

Behavior-Driven Development (BDD) is a collaborative way for a team to discover, agree on, document, and automate examples of desired system behavior. The most common BDD pitfall is treating it as a Gherkin-writing or test-automation task: without the conversations that establish shared understanding, scenarios can encode assumptions nobody agreed on.

A useful correction is to start with a small upcoming change and discuss concrete examples with product, testing, and development colleagues before automating them. Then write scenarios in business language, give each one a clear purpose, use controlled data, and keep implementation details in the automation layer. Cucumber describes discovery, formulation, and automation as iterative activities that connect shared understanding with implementation (Cucumber’s BDD overview).

1. Treating BDD as a test-writing ceremony

BDD is a practice for clarifying behavior, not a format for documenting tests after decisions have already been made. Cucumber explicitly distinguishes BDD from simply using Cucumber: teams first discuss examples, formulate them as structured documentation, and automate them to guide implementation (Cucumber introduction).

When a team skips discovery, the resulting scenarios may verify an unstated assumption rather than an agreed rule. A passing test cannot resolve a disagreement the team never surfaced.

What to do instead

  1. Pick a small, upcoming behavior change.
  2. Bring together people who understand the customer or business need, the tests, and the implementation. Cucumber describes these perspectives as the Three Amigos.
  3. Discuss examples, rules, boundaries, and unanswered questions before writing step definitions.
  4. Record the examples in language the group understands, then automate the examples that provide useful feedback.
  5. Revisit the examples as the team learns more or the product changes.

Discovery is not a one-time kickoff. It can continue while a team refines its understanding (Cucumber: Who does what?).

2. Writing Gherkin as a UI script

A scenario that describes every click and field can become tightly coupled to the current interface. If the UI changes but the promised behavior does not, a UI transcript may need editing even though the business rule is still the same.

Implementation-focused wording Behavior-focused wording
Given I visit the login page
When I enter “sam@example.test” in the email field
And I enter “correct-horse” in the password field
And I press the Login button
Given Sam has an active account
When Sam logs in with valid credentials
Then Sam sees the account overview

The second version describes a result that matters to a reader. Automation can still use the browser and perform those interactions behind the steps. Cucumber recommends describing behavior rather than implementation mechanics, while recognizing that imperative, UI-level tests can be appropriate in some contexts (Writing better Gherkin).

When UI details belong in a scenario

Include interface detail when the interface behavior itself is the requirement—for example, an accessibility interaction or a specific keyboard workflow. Otherwise, put selectors, clicks, and other mechanics in the automation layer. A useful review question is: would this wording still describe the expected behavior if the screen were redesigned?

3. Using vague or unrealistic examples

“Given a valid order” may conceal which conditions make an order valid. A concrete example makes assumptions visible: “Given Priya has a paid order worth $48.00” gives the group something specific to examine. Cucumber recommends relevant, concrete examples without unnecessary technical details (Cucumber examples).

Choose examples that explain the rule

  • Use meaningful people, amounts, dates, places, or other domain values when they clarify the behavior.
  • Include important boundaries, such as just below, at, and above a threshold, when those values determine the rule.
  • Use data that reveals the condition being discussed; avoid technical implementation data that adds no meaning.
  • Keep automated scenarios independent of mutable production records. Seed or create controlled test data so the example does not depend on a particular customer ID already existing.

Concrete does not mean hard-coded to one environment. The scenario should communicate a meaningful example, while the test setup should make that example reproducible.

4. Making one scenario explain everything

A long scenario can mix several rules, incidental setup, and unrelated outcomes. That makes it harder to understand what the example specifies and can make failures ambiguous: a scenario may fail because of a side detail rather than the behavior its name suggests.

Give each scenario an intention-revealing name and focus it on one rule or outcome. Cucumber’s Gherkin reference recommends 3–5 steps per example; practitioner Seb Rose suggests aiming for five lines or fewer for most scenarios. These are writing heuristics, not Gherkin syntax limits (Gherkin reference; Seb Rose, “Keep your scenarios BRIEF”).

Split by behavior, not by arbitrary line count

For example, a scenario about a rejected payment should not also cover a successful payment, a receipt email, and an account lockout. Give distinct outcomes their own examples. Keep essential setup, but move repeated or low-level setup into appropriate helpers or fixtures if that makes the example clearer.

Short is useful only when the reader can still understand the rule. Do not remove meaningful context just to hit a target number of lines.

5. Omitting business voices and shared language

BDD depends on shared understanding across business and technical roles. Product, testing, and development perspectives can expose different questions: what outcome matters, which cases need checking, and what implementation constraints or edge conditions need discussion. The Three Amigos is a useful way to bring those perspectives together, not a requirement to hold a formal meeting for every scenario (Cucumber: Who does what?).

Use the domain terms people already understand. If one concept is called “account,” “profile,” and “customer record” in different scenarios, agree on the intended term and use it consistently. Inconsistent language can make a shared specification look like it describes different concepts.

Have the people who understand the behavior review examples before and after automation. Scenarios are living documentation: their wording should evolve as the team’s understanding and the product change (Cucumber’s BDD overview).

6. Coupling step definitions to features or stacking actions

Step definitions that only make sense in one feature can lead to duplicated glue and growing maintenance costs. Cucumber advises organizing reusable steps around domain concepts and composing behavior with ordinary programming-language helper methods rather than calling one step definition from another (Cucumber anti-patterns).

For example, a helper such as createActiveAccount() can be used by setup code where appropriate. A step definition should translate a readable scenario step into that helper or other automation; another step definition should not invoke it as a shortcut.

Split conjunction steps when they hide actions

A step like “When I log in and update my address and place an order” hides multiple actions and makes it unclear which action failed. Split it into steps when each action matters to the example or has a separate outcome. Keep a combined step only if the phrase expresses one coherent domain action and remains easy to understand.

  • Prefer reusable steps that describe domain concepts, not selectors or page-specific mechanics.
  • Keep step definitions small enough to map clearly from scenario language to automation.
  • Put programming-language composition and shared setup in helpers, not in chains of step definitions.
  • When similar steps have drifted in wording, settle on consistent domain language before adding another near-duplicate definition.

7. Using Scenario Outlines without meaningful examples

A Scenario Outline is a template that Cucumber runs once for each row in its Examples table; the outline itself is not run as a single scenario (Gherkin reference).

Use an outline when several concrete data combinations illustrate the same rule. For instance, an outline can show that an order qualifies for free shipping at several relevant basket totals. Keep the table small enough to scan and make each row an intentional case. If rows actually demonstrate different rules or outcomes, separate them into named scenarios instead of hiding that difference in a large table.

8. A practical review checklist

  • Discovery: Did the relevant business and delivery perspectives discuss examples before automation?
  • Purpose: Does each scenario explain a behavior or rule that matters?
  • Language: Can a non-technical stakeholder follow the wording, and are domain terms consistent?
  • Specificity: Are examples concrete enough to reveal assumptions and meaningful boundaries?
  • Control: Does automation create or seed the data it needs rather than relying on mutable production records?
  • Focus: Does each scenario have one clear purpose, without unrelated outcomes or incidental detail?
  • Automation boundary: Are UI mechanics and selectors kept out of behavior wording unless the UI itself is the requirement?
  • Reuse: Are shared domain actions composed with helpers rather than step-definition chains?
  • Maintenance: Will the scenario be reviewed when the product or shared understanding changes?

9. Common BDD troubleshooting

Symptom Likely cause Correction
Scenarios pass but stakeholders disagree about what the feature should do. Automation started before the team discussed concrete examples and rules. Pause new glue work, discuss examples with the relevant roles, record decisions and open questions, then update scenarios.
A small UI change breaks many scenarios. Scenarios describe selectors, clicks, or page layout rather than behavior. Move mechanics into automation helpers and rewrite examples around the promised outcome. Keep UI-level scenarios where the interface behavior is itself under specification.
Scenarios are hard to read or failures are hard to diagnose. A scenario covers several behaviors or includes incidental setup. Give each rule a focused scenario, split independent outcomes, and keep the title aligned with the behavior being checked.
Tests pass locally but fail when data changes or in another environment. The automation depends on a particular mutable record or uncontrolled state. Create isolated, controlled test data and make setup explicit or reliably reusable.
Step definitions are duplicated or keep multiplying. Glue is coupled to feature files, domain terms vary, or shared actions are chained as steps. Normalize wording, organize reusable steps around domain concepts, and compose automation with ordinary helper methods.
A Scenario Outline table is confusing. Rows represent different rules, or the table contains too many incidental combinations. Keep only intentional examples of the same rule; use separately named scenarios for distinct behavior.
One step performs several actions and the failure location is unclear. A conjunction step bundles independent actions or preconditions. Split it into focused steps when the actions matter separately, or move mechanical composition into a helper while retaining clear scenario intent.

10. FAQ

Does every BDD scenario need to be automated?

No. BDD connects examples and shared understanding with implementation, and automation can check selected examples. The useful question is whether automating a scenario gives the team feedback worth maintaining.

Is declarative Gherkin always better than imperative Gherkin?

No. Behavior-focused wording is generally easier to keep as shared documentation, but UI-level imperative tests can be appropriate when the interface mechanics are what the team needs to verify.

Is there a maximum number of steps allowed in a scenario?

There is no hard Gherkin limit implied by the writing recommendations. The suggested 3–5 steps are a readability heuristic; clarity and a single purpose matter more than an exact count.

When should I use a Scenario Outline?

Use one when multiple data rows demonstrate the same behavior rule. Use separate scenarios when the cases need different explanations or represent different rules.

Or skip the browser setup

BDD scenarios and browser automation are separate concerns, but if your work also needs screenshots of test pages or application states, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a screenshot or PDF. The capture can accept cookie and consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before taking the shot; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say the page verdict and whether the request was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Here is the one-call cURL example; see the ScreenshotNeo documentation for the API options:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Free includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots. Sign up for free and get 1,000 screenshots a month with no card.