ScreenshotNeo

BlogEngineering

4 Times Automated Tests Passed Even Though Bugs Were Present

A green test run only confirms the checks that ran met their expectations. Here are four ways real defects can still pass unnoticed—and how to close the gaps.

By the ScreenshotNeo team4 October 20267 min read

A passing automated test run means the checks that ran satisfied their encoded expectations in that run’s environment. It does not prove that every important behavior was checked, that the expectations were correct, or that the software has no defects. Four common gaps explain why bugs can coexist with green results: the broken path was never tested, the test expected the wrong result, a test double bypassed relevant production behavior, or the checks did not examine the kind of outcome that broke.

When a test passes, ask what requirement it exercised, which code and dependencies it reached, what result it expected, and what it actually observed. That is the boundary of the evidence the pass provides. ISTQB’s testing principles likewise caution that testing cannot prove the absence of defects. ISTQB testing principles

1. The broken path has no check

A test suite can pass every scenario it contains while missing the user journey or edge case where the defect occurs. For example, tests may cover successful registration but omit expired invitation links, or cover a normal checkout but skip a particular discount and payment combination. The tested scenarios remain green because they never exercise the failure.

AxonBuild describes audited examples in which the relevant path lacked a working test. That is an illustration of a coverage gap, not evidence that adding tests by itself guarantees correctness. AxonBuild’s account and examples

How to find the gap

  1. Start with the user-visible failure or requirement, and write down the exact conditions needed to reach it.
  2. Trace those conditions through the route, handler, service, and any relevant external boundary.
  3. Find a test that supplies those conditions and asserts the intended outcome. A nearby test with different inputs does not establish that this path is covered.
  4. Reproduce the defect, add a focused regression test, and confirm that the new test fails before the fix and passes afterward.

Coverage reports can show that code ran, but execution alone does not establish that the important condition or outcome was checked. Use coverage to identify areas for review, then inspect the scenario and assertion.

2. The test expects the wrong result

A test can faithfully check a faulty requirement, an incorrect assumption, or the implementation’s existing behavior. In that case the test passes because the bug has been encoded as the expected result.

AxonBuild gives a generated-test example that asserted division by zero should return zero. Treat it as that article’s illustration, not a universal pattern. AxonBuild’s account and examples

How to check the oracle

A test oracle is the source used to decide what the correct result should be. Check it independently of the code under test: use the written requirement, a domain rule, a trusted reference calculation, or a reviewed product decision. If the test simply repeats the implementation’s logic, both can share the same mistake.

  • For boundary cases, state the intended behavior before looking at the implementation.
  • For calculations, verify expected values with an independent method or known examples.
  • For error cases, assert the specified failure behavior rather than choosing a convenient fallback.
  • Have a reviewer compare the assertion to the requirement, especially when a test was generated or copied from existing code.

A useful test review traces requirement → input and setup → behavior exercised → expected result → observed result. A green assertion is meaningful only if the expected result is defensible.

3. A test double hides the production path

Mocks, stubs, and fakes make tests faster and isolate a unit, but they can also replace the very behavior that contains the defect. The test may verify that a substitute received a call while never running the production implementation or crossing the boundary where the real failure occurs.

AxonBuild reports a checkout suite that did not call the code that created a sale. This is a specific case reported by that article, not an independently verified finding. AxonBuild’s account and examples

Choose a test boundary that reaches the risk

Keep isolated unit tests for local decisions, and add an integration or end-to-end check when correctness depends on wiring, persistence, payment behavior, or another production boundary. A test need not use a live external service to reach production code: a controlled test database or a contract-checked service substitute can preserve the relevant application path. The key is to know exactly what has been replaced.

  • Review mocks for critical methods: does the test replace the operation whose behavior is in question?
  • Assert a meaningful outcome, such as a persisted order or emitted domain event, rather than only an interaction with a mock.
  • Use contract tests or integration tests for serialization, database mappings, and service interfaces where mismatches can escape isolated tests.
  • Keep a small number of end-to-end checks for the highest-risk user journeys; these catch wiring gaps but are usually slower and more environment-sensitive.

4. The suite checks function but not appearance

A workflow can perform its functional task while presenting a broken interface. A registration dialog might accept input and submit successfully even though buttons overlap, labels are clipped, or the layout fails at a supported viewport. Tests that check only values, events, or successful navigation do not automatically verify visual correctness.

Qt describes green tests alongside broken UI layouts, including misplaced or overlapping buttons when checks validate behavior but not appearance. Qt on balancing feature and visual testing

Match the check to the requirement

If the requirement is that a dialog submits valid data, assert that behavior. If it is also that controls remain visible and usable at supported sizes, add visual assertions or a review suited to that requirement. Screenshot comparison can reveal layout changes, but it needs stable rendering conditions and review of intentional changes; it does not replace functional tests or prove usability.

For browser UI checks, record the viewport, device scale, browser version, fonts, and relevant application state. Keep dynamic content stable or mask regions that are expected to vary. Investigate diffs rather than blindly updating baselines: a baseline update can otherwise bless the visual defect.

What a green test run actually tells you

A green result is evidence about the checks that executed, the inputs they used, their assertions, and the environment in which they ran. It rules out only failures those checks were capable of detecting under those conditions.

Question What to inspect
Was the affected behavior covered? Scenario inputs, branches, user journey, and edge conditions
Did the test reach the relevant code? Mocks, stubs, fakes, service boundaries, persistence, and wiring
Was the expected result correct? Requirement source, independent calculation, and assertion review
Did the check observe the right kind of outcome? Functional behavior, appearance, performance, accessibility, or other stated requirement

Neither a high pass rate nor 100% code coverage establishes that the product has no bugs. Pass rate describes outcomes among the checks that ran; coverage describes some aspect of execution. Neither by itself proves the scenarios were sufficient or the expectations right.

A practical review for escaped defects

  1. Reproduce the problem. Record the user action, inputs, environment, and observed result.
  2. Locate the missing link. Decide whether the scenario was absent, the expected result was wrong, relevant production behavior was replaced, or the test observed the wrong outcome.
  3. Add the narrowest useful regression check. Make it fail for the defect and pass for the intended behavior.
  4. Check the requirement independently. Review the expected result against its source, not just the current implementation.
  5. Run at the needed boundary. Use a unit, integration, end-to-end, visual, or manual check according to what must be observed.
  6. Keep complementary checks. A regression test helps prevent the same escape; exploratory review can still find behavior that was not anticipated in the test design.

Automated browser screenshots can help inspect visual outcomes during this work, but a screenshot is evidence for visual review, not a universal correctness verdict. If you capture a page yourself, use a browser automation tool, wait for the relevant state, set a fixed viewport, and compare the result with an appropriate baseline.

Or skip the browser setup

For a browser screenshot in a test or review workflow, ScreenshotNeo provides a single GET request. See the ScreenshotNeo API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing status. Its MCP server lets AI agents use screenshot, page-info, and PDF tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month, with no card.

Common questions

Why do my tests pass but the app does not work?

The failing path may not run in the suite, the assertion may encode the wrong result, a test double may bypass the faulty behavior, or the check may not observe the kind of failure you see. Trace the failing user action back to the test scenario and its boundary.

Does 100% test coverage mean there are no bugs?

No. Coverage does not show that assertions are correct or that every meaningful input and requirement was checked. It is a signal for reviewing what ran, not proof of defect-free software.

Should I remove mocks to make tests more realistic?

Not categorically. Mocks are useful for isolating units. Add tests at a broader boundary when the risk depends on the production code or integration that the mock replaces.

Can screenshot tests replace functional tests?

No. Screenshots help assess rendered appearance. Functional checks establish behavior such as validation, navigation, and data changes; use each where its evidence applies.