ScreenshotNeo

BlogGuides

False Positives vs. False Negatives in Software Testing

Learn how false positives and false negatives mislead software teams, how flaky tests fit in, and how to investigate failures and find gaps in test coverage.

By the ScreenshotNeo team4 October 20269 min read

A false positive reports a defect when the tested object has no such defect. A false negative fails to identify a defect that is present. In a test suite, a red run is not by itself proof that production code is wrong, and a green run is not proof that the code has no defects. Judge each result against the intended behavior or specification and the actual behavior.

The terminology follows the [ISTQB glossary](https://glossary.istqb.org/en_US/term/false-positive-result) and [its false-negative entry](https://glossary.istqb.org/en_US/term/false-negative-result). This guide explains what each error means, how flaky tests complicate diagnosis, and how to improve confidence without treating a single test score as a guarantee.

1. What do false positive and false negative mean?

First state what “positive” means. Here, a positive result means the test reports a defect or detects a failure condition.

Reality, judged against the specification Test reports a defect (positive) Test does not report a defect (negative)
A defect is present True positive: the test detects it False negative: the defect is missed
No defect is present False positive: the test reports a defect anyway True negative: the test does not report a defect

For example, suppose a function is specified to return an empty list when there are no matching records. A test that fails because it expected an empty list but the function returned one suggests a real defect. If the function correctly returns an empty list but a test fixture or assertion incorrectly expects null, the test reports a problem even though the implementation matches the specification. Conversely, if the function incorrectly returns null but the test only checks that the call completes, the defect may pass unnoticed.

In everyday CI language, teams often use “false alarm” for an unjustified red build and “missed defect” for an unjustified green build. Keep the reference condition explicit: intended behavior and actual behavior. Test-runner color alone cannot tell you which side is wrong.

2. False positives: a test reports a defect that is not there

A false positive can come from faulty test code, a mistaken expected result, bad test data, or conditions outside the implementation under test. The production code may be correct while the test environment, fixture, or expectation is not.

Common causes

  • Incorrect expectation: the assertion encodes behavior that the specification does not require.
  • Bad or stale fixture: test data is invalid, inconsistent, or no longer represents the scenario.
  • Shared state or order dependency: another test changes a global, database row, clock, or file that this test assumes is untouched.
  • Timing assumptions: the test expects an asynchronous operation to finish within an unrealistically tight bound.
  • Uncontrolled external systems: a test depends on a network service, queue, browser, or third-party response that varies.
  • Overly strict numeric comparison: floating-point calculations differ by a small precision amount even though the result is acceptable.
  • Infrastructure or resource pressure: parallel load, a slow worker, or a transient environment problem causes a failure unrelated to the code change.

How flaky tests fit

A flaky test passes and fails intermittently without a clear deterministic cause. A failure on unchanged code may be a false alarm, but “flaky” describes the inconsistent behavior; it does not prove that every failure is false. A real defect can also be intermittent, especially when it depends on concurrency or timing. The [pytest documentation on flaky tests](https://docs.pytest.org/en/stable/explanation/flaky.html) notes that unreliable signals can erode trust, hide real failures, and waste investigation time.

pytest lists uncontrolled system state, test-order dependencies, parallel execution, overly strict timing assertions, and floating-point comparisons among potential causes. Improve isolation and cleanup, make order dependencies visible by varying order, use appropriate tolerances for numeric values, and reproduce the failing conditions. Reruns help determine whether a failure is intermittent; they do not identify its cause.

Terminology can vary between organizations. Chromium’s [CQ documentation](https://chromium.googlesource.com/chromium/src/+/main/docs/testing/flakiness.md) uses “false negative” locally for a flaky failure that should have passed, and describes retries as a way to reduce disruption while allowing flaky tests to land. That usage differs from the ISTQB definition above. When communicating across teams, define what a positive result means instead of assuming the labels are universal.

3. False negatives: a real defect passes the suite

A false negative occurs when the tested object contains a defect but the test does not identify it. This often happens because a test does not exercise the relevant case, or because its assertions do not distinguish correct behavior from incorrect behavior.

Typical coverage and assertion gaps

  • Missing scenario: tests cover ordinary inputs but omit empty, boundary, malformed, or unusually large inputs.
  • Weak assertion: the test checks that a function ran, returned any value, or did not throw, rather than checking the required result.
  • Wrong observation point: the test asserts an internal detail while the defect affects externally visible behavior, or it never checks a critical side effect.
  • Unrepresentative fixture: test data never exercises the branch, role, permission, state transition, or failure mode where the defect occurs.
  • Mock hides behavior: a stub replaces the code or integration that contains the defect.
  • Test not run: the relevant test is excluded, misconfigured, skipped, or absent from the CI job that gates the change.

A passing suite establishes only that the executed assertions passed for the inputs and conditions exercised. It says nothing directly about untested behavior.

4. How to compare the risks

Neither error type is always more costly. The consequences depend on what the test gates and what a mistake would affect. Consider these questions when deciding how much investigation or test strength a check needs:

Decision axis Questions to ask
Impact What happens if the defect ships? What is the cost if a correct change is blocked?
Likelihood and other detection How plausible is this defect class? Do later checks, monitoring, or review catch it?
Decision point Is the test for quick local feedback, a merge gate, or a release or safety gate?
Investigation cost How long does diagnosis take, and can the failure be reproduced?
Recovery Can a defect be rolled back or detected downstream, or would its effects be difficult to reverse?

Use these axes to set practical priorities, not as a standardized score. There is no universal numeric ranking for the cost of false positives versus false negatives: a noisy unit test and a missed defect in a safety-critical release have different stakes.

5. Investigate a suspicious CI failure

  1. Preserve the first failure. Keep its logs, test order, inputs, environment details, and commit identifier. Establish whether code and conditions really stayed constant.
  2. Reproduce the run. Run the failing test in isolation, then under the original suite and concurrency conditions. Check whether order, shared state, or external dependencies change the result.
  3. Rerun with a purpose. Repeated failures under the same conditions suggest a deterministic issue; mixed results suggest intermittency. Record the original failure even if a retry passes.
  4. Check the test and the specification. Compare the expected result with the documented behavior, then inspect implementation, fixtures, cleanup, timing, and assertions.
  5. Fix the source of the mismatch. Change production code when it violates the intended behavior; correct the test when its expectation or setup is wrong; clarify the specification when behavior is ambiguous.
  6. If you must quarantine it, make that temporary and visible. Assign an owner and follow-up. pytest warns that permanently marking unreliable tests as non-strict expected failures can hide real regressions.

A retry policy is a containment tool. Chromium documents the tradeoff: retries can reduce disruption from flaky tests but can also let failures pass through. Keep retry results observable and use them alongside reproduction and root-cause work.

6. Find false negatives with stronger tests and mutation testing

Start from behavior that matters, then add checks that would fail if that behavior regressed. For each requirement, ask what incorrect output or side effect could still satisfy the current assertions. Include meaningful boundary and failure cases, and ensure the test actually runs in the relevant CI gate.

Mutation testing probes test sensitivity by making small deliberate changes to code and checking whether the suite catches them. A mutant is “killed” when a test fails because of the change; it “survives” when the suite still passes. Microsoft’s [Stryker.NET guidance](https://learn.microsoft.com/en-us/dotnet/core/testing/mutation-testing) recommends reviewing surviving mutants for gaps or weak assertions and focusing on higher-risk code rather than chasing a perfect score. Google’s [Testing Blog article on mutation testing](https://testing.googleblog.com/2021/04/mutation-testing.html) emphasizes that tests added to kill mutants should themselves be valuable.

Use mutation results carefully

  • A surviving mutant is a prompt to inspect whether a meaningful behavior is untested; it is not automatically proof of a production defect.
  • Some mutants are equivalent with respect to observable behavior, so no useful test can distinguish them.
  • Mutation operators sample possible changes; they do not enumerate every defect a program could contain.
  • A mutation score is not the probability that the suite will catch real defects. Use it to guide review, especially around important logic and business rules.

The [ISO overview of software testing concepts](https://www.iso.org/standard/79428.html) identifies ISO/IEC/IEEE 29119-1:2022 as general concepts and says Part 1 is informative. This overview is terminology context; citing it does not certify a test suite or prescribe a universal error-cost calculation.

7. Practical checklist

  • Write down the intended behavior before labeling a red result false.
  • For a green result, ask whether the assertions would detect the important wrong outcomes.
  • Keep flaky failures visible; record the first result and investigate intermittent causes.
  • Isolate tests from shared state, ordering, external services, and uncontrolled time where practical.
  • Use assertions that express outcomes and use suitable tolerances for approximate values.
  • Apply mutation testing selectively to critical logic; inspect survivors rather than optimizing a score.
  • Make retries and quarantines visible, owned, and reviewable.

8. Or skip the browser setup

If your test workflow needs reference screenshots of web pages, you can run a browser capture yourself or make one API request. ScreenshotNeo is a website screenshot API and MCP server for developers by Yorker Media. The API returns a PNG, JPEG, WebP, or PDF; see the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers indicate the page verdict and whether the shot was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, no card required.

9. Frequently asked questions

Can one test be both a false positive and a false negative?

For one result under one stated condition, it is one or the other relative to the chosen defect condition. Across different inputs, runs, or interpretations of the specification, a test can produce both kinds of error.

Does a flaky test always create false positives?

No. Intermittency makes results less dependable, but a failure may expose a real timing or concurrency defect. Reproduce it and compare the behavior with the specification.

Does a 100% mutation score prove a suite catches every defect?

No. Mutation testing samples particular code changes, and its score does not measure all possible defects or guarantee detection in production.

Should teams retry every failed test?

Retries can reduce disruption from intermittent failures, but a later pass does not explain the first failure. Retain the initial result and make retry outcomes visible.

Is a red test always a false positive if the code did not change?

No. Environment, inputs, order, and timing may have changed, or the defect may be intermittent. Unchanged code alone does not establish that the failure is false.

Terminology note: The FDA-hosted software terminology glossary dates to August 1995 and is historical terminology context, not current regulatory guidance. See the FDA software validation resources for current context.