ScreenshotNeo

BlogGuides

Common Types of Software Bugs and How to Find Them

Learn to recognize common software bugs, reproduce them with focused tests, and combine static and dynamic checks to find and fix defects.

By the ScreenshotNeo team4 October 202610 min read

Software bugs are defects that make a program behave differently from its intended behavior. They can come from incorrect requirements or design as well as from code. The practical way to find them is to describe the expected behavior, reproduce the failure with a small test, and combine code review, static checks, and safe tests of the running program. No single technique finds every defect.

This guide groups common bugs by the failure a developer can observe, explains how to investigate each, and gives a repeatable discovery and verification workflow. The categories are useful for triage, not an exhaustive or mutually exclusive taxonomy. For a formal classification of security weaknesses and vulnerabilities, see NIST’s Bug Framework.

1. Logic and requirements bugs

A logic bug occurs when the implementation produces the wrong result for a valid situation. A requirements bug occurs when the system implements a rule that is incomplete, ambiguous, or wrong. These can look identical from a user’s perspective: the software does something other than what the user or business rule requires.

  • Examples: a discount applies twice; a permission check allows the wrong role; a report excludes a valid record; a workflow skips a required approval.
  • Look for them: write down the intended rule in observable terms, then test ordinary cases, invalid cases, and boundaries. Use black-box tests to check externally visible behavior, and structural tests when a particular code path needs coverage.
  • Investigate the specification: ask whether the expected behavior is actually correct and whether important misuse cases or trust boundaries were considered. A test that faithfully asserts the wrong requirement will preserve the bug.

When a defect is confirmed, keep a regression test that fails before the fix and passes after it. NIST’s developer verification guidance includes black-box testing and historical test cases among its recommended techniques.

2. Input and boundary bugs

Input bugs appear when software handles empty, malformed, unexpected, or extreme values incorrectly. Many defects occur at transitions: zero versus one, just below versus at a limit, or the first and last item in a collection.

  • Examples: an empty list causes an exception; a maximum-length field is truncated incorrectly; a negative quantity bypasses validation; a timestamp at a daylight-saving transition is interpreted incorrectly.
  • Look for them: test minimum and maximum valid values, values just outside those limits, missing fields, unusual encodings, malformed data, and combinations of optional fields.
  • Use fuzzing when appropriate: send generated, unexpected, or specially crafted inputs through a parser or interface and watch for crashes, hangs, invariant failures, or unexpected output.

Keep fuzzing and other dynamic tests in an isolated test environment. NIST describes fuzzing as dynamic analysis that exercises software with random and crafted input, and advises running dynamic analysis in test environments rather than live systems. See the NIST analysis overview.

3. State, ordering, and concurrency bugs

These failures depend on the order or timing of events. A program may work when operations happen one at a time, then produce stale data, duplicate work, or corrupted state when requests overlap.

  • Examples: two workers update the same counter and one update is lost; a retry creates a duplicate payment; a UI displays a response from an older request after a newer one has completed.
  • Look for them: record the sequence of operations and relevant timestamps; reproduce with controlled ordering; repeat the scenario under varied scheduling; inspect shared mutable state, locking, retries, and idempotency.
  • Use suitable tools: static analysis that checks races can help with parallel software, while runtime tests can expose failures that depend on actual execution. Neither result alone establishes that all schedules are safe.

Make concurrency tests deterministic where possible by controlling barriers, fake clocks, or test doubles. A test that merely sleeps for a guessed duration can be flaky and may miss the race entirely.

4. Performance and resource bugs

Performance defects include slow responses, excessive memory or CPU use, resource leaks, hangs, and failure under load. They may only appear with realistic data volumes or sustained traffic, so a small happy-path test is not enough.

  • Examples: a query becomes slow as a table grows; memory rises on every request; a queue grows faster than workers can drain it; a request never returns after a dependency stalls.
  • Look for them: measure representative workloads, include larger-than-usual inputs, observe resource use over time, and exercise overload or dependency failure in a controlled environment.
  • Compare against a baseline: a measurement is useful when the workload and environment are recorded and can be repeated. Avoid treating one run as a universal performance guarantee.

NIST’s verification guidance includes dynamic testing and tests for denial-of-service and overload scenarios. Apply load carefully in an environment you control.

5. Security weaknesses

A security bug can allow unauthorized access, unsafe data flow, secret exposure, or exploitation through an included dependency. Security defects can originate in requirements and architecture, not only in a line of code.

  • Examples: an authorization check is missing on one endpoint; untrusted input reaches a command or query; a credential is committed to source control; a dependency with a known weakness is shipped.
  • Look for them: threat model important assets and trust boundaries; review access control and data flow; run static checks and secret detection; test relevant interfaces dynamically; inspect included components.
  • Validate findings: confirm whether the reported path is reachable and whether the alleged impact is possible in the actual deployment. Prioritize by impact and exploitability, not scanner severity alone.

Static analysis can prioritize review, but it has blind spots. OWASP notes that high-confidence automatic detection of many application security flaws remains beyond the state of the art. Read the finding, inspect its context, and consider code the tool cannot reliably classify. See OWASP’s static code analysis overview.

6. Choose detection methods by what they can observe

Static analysis inspects code without running it. Dynamic analysis executes software and observes its behavior. They answer different questions and are more useful in combination than as substitutes.

Method What it can reveal Limits and good use
Manual review Incorrect assumptions, confusing logic, unsafe data flow, and design gaps Depends on reviewer context and attention; focus review on changed, complex, or high-risk paths.
Unit and structural tests Expected behavior of functions and selected code paths Only cover the cases and paths represented by tests; include boundaries and confirmed regressions.
Black-box tests Whether externally visible behavior matches a requirement May not identify the underlying cause; pair a failing case with focused diagnosis.
Static analysis Patterns and properties detectable from source without execution Can produce false positives and miss defects; inspect findings and tool blind spots.
Dynamic tests and runtime observation Failures that occur along executed paths, including runtime behavior and resource problems Unexecuted paths remain unobserved; use a representative, non-production environment.
Fuzzing Unexpected input cases that can trigger crashes, hangs, or violated invariants Needs a useful harness and oracle for recognizing bad behavior; run safely outside production.
Dependency and service review Risks in components included or called by the application Inventory what is actually shipped and configured; a clean application scan does not validate every dependency or service.

NISTIR 8397 recommends a range of broadly applicable verification techniques and explicitly does not claim to cover every possible verification method. Tool performance also varies by bug class and complexity. NIST’s SATE VI report found lower recall and discrimination in its more complex C track than in its less complex Java track; that is a finding from that evaluation, not a universal ranking of languages or tools. See the SATE VI report.

7. A repeatable bug-finding workflow

  1. Describe the failure. Record expected behavior, actual behavior, environment, input, and exact reproduction steps. Keep observations separate from guesses about root cause.
  2. Check the requirement and design. Verify that the expected behavior is specified correctly. For security-sensitive behavior, identify assets, trust boundaries, and misuse cases.
  3. Make the smallest reproducer. Reduce the input and steps until the failure is easy to repeat. Prefer a focused black-box case for user-visible behavior, or a structural test when a code path must be isolated.
  4. Run low-cost checks early. Review the changed code and run relevant unit tests, static analysis, and secret checks. Triage findings rather than accepting or dismissing them automatically.
  5. Exercise runtime behavior safely. Run integration, load, or fuzz tests in a representative test environment. Check dependencies and services involved in the path.
  6. Fix the cause, then verify. Rerun the reproducer, regression suite, and checks relevant to the defect. Confirm the test fails without the fix when practical and passes with it.
  7. Keep the regression case. Preserve the test and enough context to diagnose a future recurrence. A clean scan is not proof that this defect is fixed or that the system has no other bugs.

When choosing between methods, compare whether they inspect code or exercise running behavior, which paths and bug classes they reach, feedback timing, setup effort, and how findings are validated. Combining techniques is useful because their coverage differs.

8. Browser-visible bugs: reproduce the page state

Some defects only appear in the rendered page: a consent overlay hides a control, a popup shifts a layout, or a responsive state breaks a component. A screenshot is useful evidence of what the browser displayed at a particular viewport and time, but it cannot establish the underlying cause or replace functional and accessibility tests.

For a manual check, open the page in the target browser and viewport, reproduce the state, capture the screen, and record the URL, viewport, browser, and steps alongside the image. For repeatable checks, automate the browser capture as part of a test workflow and compare only after controlling dynamic content, timing, and viewport differences.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. For example, this cURL request saves a WebP screenshot of a page:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp

See the ScreenshotNeo API documentation for request details. ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents using Claude, Cursor, or another MCP client. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo and get 1,000 screenshots a month free with no card.

10. Troubleshooting: why bug checks miss or misreport defects

Symptom Likely cause What to do
The bug cannot be reproduced Missing environment detail, timing dependency, stale data, or an unrecorded input Capture versions, configuration, data shape, timestamps, and exact steps; reduce the case and control dependencies.
The test passes locally but fails elsewhere Environment or configuration differs, or a timing assumption is exposed Compare runtime versions, locale, timezone, feature flags, services, and resource limits; make setup explicit.
A regression test is flaky Uncontrolled time, randomness, ordering, shared state, or external services Inject a clock or deterministic seed, isolate state, control event order, and replace unreliable dependencies in the focused test.
Static analysis reports many irrelevant findings Rules may not understand project context or the path may be unreachable Validate each finding, tune rules narrowly, document accepted risk, and keep checks for relevant high-risk patterns.
The scanner reports nothing, but a defect remains The bug class, path, or context is outside the tool’s detection ability Add behavior-focused tests and review; use another technique suited to the failure. A clean scan is not proof of correctness.
Fuzzing finds crashes that do not matter—or finds nothing The harness lacks a useful oracle, input space is constrained, or the tested path is narrow Define invariants and expected failures, verify the harness reaches meaningful code, and seed it with representative inputs.
Performance tests disagree between runs Workload, warm-up, host load, cache state, or measurement method changed Record and stabilize the workload and environment, repeat measurements, and compare like with like.
A security finding is hard to prioritize The report omits reachability, privilege, data sensitivity, or deployment context Trace the data or control path and assess plausible impact in the actual configuration; then record the rationale.

11. Performance, reliability, and cost of detection

Run fast, deterministic checks frequently, such as focused tests and suitable static checks. Reserve slower integration, load, and fuzzing work for scheduled or targeted runs when setup and compute costs are higher. The right balance depends on the codebase and risk; the sources do not establish one universal schedule.

  • Performance: measure checks in your own environment. Large scans, broad integration suites, and long fuzzing campaigns consume more time and compute than a focused test.
  • Reliability: keep tests repeatable by controlling time, randomness, external services, and shared state. Treat findings as evidence to investigate, not a complete inventory of defects.
  • Cost: account for developer review time, test infrastructure, and the cost of failures escaping into use. Prioritize checks around impact and likelihood rather than maximizing tool count.
  • Coverage: keep a record of tested paths and known limitations. Multiple techniques reduce blind spots, but none guarantees the absence of bugs.

12. FAQ

Are bugs always caused by coding mistakes?

No. Requirements, design choices, configuration, dependencies, and interactions between components can all produce defects.

Can static analysis prove software is bug-free?

No. It can detect some classes of issues without executing the program, but it has limits and may miss defects or report findings that need human validation.

Should fuzzing run against production?

Use a controlled test environment. Fuzzing deliberately sends unusual inputs and can crash or overload the target.

How should a team prioritize a bug?

Consider user impact, security impact, reachability, frequency, and the consequences of failure, then document the decision and verify the fix with a regression case where feasible.

Does a screenshot prove a browser bug is fixed?

It records a rendered state. Pair it with the functional or visual regression check that matches the defect and verify relevant viewport and timing conditions.

References