Why Complexity Makes Test Automation Harder
Complex systems create more combinations, timing conditions, and failure causes for automated tests to handle. Learn how to control test scope without overstating coverage.
Complexity makes test automation harder because it expands the behavior a test suite must represent: more input combinations, states, dependencies, configurations, and timing conditions. Testing every possibility quickly becomes impractical. Teams must decide which combinations matter, keep checks maintainable as systems change, and determine whether a failure came from the product, the test, or its environment.
The practical response is to model important conditions, select representative values deliberately, and target interactions at a level that fits the system’s risk. Combinatorial testing can reduce the number of cases when faults are triggered by interactions among a limited number of parameters, but it does not prove every possible behavior is correct. NIST’s research on software fault complexity explains the conditional reasoning behind this approach.
1. Complexity expands the test space
A parameter is any condition that can change the behavior under test: an input value, user role, browser, configuration flag, service response, or state such as a signed-in session. Each parameter can take multiple values. The combinations multiply as parameters are added.
For example, a page might vary by three user roles, two authentication states, four viewport sizes, and three backend response states. Exhaustive coverage of just those dimensions would require 3 × 2 × 4 × 3 = 72 combinations. Real systems often have many more dimensions, constraints, and sequences of actions.
NIST’s 2004 paper puts the fundamental limit plainly: “Exhaustive testing of computer software is intractable.” The paper also reports empirical evidence that failures in studied domains were often caused by combinations of relatively few conditions. Under the explicit assumption that faults are triggered by combinations of no more than t parameters, covering all t-way combinations can be effectively exhaustive for discrete values. That assumption is not a guarantee that a particular system has no higher-order faults.
2. Modeling and choosing values takes judgment
A test generator can create cases from a model, but people still have to decide what the model represents. They must identify parameters, choose meaningful values, describe constraints, and select how many-way interactions to cover. An incomplete or unrealistic model can generate many tests while missing the conditions that matter.
A NIST case study of its ACTS test-generation tool found that modeling the input space was a significant undertaking. The study reported effective coverage and fault detection in the system it examined; it is evidence that combinatorial testing can be useful, not a universal benchmark or a promise for every project. See the ACTS case study.
Handling continuous inputs
Numbers such as distance, price, or elapsed time can have enormous or effectively continuous value ranges. Testing every value is impossible. Partition the range according to requirements and behavior, then choose representative values from each partition. Include boundary-value cases where behavior changes, such as zero, a minimum or maximum, just below a threshold, and just above it. NIST’s combinatorial testing resources discuss selecting and modeling test inputs.
Write down why the chosen values represent the relevant behavior. A generated test set is only as informative as its input model and value selection.
3. Interaction coverage is a tradeoff, not a proof
When exhaustive enumeration is infeasible, t-way testing aims to cover combinations of values across every selected group of t parameters. Pairwise testing is the special case t = 2. It can be a reasonable starting point when pair interactions are important, while higher interaction strength may be justified for risky features or known multi-condition behavior.
| Choice | What it covers | Tradeoff |
|---|---|---|
| Exhaustive combinations | Every modeled combination | Can become too large to generate, run, and maintain. |
| Pairwise (2-way) | Every pair of values across modeled parameters | Smaller than exhaustive coverage, but can miss faults that require three or more conditions. |
| Higher-strength t-way | Interactions among groups of up to the selected strength | More coverage of interactions, with a larger test set and execution cost. |
| Risk-based selected cases | Conditions chosen from domain knowledge, incidents, and requirements | Can focus effort on consequential behavior, but relies on judgment and can leave unanticipated cases out. |
Use constraints to exclude impossible combinations, and variable interaction strength when some parameters warrant more coverage than others. State the selected strength and assumptions. Do not call a pairwise or t-way suite exhaustive unless the relevant fault-interaction assumption has been justified.
4. More tests create execution and maintenance costs
As a system and its suite grow, test execution can take longer, feedback arrives later, and failures become harder to diagnose. Checks also need maintenance when interfaces, dependencies, data models, or expected behavior change. Brittle assertions and tests that depend on implementation details can turn harmless refactoring into broad test repair work.
A 2026 survey of Selenium-based automation reported average challenge ratings of 3.43 for assertability, 3.24 for asynchrony, and 3.15 for brittleness. The excerpted results do not state the rating scale, so these values should not be read as percentages or prevalence estimates. The survey describes challenges in scaling and maintaining suites, long execution, and distinguishing failure causes. See the survey in Information and Software Technology.
Keep the suite useful by making each test’s purpose clear, choosing stable assertions tied to requirements, and running tests at the appropriate scope. This dossier does not establish a universal ranking of unit, integration, API, or end-to-end frameworks; choose layers based on the behavior and risk being checked.
5. Timing and shared state make automation flaky
Asynchronous work, concurrency, test order, network behavior, clocks, randomness, and shared resources can make the same test pass on one run and fail on another without a relevant product change. This is flakiness. It makes feedback less trustworthy and can delay releases. A 2023 multivocal review identifies test-order dependency and concurrency among widely studied causes of flaky tests. See the review in the Journal of Systems and Software.
Complexity can also make a flaky failure difficult to reproduce and localize because more interacting components and environmental conditions may be involved. That is a practical inference about diagnosis, not a quantified causal finding from the review. Make runs reproducible where possible: record inputs, configuration, browser or runtime version, test order, timestamps, and relevant service responses.
6. A practical way to control the work
- Define the behavior and risk. Start from requirements, user-visible outcomes, incidents, and costly failure modes. Avoid modeling every implementation detail by default.
- List parameters and values. Include relevant inputs, states, configuration, dependencies, and environment differences. Use representative partitions for continuous values.
- Write down constraints. Exclude combinations the system cannot reach, and document assumptions so the model can be reviewed when the product changes.
- Select interaction strength. Use pairwise coverage as a deliberate choice when two-way interactions are a reasonable target. Increase strength for high-risk behavior or where evidence suggests faults depend on more conditions.
- Review the generated cases. Check that important boundaries, state transitions, and known failure scenarios appear. Generation does not replace domain review.
- Keep tests diagnosable. Make failures identify the case, values, environment, and relevant artifacts. Separate product assertions from setup and infrastructure checks.
- Investigate inconsistent results. Reruns can help expose a flaky test, but a passing rerun does not make the original failure irrelevant. Find the unstable dependency, timing condition, shared state, or environment difference.
- Revisit the model as the system changes. Add parameters when new behavior warrants them, remove obsolete dimensions, and update the selected coverage based on risk and observed failures.
7. Troubleshooting common automation problems
| Symptom | Likely cause | What to do |
|---|---|---|
| The suite is too slow to give useful feedback | Too many redundant combinations, expensive setup, or every check running at the same scope. | Review the model and constraints, use an appropriate interaction strength, and identify repeated setup or checks that do not add distinct evidence. |
| A generated suite misses an important defect | The parameter model, values, constraints, or interaction strength did not represent the triggering conditions. | Add the missing condition and a regression case; review whether similar behavior requires stronger interaction coverage. |
| The same test passes and fails across runs | Async waits, races, order dependency, shared state, time, randomness, network, or machine variation. | Capture run metadata, isolate state, control clocks and seeds where possible, and replace arbitrary sleeps with condition-based waits. |
| Many tests fail after a small UI change | Assertions or selectors are coupled to fragile presentation details. | Anchor checks to stable user-visible behavior and intentional contracts; update shared helpers carefully and review what each assertion protects. |
| A failure is difficult to classify | Product behavior, test code, assertion logic, synchronization, or infrastructure can produce similar symptoms. | Preserve logs and artifacts, reproduce with the same inputs and environment, and inspect setup, waits, and dependencies before assigning cause. |
| Continuous values create too many cases | The test plan treats a large numeric range as a list of individual values. | Partition by requirements, use representative equivalence classes, and include boundaries and values around thresholds. |
8. Performance, reliability, and cost
Automation has costs beyond the number of generated cases: modeling and review effort, execution time, environment capacity, test maintenance, and time spent investigating failures. Raising interaction strength generally increases coverage of modeled combinations and can increase the suite size. The exact size depends on parameters, values, constraints, and the generator; this research dossier does not provide a general cost formula or a universal return-on-investment figure.
Reliability depends on stable test behavior and understandable evidence. A large suite that frequently produces ambiguous failures can provide weaker practical feedback than a smaller, well-modeled suite whose failures are reproducible and actionable. Track execution duration, failure categories, rerun outcomes, and maintenance effort so coverage decisions can be revisited against real project needs.
9. Capture visual evidence for browser checks
When browser automation checks a visual state, a screenshot can help a developer inspect what the page rendered at the point of failure. Capture evidence with the same viewport, state, and timing conditions as the test, and treat the image as diagnostic context rather than proof that every behavior is correct.
For a do-it-yourself capture, use the screenshot capability of your browser automation framework and save the artifact when an assertion fails. Keep the page state and viewport consistent between runs so comparisons are meaningful.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation. For example, save a WebP screenshot with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides screenshot and PDF tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month with no card.
FAQ
Does pairwise testing find every bug?
No. It covers pairs of modeled parameter values. It can miss faults requiring three or more conditions, omitted parameters, unmodeled values, or behavior outside the model.
Should every test be automated?
No. Automate checks where repeatability and feedback justify the modeling and upkeep. Choose coverage based on risk and the evidence the check can provide.
Does a flaky test mean the product is defective?
Not by itself. The cause may be product behavior, test logic, timing, shared state, or the environment. Preserve evidence and investigate before drawing that conclusion.


