Model-Based Testing: What It Is and How It Works
Model-based testing derives tests from a model of expected system behavior. Learn the workflow, model choices, test selection, trade-offs, and practical examples.
Model-based testing (MBT) uses a model of a system or its expected behavior to design, select, or generate tests. The model describes relevant states, actions, rules, inputs, and expected responses. A tool can explore that model to produce test sequences and, where supported, an oracle: checks that compare the system under test (SUT)’s observed behavior with the model’s expected behavior.
In practice, MBT is a workflow rather than a particular diagram, programming language, or product. You define what behavior matters, model it, choose which behaviors to exercise, generate or derive testware, connect it to the SUT, run it, inspect results, and update the model as requirements change. The approach can help with complex, stateful behavior, but generated tests are only as useful as the model and selection criteria behind them.
1. What model-based testing means
A model is a testable representation of behavior. Depending on the goal and tool, it might be a state machine, a set of behavioral rules, a transition system, or another representation. It need not reproduce the implementation’s internals. It should capture the behavior and distinctions that tests need to check.
For a login flow, a small model might include these states:
- Signed out: a valid username and password may lead to Signed in; invalid credentials leave the user Signed out and show an error.
- Signed in: signing out returns to Signed out.
- Locked: after a defined number of failed attempts, login is rejected until the specified recovery action.
A generated sequence could try invalid credentials, then valid credentials, then sign out. Its oracle could check the returned status, visible error, and resulting authentication state at each step. The system’s actual interface may be a web UI, API, device, or service; adapters connect abstract model actions to those concrete operations.
MBT does not mean every testing activity is automated. Standards guidance describes automated testware generation and assumes test execution is automated, but teams still need to decide what to model, review generated tests, investigate failures, and maintain the model.
2. The workflow from model to results
- Set the objective. Identify the requirements, risks, interfaces, and behaviors the tests should address. Resolve ambiguous or conflicting requirements where possible; encoding ambiguity in a model does not resolve it.
- Build a testable model. Represent relevant states, actions, inputs, constraints, and expected responses. Keep the model focused on observable behavior needed for the objective.
- Choose test-selection criteria. A model may allow many paths. Decide which model elements or paths matter, such as reaching each state, exercising each transition, or covering selected combinations. The tool and model determine which criteria are available.
- Generate or derive testware. Produce abstract sequences, executable tests, or tests generated during execution. Generated testware may include actions and expected-result checks, or may need an oracle and data supplied separately.
- Adapt it to the SUT. Map model actions to API calls, UI operations, device commands, or other test harness actions. Configure test data, environment, reset behavior, and any required adapters.
- Execute and inspect. Run tests offline from a saved test repository or generate and execute them on the fly. Compare observed behavior with expected behavior, inspect failures and coverage, and determine whether a failure indicates a product defect, environment issue, or model problem.
- Revise as the system changes. Update the model and regenerate or adapt the tests when requirements, interfaces, or implementation behavior change. Review the model and selected suite again; old generated tests do not automatically stay relevant.
The exact model language, generation algorithm, selection technique, and integration depend on the tool and approach. ISO/IEC/IEEE 29119-8 describes process guidance; it does not prescribe the generation algorithm or select a tool for a team.
3. A small runnable example
This Python example models a simple lockout rule and walks a selected path. It is deliberately self-contained: the model’s step function supplies both the expected transition and a tiny simulated SUT. In a real project, replace the simulated system with calls to the actual API or UI, while retaining an independently specified model as the oracle.
from dataclasses import dataclass
@dataclass
class Model:
state: str = "signed_out"
failed_attempts: int = 0
lock_after: int = 3
def expected(self, action):
if self.state == "locked":
return "locked"
if action == "valid_login":
return "signed_in"
if action == "invalid_login":
return "locked" if self.failed_attempts + 1 >= self.lock_after else "signed_out"
if action == "sign_out" and self.state == "signed_in":
return "signed_out"
return self.state
def apply(self, action):
next_state = self.expected(action)
if action == "invalid_login" and self.state == "signed_out":
self.failed_attempts += 1
if action == "valid_login" and self.state == "signed_out":
self.failed_attempts = 0
self.state = next_state
return next_state
class FakeSystem:
"""Stand-in for the real system under test."""
def __init__(self):
self.state = "signed_out"
self.failed_attempts = 0
def perform(self, action):
if self.state == "locked":
return self.state
if action == "invalid_login":
self.failed_attempts += 1
if self.failed_attempts >= 3:
self.state = "locked"
elif action == "valid_login" and self.state == "signed_out":
self.failed_attempts = 0
self.state = "signed_in"
elif action == "sign_out" and self.state == "signed_in":
self.state = "signed_out"
return self.state
model = Model()
sut = FakeSystem()
sequence = ["invalid_login", "invalid_login", "invalid_login", "valid_login"]
for action in sequence:
expected = model.apply(action)
observed = sut.perform(action)
print(f"{action}: expected={expected}, observed={observed}")
assert observed == expected, f"Mismatch after {action}"
Save as mbt_demo.py and run with python mbt_demo.py. The final valid login remains locked because the chosen sequence has reached the lockout state. The example demonstrates a model and oracle check; it does not generate paths automatically or demonstrate a production test framework. In a real model, avoid deriving expected results from the same implementation logic under test, since that can make the test repeat the defect.
4. Models, paths, and coverage
Model choice should follow the behavior and test objective. A state machine is a natural fit when actions cause transitions among meaningful states. Rule-based models can be useful when decisions depend on combinations of conditions. Different tools support different representations, so evaluate whether the model can express the constraints and outcomes you need and whether reviewers can understand it.
Test selection controls which portion of the model becomes a practical suite. Common objectives include reaching modeled states, exercising transitions, or covering chosen paths and input classes. These are examples of goals, not guarantees that a particular tool supports every criterion. Longer paths and more combinations can increase execution cost quickly, particularly when state and data interact.
Report coverage against the chosen model and criterion, and say exactly what was measured. Model coverage is evidence that the generated suite exercised parts of the model; it is not proof that the model contains every requirement, that the model is correct, or that the product is defect-free. Review uncovered behavior and the rationale for excluding it.
5. When MBT is a good fit
MBT is worth evaluating when behavior has meaningful state, many interacting conditions, or sequences that are tedious to enumerate and maintain by hand. It may be especially plausible for reactive or distributed systems, asynchronous or nondeterministic interactions, and methods with complex parameters. A large or effectively unbounded state space with multiple possible ways to cover requirements can also signal that model-based selection may help.
Formalizing behavior can expose requirements that are vague, contradictory, or incomplete. Reusing a model to regenerate tests may make changes easier to manage than updating many hand-written cases individually. These are fit heuristics, not promises of reduced effort or better quality in every project.
Be cautious when the system is small and stable, when the behavior is simple, or when there is no capacity to maintain the model and its SUT adapters. The initial modeling effort, learning curve, integration work, process changes, and continued maintenance are real costs. MBT complements other testing methods; it should not be applied blindly.
6. Tool and approach evaluation
The sources do not establish a current product ranking for MBT tools. Evaluate candidates against the work your team must do:
- Model language and expressiveness: can it represent the behavior, constraints, and outcomes you need?
- Selection and coverage: what test-selection criteria and coverage measures are supported, and can you bound the suite?
- Generated tests and oracle: are sequences readable and reviewable? How are expected results represented and checked?
- Generation and execution: are tests generated offline, on the fly, or both, and how does execution fit your environment?
- SUT integration: what adapters, data setup, cleanup, and existing test framework support will be needed?
- Model maintenance: can the team review changes and keep the model aligned with evolving requirements and behavior?
- Adoption effort: what training, process changes, and deployment work are required?
ISO/IEC/IEEE 29119-8 provides requirements and guidance for applying MBT within the ISO/IEC/IEEE 29119-2 test process. Its scope covers definitions and links to test documentation and applies across development lifecycle models. The ISO listing accessed for this article described the edition as in the final publication process / under publication; check the listing for its current status before relying on that publication-stage detail.
7. Learning and practice resources
Microsoft’s Model Based Testing — An Introduction to Model-Based Testing and Spec Explorer explains a behavioral approach in which a model captures requirements and expected behavior, and tools generate test sequences and an oracle. It reports a specific historical Blueline protocol-compliance project: around 50 person-years saved, about 40% of effort versus a traditional approach, within work described as hundreds of protocols and approximately 250 person-years of testing. Those figures apply to that project and are not a general estimate for MBT.
ETSI’s Model-Based Testing overview describes MBT use in ICT, IT, embedded, and medical systems. It also reports historical 2012 work: four commercial tools across three case studies generated twelve models with tests for standards-related IMS and ITS work. This is dated case-study evidence, not a current tool comparison. ETSI also identifies a guide covering model creation, test generation and selection, and review of models and generated tests.
The ISTQB Certified Tester Model-Based Tester (CT-MBT) page describes an advanced MBT certification for testers, analysts, managers, developers, and architects. It lists the Certified Tester Foundation Level certificate as a prerequisite and covers MBT activities, artifacts, modeling, model languages, selection criteria, implementation, execution, adaptation, and deployment evaluation. The published exam structure is 40 questions, 26 required to pass, and 60 minutes, with 25% extra time for non-native-language candidates; verify current exam and provider details with ISTQB before planning certification.
8. Or skip the browser setup
For a web test that needs a screenshot of a page or state, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It can capture a page as PNG, JPEG, WebP, or PDF with one GET request. For example, capture a test report or a browser-visible state for review:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, no card required.
9. Troubleshooting model-based tests
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Generated tests miss an important behavior | The behavior or constraint is absent from the model, or the selection criterion does not reach it. | Trace the requirement to model states and transitions, then review the selected paths and exclusions. |
| Many tests fail at the same step | A model expectation is wrong, the SUT adapter maps an action incorrectly, or the environment is not in the assumed state. | Inspect the first divergence, verify setup and reset behavior, and compare the requirement, model transition, adapter action, and observed result. |
| Tests pass despite a known defect | The model omits the relevant requirement, the oracle is weak, or expected results mirror the implementation. | Strengthen independent expected-result checks and add the missing behavior to the model and selection objectives. |
| The suite is too large or slow | The model permits many paths or input combinations, or setup and execution are expensive. | Use explicit selection criteria and risk-based bounds, reduce redundant sequences where justified, and measure setup, action, and cleanup time separately. |
| Results vary between runs | The behavior is asynchronous or nondeterministic, timing assumptions are brittle, or test data/state is shared. | Model allowed outcomes and synchronization conditions, isolate test data, and avoid treating an allowed outcome as a failure. |
| Model and implementation drift apart | Requirements or behavior changed without corresponding model updates. | Include model review in change work, regenerate affected tests, and verify the model against current requirements. |
10. Performance, reliability, and cost
MBT adds work before the first useful run: modeling, learning the tool, building adapters, and setting up dependable test data and environments. Later, model reuse and regeneration may reduce the burden of maintaining many manually enumerated sequences, but that payoff depends on model quality and how often behavior changes.
Execution cost depends on the number and length of selected paths, the SUT, environment setup, and whether tests can run in parallel safely. A broad model can generate more tests than a team can afford to run on every change. Bound selection deliberately, preserve representative regression suites, and schedule broader exploration where it fits the delivery process.
For reliability, make state initialization and cleanup explicit, isolate data where possible, and distinguish product failures from environment and adapter failures. For asynchronous and nondeterministic behavior, model the permitted outcomes and synchronization rules instead of relying on arbitrary sleeps or a single expected ordering. Keep a record of the model version, selection criterion, and generated test version associated with results so failures can be reproduced.
Do not use raw generated-test counts or model coverage as a quality score. The meaningful question is whether the model represents the required behavior and whether the chosen suite exercises the risks that matter.
11. Frequently asked questions
Does model-based testing replace manual testing?
No. It can automate generation and execution for modeled behavior, while exploratory testing, reviews, and other test approaches may cover risks that are difficult to model.
Does MBT require a formal specification?
It requires a model suitable for the objective, but not one universal formal language. The representation and rigor vary by tool and project.
Are generated tests automatically correct?
No. They reflect the model, selection criteria, generator, and SUT integration. Review the model and oracle, and investigate failures rather than assuming every generated result is meaningful.
Is MBT only for large systems?
No, though the extra modeling and integration cost is easier to justify when state, interactions, or change make manually maintained sequences difficult.


