How to Use Software Tests as Documentation
Turn tests into clear, runnable examples of software behavior. Learn which test types answer which questions, how to write readable tests, and where prose still matters.
Use software tests as documentation by writing them as readable, runnable examples of observable behavior. Give each test a clear name, focused setup, action, and expected outcome; choose a test level that answers a specific reader’s question; and keep the suite reliable and current. Tests document the cases they exercise, not every possible behavior, so pair them with prose for rationale, constraints, and gaps.
A reader should be able to understand the claim from the test name, then verify it by following the setup and assertion. A passing suite means its assertions passed for the exercised cases; it does not prove that every requirement or input is covered.
1. Decide what the test should explain
Start with the question a future reader needs answered. The test form should fit that question:
| Reader’s question | Useful test | What it documents | Limit |
|---|---|---|---|
| What does this rule do for these inputs? | Focused unit test | Local behavior and boundary examples | An isolated component or mock does not establish whole-system behavior. |
| What does this business process mean? | Acceptance test or BDD scenario | Domain-language examples that stakeholders can discuss | Scenarios need to stay concise and connected to executable checks. |
| What does one service expect from another? | Contract test | Agreed request, response, or message shape at a boundary | A contract alone does not prove the deployed system works end to end. |
| Can a user complete an important workflow? | A small set of UI or end-to-end tests | High-level behavior through integrated components | They tend to be slower and more exposed to environmental variables. |
Use a mix based on the system and its risks. The UK Home Office describes the test pyramid as guidance to adapt, not a universal quota; Apple’s testing guidance distinguishes fast, isolated unit tests from integration and UI tests, noting that UI tests take longer and can be affected by multiple app variables. Home Office test pyramid guidance · Apple testing guidance.
2. Write tests that read like examples
- Name the behavior or rule. Describe the condition and expected result, rather than repeating only a method or class name.
- Keep one test focused. A reader should not have to untangle several unrelated claims from one test.
- Show the relevant setup, action, and outcome. Keep fixtures proportionate; hide repetitive mechanics in helpers, but make behavior-shaping inputs visible.
- Include representative normal and edge cases. Examples should clarify meaningful boundaries, not imply exhaustive coverage.
- Use the domain’s language where readers share domain knowledge. For acceptance scenarios, terms a product or business stakeholder recognizes make examples easier to review.
- Explain the unusual part, not every line. A short comment can say why an edge case matters or why an assertion exists.
- Keep tests independent, repeatable, and easy to run. A stale or flaky test is unreliable documentation.
NHS Digital’s engineering guidance says tests should be clear enough to act as documentation, focused, independent, idempotent, and runnable from the command line. NHS Digital testing guidance.
Example: document a local rule
Here is a runnable Python example using only the standard library. Save it as test_discount.py and run python -m unittest -v. The test names say what a reader should learn; the function is deliberately small so the examples remain easy to inspect.
import unittest
def apply_discount(subtotal_cents: int, discount_percent: int) -> int:
if subtotal_cents < 0:
raise ValueError("subtotal cannot be negative")
if not 0 <= discount_percent <= 100:
raise ValueError("discount must be between 0 and 100")
return subtotal_cents * (100 - discount_percent) // 100
class ApplyDiscountTests(unittest.TestCase):
def test_twenty_percent_discount_reduces_subtotal(self):
self.assertEqual(apply_discount(1000, 20), 800)
def test_full_discount_makes_subtotal_zero(self):
self.assertEqual(apply_discount(1000, 100), 0)
def test_zero_discount_preserves_subtotal(self):
self.assertEqual(apply_discount(1000, 0), 1000)
def test_negative_subtotal_is_rejected(self):
with self.assertRaisesRegex(ValueError, "subtotal cannot be negative"):
apply_discount(-1, 20)
def test_discount_over_one_hundred_is_rejected(self):
with self.assertRaisesRegex(ValueError, "discount must be between 0 and 100"):
apply_discount(1000, 101)
if __name__ == "__main__":
unittest.main()
These tests document selected rules: the result for a representative discount, behavior at the endpoints, and invalid inputs. They do not answer questions about currency rounding policy, taxes, or whether discounts apply to shipping. Those need product decisions, more examples, or prose.
3. Use shared language for domain behavior
When the important meaning is in a business process, write examples in terms of that process. A BDD scenario can make the example discussable by people who do not work in the implementation. Cucumber describes collaboratively written executable specifications as a way to establish shared language; the scenarios should still be linked to checks that run against the system. Cucumber’s BDD overview · Cucumber introduction.
Feature: Applying a discount
Scenario: A customer uses a valid discount
Given a basket subtotal of 10 dollars
When a 20 percent discount is applied
Then the payable subtotal is 8 dollars
Scenario: A full discount is applied
Given a basket subtotal of 10 dollars
When a 100 percent discount is applied
Then the payable subtotal is 0 dollars
This plain-language format is useful only if the steps are implemented and exercised. Keep scenarios about meaningful outcomes rather than turning every low-level function call into a feature story.
4. Document service boundaries with contracts
For a service integration, capture the expectations that both sides agree on: the message shape, required fields, and relevant response behavior. Contract tests check messages against a shared contract. Pact describes this as a narrower form of integration assurance than deploying and testing the complete system end to end. Pact documentation.
State what the contract does not cover. A provider conforming to a contract does not, by itself, prove that every consumer uses it correctly or that production infrastructure, authentication, and deployment configuration work together.
5. Keep end-to-end tests for important workflows
UI and end-to-end tests can show that a critical workflow works through integrated components, which makes them valuable examples of user-visible behavior. Keep the set focused: these tests are generally slower and can be affected by more variables than isolated tests. Use them for important workflows and high-risk areas, then cover detailed input rules closer to the relevant component. The appropriate balance depends on the project; do not treat a pyramid shape or fixed percentage as a requirement.
6. Know what tests cannot document alone
- They are selective. A green suite shows that written assertions passed for exercised cases. It cannot establish untested behavior. ISO/IEC/IEEE 29119-1:2022 notes that exhaustive testing is infeasible in nearly all non-trivial situations. ISO/IEC/IEEE 29119-1:2022.
- They can preserve the wrong expectation. If a test encodes a bug or a mistaken requirement, it may pass consistently while misleading readers. Confirm intent against product requirements and domain owners.
- They may explain implementation, not user behavior. A unit test can precisely document a local rule while saying little about the integrated workflow.
- They do not explain every reason. Use prose for design rationale, operational constraints, setup instructions, security assumptions, and known omissions.
- They age with the code. When behavior changes, update or remove affected examples in the same change and review whether the test name still describes its assertion.
7. Keep the suite useful as maintained documentation
- Make the normal test command discoverable in the repository’s contributor instructions.
- Keep test data small and meaningful; name fixtures after the behavior they represent.
- Avoid assertions that depend on order, wall-clock timing, shared mutable state, or external services unless that dependency is the behavior under test.
- When mocks are used, make clear which boundary is simulated. A mock-based unit test documents the local expectation, not the real remote service.
- Review behavior examples during code review: does the test name state a useful claim, and would a new maintainer understand why the outcome follows?
- Use prose alongside tests for rationale and for requirements that have no automated check.
8. Troubleshooting: when tests make poor documentation
| Symptom | Likely cause | Fix |
|---|---|---|
The test name says only test_process. |
The name identifies code rather than the behavior readers need. | Rename it to state the condition and expected outcome. |
| A reader must inspect many helpers to understand the assertion. | Setup hides the behavior-shaping inputs or combines too many cases. | Expose the meaningful inputs and expected result; split distinct claims and keep mechanical helpers small. |
| The test passes but the documented behavior is wrong. | The test expectation drifted from product intent or captured a bug. | Check requirements and domain intent, correct the implementation or test, then update related examples. |
| The suite passes locally but fails intermittently elsewhere. | Tests depend on timing, ordering, shared state, or unstable external systems. | Remove accidental dependencies, isolate state, and make external boundaries explicit. |
| A unit test is cited as proof of an entire workflow. | The test level does not match the claim. | Add a focused integration or UI check for the important workflow; retain unit tests for local rules. |
| A contract test is presented as proof that production integration works. | The contract check is being asked to cover deployment and runtime conditions it does not exercise. | Describe the contract’s scope and add appropriate checks for critical deployment or end-to-end concerns. |
| Scenarios read like technical scripts. | Implementation details displaced domain meaning. | Rewrite scenario steps in stakeholder language and keep low-level detail in unit tests. |
| Tests duplicate every sentence in the manual. | The suite is being used to replace explanations that tests cannot provide. | Keep tests for executable behavior examples; use concise prose for rationale, constraints, and coverage gaps. |
9. Performance, reliability, and maintenance cost
Readable tests have a maintenance cost: they must be run and updated when behavior changes. Keep feedback fast by putting detailed local rules in focused unit tests, and spend slower integration and UI checks on boundaries and critical workflows. That is a strategy, not a fixed ratio; complex integrations, safety needs, system architecture, and available resources can change the right mix.
Reliability affects the documentation value directly. If a test is flaky or difficult to execute, maintainers may distrust or skip it. NHS Digital recommends repeatability and command-line execution. Prefer deterministic setup, isolated test data, and explicit external dependencies. For performance-sensitive code, add tests that assert the relevant performance expectation only when the requirement and measurement conditions are clear; avoid turning unstable timing into a misleading promise.
Cost also includes reader time. A few precise examples are often more useful than a large set of redundant assertions. Choose cases that demonstrate rules and important boundaries, and leave exhaustive combinations to suitable techniques such as property-based or generated testing where the project uses them. Do not imply that a representative example covers every possible input.
Or skip the browser setup
If your tests or docs need screenshots of rendered pages, ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API returns an image or PDF; the browser-based examples above remain useful for documenting application behavior, while this handles page capture.
See the ScreenshotNeo API documentation for parameters and configuration. cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
FAQ
Can automated tests replace documentation?
No. They are maintained, executable examples of selected behavior. Keep prose for rationale, constraints, setup, and anything the tests do not cover.
Should test names describe implementation or behavior?
Prefer behavior. A reader should learn the condition and expected outcome without needing to know the internal method name.
Does a passing suite prove the system meets every requirement?
No. It means the assertions that ran passed. Coverage depends on the selected cases, and exhaustive testing is infeasible for nearly all non-trivial systems.
Should every test be written in BDD style?
No. BDD is useful when shared domain language helps people discuss behavior. Focused unit tests are usually clearer for local implementation rules.


