7 Ways to Clean Up and Improve Your Test Code
Improve test readability without weakening coverage. Learn seven practical refactoring techniques, safer assertions, and ways to keep tests repeatable.
Clean test code makes failures easier to understand and changes safer to make. Improve it by naming expected behavior, focusing each case, removing only confusing duplication, clarifying setup, writing legible assertions, controlling state and dependencies, and refactoring in small verified steps. These changes should preserve the behaviors the suite checks.
Google Testing Blog describes tests as readable documentation and recommends describing behavior through public APIs. It also raises a useful question when refactoring: “How do you know that your refactoring of the tests was safe and you didn’t accidentally remove one of the assertions?” Google Testing Blog: What Makes a Good Test? Google’s guidance on refactoring tests
1. Name the behavior the test promises
A test name should tell a reader what a user or caller can observe. Names tied to private methods or incidental implementation details make harmless refactors appear to change behavior and can hide what the test protects.
# Vague
def test_process():
...
# Describes an observable outcome
def test_expired_session_redirects_to_sign_in():
...
Use the naming style your test framework and codebase already use. A useful name usually identifies the relevant condition and outcome. Avoid encoding every setup detail in the name; keep those in the test body.
2. Keep each test focused on one scenario
A focused test has one clear reason to pass or fail. When a test exercises several unrelated scenarios, the first failure can prevent later checks from running, and a failure message may not reveal which behavior is broken. Home Office developer-testing guidance similarly emphasizes clear intent and one test case.
def test_discount_is_applied_to_eligible_order():
order = Order(subtotal=100, customer_is_member=True)
total = calculate_total(order)
assert total == 90
def test_discount_is_not_applied_to_guest_order():
order = Order(subtotal=100, customer_is_member=False)
total = calculate_total(order)
assert total == 100
Parameterized tests can keep closely related input/output examples together when each row represents the same rule. Split cases when they exercise materially different behavior or need different explanations. Do not split mechanically if doing so makes a simple rule harder to see.
3. Remove duplication only when it improves clarity
Repeated construction, setup, and assertions accumulate as suites grow. A small helper can make intent clearer when it removes irrelevant mechanics. But an abstraction that hides important inputs or outcomes makes a test harder to review.
def make_member_order(subtotal=100):
return Order(subtotal=subtotal, customer_is_member=True)
def test_member_gets_discount():
order = make_member_order()
assert calculate_total(order) == 90
Keep the helper close to the tests that use it when it is local to one suite. Name helpers for the scenario they create, and allow meaningful values to be passed in rather than burying them in a large fixture. HMRC’s test-automation guidance recommends managing test-pack size and reducing duplication across testing levels; that does not mean every repeated line needs an abstraction.
4. Make setup and fixtures understandable
Setup is part of the explanation of a test. Prefer the smallest fixture and dataset that show why the expected result follows. Critical conditions should be visible in the test or in a clearly named helper, rather than hidden in distant global setup.
def test_archived_record_is_omitted_from_active_results(repository):
repository.add(Record(id="active", archived=False))
repository.add(Record(id="old", archived=True))
results = repository.list_active()
assert [record.id for record in results] == ["active"]
Use shared fixtures for stable, genuinely common mechanics such as a temporary directory or a test client. Use case-specific data for the facts that make a scenario meaningful. If changing a fixture has surprising effects across many tests, narrow its scope or make the dependency explicit.
5. Make assertions communicate the contract
Assert the observable behavior that matters, and make failures easy to interpret. Prefer checking a returned value, state transition, or public interaction over details that merely mirror the implementation. Avoid assertions so strict that harmless variation causes failures, but do not weaken them until meaningful regressions can pass.
# Less informative
assert response
# Makes the expected behavior explicit
assert response.status_code == 201
assert response.json()["state"] == "created"
For collections, assert the relevant contents and ordering only when ordering is part of the contract. For floating-point values, use an appropriate tolerance. For time-dependent behavior, control the clock where possible instead of asserting a timing window narrower than the behavior requires. pytest’s guidance on flaky tests discusses overly strict assertions as one source of instability.
When simplifying a test, compare the old and new assertions. Ask what regression each assertion would catch and whether the revised test still catches it. Google’s test-refactoring article describes a deliberate technique: make the implementation under test wrong, verify the expected assertions fail as the test is restructured, restore the implementation, and confirm the tests pass. Use this carefully on code where deliberately changing behavior is safe; it is a focused technique, not a requirement for every edit.
6. Control state and external dependencies
Repeatable tests should not depend on execution order, leftover records, current environment, or live third-party services. The Home Office advises against environment-varying test values and recommends avoiding external dependencies such as third-party APIs in unit tests. pytest identifies uncontrolled state, ordering, missing cleanup, and strict assertions among contributors to flaky tests.
- Use temporary databases, directories, or containers where appropriate, and clean them up reliably.
- Reset shared state between tests; do not rely on a previous test’s setup.
- Freeze or inject clocks, random sources, and environment configuration when they affect outcomes.
- Stub network boundaries in unit tests. Cover real integrations in a smaller, deliberate integration suite.
- Make parallel execution safe by giving tests independent resources and unique data.
Mock only the boundary needed to control nondeterminism or isolate a unit. If a test asserts the calls to many internal collaborators, it may be coupled to implementation changes rather than behavior.
7. Refactor in small steps and preserve the signal
Change one aspect at a time: rename a test, extract a helper, or make an assertion more readable, then run the relevant tests. Keep the suite passing while refactoring production code. For a test-only refactor, verify that the old and new versions protect the same behavior, not just that the current implementation passes both.
- Record the behavior and assertions the test currently protects.
- Make one structural change.
- Run the narrow test or suite and inspect failures.
- Where safe and useful, introduce a temporary wrong result to confirm the key assertion detects it; restore the implementation immediately.
- Run the relevant suite again, then the broader suite if shared fixtures or helpers changed.
Keep temporary mutations local and ensure they are not committed. If a refactor removes an assertion, replace it only when the behavior remains checked elsewhere in a clear, intentional way.
Choosing the right test level
Unit, integration, and UI-driven tests provide different kinds of confidence and have different execution and maintenance costs. HMRC recommends preferring faster unit tests where they provide the needed confidence, reducing duplication across test levels, and maintaining test packs to reduce flakiness. It also recognizes that the right levels depend on the software; unit tests do not universally replace integration or UI tests.
| Level | Useful for | Trade-off |
|---|---|---|
| Unit | Fast checks of isolated logic and edge cases | Requires controlled dependencies; does not establish that all external components work together |
| Integration | Contracts between components, databases, or service boundaries | More setup and execution cost; isolate external services where practical |
| UI or end-to-end | Important user journeys across the assembled system | Often more operationally sensitive; keep scenarios focused and maintain selectors and test data |
Put a behavior at the lowest level that gives adequate confidence, then retain higher-level checks for risks that only emerge across boundaries. Avoid repeating the same assertion at every layer without a distinct reason.
Common cleanup problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Many unrelated failures after extracting a fixture | The fixture is shared too broadly or carries mutable state | Narrow its scope, create fresh objects per test, and make dependencies explicit. |
| A test passes alone but fails in the suite | Order dependence, leaked state, or missing cleanup | Run it in different orders, isolate resources, and add reliable teardown. |
| A test is flaky around time or numeric values | Timing assumptions, real clocks, or exact floating-point equality | Control time and use tolerances that reflect the contract. |
| A refactor leaves fewer meaningful checks | Assertions were removed while consolidating code | Map each removed assertion to a retained behavior check; use a safe temporary mutation if appropriate. |
| Failures point into a generic helper | The helper hides scenario-specific expectations | Move the assertion back to the test or pass the expectation explicitly. |
| Tests are slow after adding more coverage | Expensive setup is repeated or the same behavior is checked at several levels | Reuse safe setup, move fast checks to lower levels, and keep boundary tests for distinct confidence. |
Performance, reliability, and maintenance
Test cleanup should reduce the time needed to understand and change the suite. A helper that saves lines but forces readers to jump through layers can raise maintenance cost. A test that depends on a live service can add latency and intermittent failures; isolate that dependency for unit checks and reserve live integration coverage for cases where it adds confidence.
Prefer a fast, focused suite during inner development and run broader checks when changes affect shared infrastructure, fixtures, or integration boundaries. Keep a flaky test visible and investigate its cause rather than repeatedly rerunning it until it passes. A test suite is maintained code: its clarity, repeatability, and cost all affect how reliably it catches regressions.
Or skip the browser setup
If browser screenshots are part of visual regression or UI test work, ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API returns a screenshot or PDF, while handling common page cleanup before capture.
See the ScreenshotNeo API documentation for the available parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free and get 1,000 screenshots a month with no card.
Frequently asked questions
Should every test have exactly one assertion?
No. Keep assertions that together describe one scenario’s outcome. Split a test when checks represent independent behaviors or make failures hard to diagnose.
Should duplicated test setup always become a fixture?
No. Extract it when the helper makes intent easier to read and the shared setup is genuinely stable. Keep scenario-defining data visible.
How can I tell whether a test refactor preserved coverage?
Review the behavior each assertion protects, compare the resulting checks, and run the tests. For a safe, isolated case, a temporary deliberate fault can show that the expected check still fails.
Are flaky tests just a CI problem?
No. Flakiness can reveal uncontrolled state, environment dependence, order coupling, timing assumptions, or cleanup gaps that also make local results less trustworthy.


