Common Automation Testing Mistakes and How to Avoid Them
Flaky, slow test suites often point to choices about test scope, data, timing, and diagnostics. Learn how to make automation more reliable and useful.
If your automated tests are flaky, slow, or break whenever the interface changes, the cause is often the test strategy: too much coverage at the UI layer, uncontrolled data, weak synchronization, or failures without enough diagnostic evidence. Choose the test level that matches the risk, keep end-to-end (E2E) coverage focused on essential journeys, isolate state, and investigate recurring failures instead of accepting them as background noise.
Automation works best as one part of a testing strategy. It protects repeatable behavior and catches regressions, while exploratory testing can expose usability issues, surprising edge cases, and design problems that scripted checks may miss.
1. Automating too much through the UI
A browser-driven test can exercise a complete user journey, but it also depends on more layers: the browser, application rendering, network, test data, and any services involved. That adds runtime and more potential failure points. A small interface change can also force updates to many tests, even when the underlying behavior still works.
Use different test levels for different questions:
| Test level | Good fit | Typical trade-off |
|---|---|---|
| Unit | Focused logic, boundary conditions, and calculations | Fast feedback, but limited evidence about integration with real components |
| Service or API | Request handling, business flows, and service contracts | Broad behavior coverage without exercising the full interface |
| Integration | Interactions between components or with a dependency | More realistic, but can require careful dependency and state management |
| UI or E2E | A small number of high-value customer journeys across the running system | High fidelity, with greater runtime and maintenance cost |
The test pyramid is a useful design heuristic: have more focused lower-level checks than broad-stack GUI tests. It is not a required percentage. Teams differ in architecture and risk, and test-level definitions vary. Selenium’s guidance makes the same practical point: “No one approach works for all situations.” Adapt the mix to your system.
2. Treating a test pyramid ratio as a target
A fixed ratio can distract from the reason for having multiple test levels. A 2015 Google article offered 70/20/10 as a first guess, but the appropriate distribution depends on the system and team. Ask instead:
- Can this behavior be checked reliably and more quickly below the UI layer?
- Does this test verify an important integration or customer journey?
- Will a failure identify a likely component, or only say that the whole journey failed?
- How costly is a missed defect here compared with the cost of maintaining the test?
Martin Fowler describes the pyramid as a rule of thumb and explains the trade-offs between test levels in his Practical Test Pyramid. Use the model to spot an unbalanced suite, not to satisfy a numeric quota.
3. Letting flaky failures accumulate
A flaky test changes between pass and fail without a relevant code change. John Micco’s 2016 account reported that about 1.5% of Google test results were flaky in the context of that organization and period. This is historical, organization-specific data—not a current industry rate. The underlying lesson is that inconsistent results erode trust in the suite.
Retries can help distinguish a transient failure from a repeatable one, but they add delay and can make a real race or defect less visible. Quarantine can keep a test from blocking a release while it is investigated, but quarantined coverage is easy to forget. Micco discusses these trade-offs in Flaky Tests at Google and How We Mitigate Them.
- Record the test, environment, commit, and failure details for each retry.
- Track repeat failures and assign an owner and a follow-up date.
- Use quarantine only with visible reporting, so the missing signal is clear.
- Investigate causes such as timing assumptions, shared state, unstable dependencies, or environment differences.
- Remove temporary retry or quarantine settings once the cause is fixed.
4. Using fixed sleeps or asserting before the page is ready
A fixed delay assumes the application will reach the needed state within a chosen number of seconds. That can waste time on fast runs and still fail on slow ones. Instead, wait for the condition that makes the next action valid: a page state, element, or response relevant to the scenario.
Keep waits bounded, and make a timeout report what condition was expected. Avoid waiting for broad signals, such as all network activity to stop, when the application intentionally keeps connections open or polls continuously. Google’s guidance on good E2E tests covers waiting practices and focused scenarios in What Makes a Good End-to-End Test?
5. Testing details that change often
A test that asserts exact copy, CSS classes, or internal DOM structure may fail after a harmless redesign or refactor. Prefer stable signals tied to the scenario’s purpose: the user can complete the action, the important outcome appears, or the relevant contract holds.
Presentation details are appropriate assertions when presentation itself is the requirement. For example, a visual regression check should constrain the viewport and compare the intended region, rather than accidentally treating every unrelated page change as a defect. Keep behavioral assertions and visual assertions focused on different questions.
6. Sharing mutable state and persistent test data
Tests can contaminate one another when they reuse accounts, records, files, or external resources. A test may pass alone and fail in a suite because an earlier run changed the state it expects.
- Create data specifically for a run or test where practical.
- Use unique identifiers for records created in parallel.
- Clean up owned data even when a test fails, or use disposable environments.
- Avoid relying on a shared account whose state can be changed by another run.
- Keep fakes and stubs aligned with real dependency behavior; stale doubles can make tests pass while production behavior differs.
When isolation is impossible, document the shared resource and serialize only the tests that truly depend on it. Do not use serialization to conceal an unexplained race.
7. Making failures hard to diagnose
A test result that says only “journey failed” forces someone to reproduce the issue before they can even begin to understand it. Preserve evidence that helps distinguish an application defect from a test or environment problem.
- Include the failed assertion, expected condition, and actual result.
- Keep useful application and test-runner logs, with timestamps where timing matters.
- Capture a screenshot for browser failures when it adds context.
- Preserve relevant test data or a database snapshot when safe and practical.
- Record the browser, viewport, environment, and build or commit identifier.
Be mindful of credentials and personal data in artifacts. Redact secrets and limit retention to what your debugging process needs. Documentation of known failures helps with handoffs, but it does not replace fixing recurring instability.
8. Treating automation as the whole testing strategy
Automated checks are repeatable and useful for regression protection, but they cannot answer every question about usability, clarity, or unexpected behavior. Reserve time for exploratory testing: follow real workflows, vary inputs, and investigate behavior that scripted expectations did not anticipate. Turn discoveries into regression tests when the new check can reliably protect against recurrence.
9. Building a more dependable suite
- Start from risk. List the customer and system behaviors where a defect would matter most.
- Choose the cheapest reliable test level. Put focused logic in unit tests, component interactions in service or integration tests, and only essential cross-system journeys in UI tests.
- Make state explicit. Create isolated data and identify dependencies that cannot be controlled.
- Synchronize on conditions. Replace arbitrary sleeps with bounded waits for relevant state.
- Make failures informative. Preserve the logs, screenshots, and run context needed to diagnose failures.
- Review the failure signal. Track flakes, retries, and quarantined tests; assign follow-up rather than normalizing them.
- Keep human testing in the loop. Use exploratory sessions to find gaps that automation does not describe.
10. Troubleshooting common automation failures
| Symptom | Likely cause | Practical fix |
|---|---|---|
| Passes locally, fails in CI | Environment differences, timing, resource pressure, or hidden local state | Capture environment and build details, remove dependence on local data, and wait for explicit conditions. |
| Fails only when the suite runs in parallel | Shared accounts, reused records, common files, or global mutable state | Give each run isolated data and unique resources; serialize only unavoidable shared dependencies. |
| Browser test times out intermittently | Fixed sleeps, wrong readiness signal, or a dependency that is slower than expected | Wait for the specific expected state, set a bounded timeout, and include the unmet condition in the failure. |
| Many tests fail after a UI refactor | Assertions depend on DOM structure, styles, or copy that was not part of the behavior contract | Assert user-visible outcomes or stable contracts; keep visual checks targeted to visual requirements. |
| Retry passes but first attempt fails | Possible timing race, resource contention, or unstable dependency | Keep the retry evidence and investigate the first failure; do not treat the eventual pass as a fix. |
| Failure cannot be reproduced | Missing run context, transient data, or an uncontrolled external dependency | Preserve timestamps, environment, logs, screenshots, and relevant state; reduce uncontrolled dependencies where possible. |
| Tests fail after another test changed data | Persistent shared data or incomplete cleanup | Create ephemeral fixtures and ensure cleanup runs on both success and failure. |
11. Performance, reliability, and cost
Suite cost includes more than runner minutes. Slow feedback can delay development; flaky results consume investigation time; brittle tests require maintenance after routine changes. A layered portfolio can shorten feedback by catching focused problems lower in the stack, while a small set of E2E tests checks important integrated behavior.
- Measure useful signals: duration by test level, repeat failure rate, time to diagnose, and how often failures reveal product defects.
- Parallelize carefully: parallel execution reduces elapsed time only when data and dependencies are isolated. Otherwise it can create new races.
- Keep retries visible: report initial failures as well as final outcomes so retries do not hide instability.
- Control artifact volume: retain detailed evidence for failures and enough metadata for trends, while avoiding unnecessary sensitive or oversized artifacts.
- Review maintenance cost: remove redundant checks and move assertions to a lower, more stable test level when they do not require a full browser journey.
There is no universal test count or percentage that guarantees a dependable suite. Optimize for fast, trustworthy feedback on the risks that matter in your product.
12. Capturing screenshots for browser-test diagnosis
A screenshot can show what the browser displayed at the moment of failure, which is useful when a test times out, a layout changes, or a page renders unexpectedly. Keep capture tied to a useful point in the test, such as a failed assertion or a named checkpoint, and pair it with logs and run metadata. A screenshot alone will not reveal hidden application state or explain why an action failed.
For a page you can access directly, a local browser automation setup can capture a screenshot. For example, using Playwright with Node.js:
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.locator('body').waitFor({ state: 'visible' });
await page.screenshot({ path: 'failure-context.png', fullPage: true });
} finally {
await browser.close();
}
Install Playwright and its browser binaries according to the official Playwright documentation. Replace the example URL with a page you are authorized to access. In an actual test, capture from the existing page and test context so the screenshot reflects the same state as the failure.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns an image or PDF; see the API documentation for options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
For Node.js environments without Bun, read the response as an array buffer and write it with your runtime’s filesystem API. These examples show the core request; use the docs for output and capture options. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo for product details and the docs for configuration. Sign up free for 1,000 screenshots a month, no card required.
FAQ
Should every flaky test be quarantined?
No. Quarantine can remove a known unstable test from a critical path while it is investigated, but keep its status visible and assign follow-up so it does not become permanently ignored.
Are UI tests a bad idea?
No. They are valuable for a small set of important journeys that need to be verified across the running system. The problem is relying on them for every behavior.
How many end-to-end tests should a team have?
There is no universal count. Select tests based on customer impact, system risk, and whether lower-level checks can answer the same question more reliably.
Can a screenshot prove why a test failed?
It can show visible page state, but it does not capture all relevant causes. Pair it with logs, test data, and environment details.


