Common Challenges in Automation Testing and How to Solve Them
Diagnose flaky, brittle, slow, and costly browser tests with practical fixes for timing, test data, parallel runs, and test scope.
Automated tests usually become flaky, brittle, slow, or expensive because they depend on uncontrolled state, timing assumptions, shared data, execution order, or too much browser-level coverage. Start by reproducing the failure and checking its state and timing, then make each test independent, wait for meaningful conditions, isolate parallel work, and keep browser tests for behavior that actually needs a browser.
Browser automation has a real cost in runtime and infrastructure. Selenium’s guidance is to ask whether a real browser is needed for the behavior under test; a lower-level check may answer some questions faster. Browser tests remain useful for critical user-visible flows where browser behavior is part of the requirement. Selenium’s test automation overview describes tests as data setup, a discrete action, and result evaluation, and recommends keeping those steps short.
1. Diagnose the failure before changing the test
First determine whether the failure is repeatable, intermittent, order-dependent, or limited to parallel or CI execution. A retry that passes is evidence of an intermittent symptom, not proof of its cause.
- Record the failing test, environment, worker or shard, and whether it passed on a retry.
- Preserve the relevant browser state: URL, console and network errors, screenshot or trace when available, and the state of records the test touched.
- Run the test alone, then with its normal suite order, then with parallel execution if that is where it fails.
- Change one suspected dependency at a time: timing, setup, shared data, order, or resource pressure.
This sequence helps separate an application race from a test dependency or an overloaded CI environment. Google’s testing guidance identifies execution-time dependence, asynchronous event-order assumptions, waits without timeouts, and races between tests and the application as common sources of flakiness. Google’s guidance on test flakiness.
2. Fix flaky waits and timing races
A fixed delay assumes the page or application will always become ready within that exact interval. If it is too short, the test races ahead; if it is too long, every successful run pays the delay. Replace arbitrary sleeps with a wait for the state the next action actually requires, and always bound the wait with a timeout.
Playwright example
import { test, expect } from '@playwright/test';
test('shows the saved confirmation', async ({ page }) => {
await page.goto('https://example.com/settings');
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByRole('status')).toHaveText('Saved', {
timeout: 10_000,
});
});
The assertion waits for a meaningful visible result. Prefer locators and assertions that describe the expected condition over waiting for a guessed number of milliseconds. If the condition does not appear before the timeout, inspect whether the action failed, the application returned an error, the locator is wrong, or the app is genuinely slow.
Timeouts are limits, not a fix for a missing condition. Choose them to reflect the expected environment and capture diagnostics when they expire. A large timeout can hide a slow or broken dependency and lengthen feedback; a tiny timeout can create false failures on loaded CI workers.
3. Make tests independent of execution order
A test should establish its own prerequisites rather than assume another test created a user, left a session active, or populated a database. Selenium warns against depending on a particular test order, and pytest notes that leftover state can make tests fail when run in parallel.
- Create the required records in the test or a deliberately scoped fixture.
- Use cleanup where appropriate, while ensuring cleanup failure does not hide the original failure.
- Use unique identifiers for records that may be created or modified concurrently.
- Do not rely on a test’s prior browser cookies, local storage, or server-side effects unless those are explicitly part of the scenario.
A useful check is to run a failing test by itself and in a different order. If its result changes, look for setup or cleanup that leaks state. Selenium’s guidance on avoiding shared state and pytest’s explanation of flaky tests discuss these dependencies.
4. Make parallel execution safe
Parallelism can shorten feedback, but it increases the chance that workers collide through shared backend records, fixed filenames, accounts, queues, or rate-limited external services. Playwright runs test files in parallel by default; its workers are separate processes, which does not isolate state outside those processes.
- Assign unique data and output paths per test or worker.
- Use worker-scoped setup only when sharing within that worker is intentional.
- Keep external services and shared test accounts in the isolation review.
- Increase worker count gradually while watching failures, service limits, and CI resource pressure.
- Use sharding when splitting work across CI jobs, and ensure each shard has isolated state.
There is no universally correct worker count. Choose it based on the available CI resources, the application’s ability to handle concurrent test traffic, and any external-service limits. See Playwright’s parallelism guide.
5. Choose the right level of test
Use a browser when the browser itself is part of what must be verified: navigation, rendering, user interaction, browser storage, or an end-to-end critical flow. Use a lower-level test when it can establish the same behavior more directly and with less setup.
| Question | Browser test is a good fit when… | Consider a lower-level check when… |
|---|---|---|
| Does browser behavior matter? | The result depends on rendering or real user interaction. | The behavior is logic or data handling that can be checked without a browser. |
| What does the test cost? | The extra runtime and browser infrastructure are justified by user-visible confidence. | A browser adds setup and runtime without improving the evidence. |
| Can failure be diagnosed? | You can retain enough state and timing evidence to reproduce it. | A simpler test gives clearer failure localization. |
A focused browser suite for critical flows can complement faster checks below it. Avoid turning every requirement into a full end-to-end flow. Selenium’s overview directly emphasizes browser necessity and infrastructure cost; the isolation and diagnosis comparisons above follow from the cited test guidance.
6. Use retries as a diagnostic signal
Retries can reveal that a failure is intermittent and can help teams keep a signal visible while investigating. They do not remove the underlying timing race, shared data, order dependency, or environmental limit. Playwright supports retries and starts a fresh worker after a failure, so a retry may also run under a different worker state.
Track which tests fail first and pass later, and investigate their common dependencies. Do not treat a green retry as evidence that the test is reliable. Playwright’s retry documentation explains retry behavior.
7. Troubleshooting common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Element lookup or click fails intermittently | The test acts before the relevant UI state is ready, or the locator is ambiguous. | Wait for the expected state with a bounded assertion; use a locator tied to the intended element and inspect the page state on failure. |
| Test passes alone but fails in the suite | Order dependence, leaked state, or cleanup behavior. | Run it in different positions, initialize its prerequisites, and remove reliance on another test’s side effects. |
| Test fails only with multiple workers | Shared records, accounts, files, or external-service limits. | Isolate data and output, then lower worker count temporarily to confirm whether concurrency is involved. |
| Test passes on retry | An intermittent race or environment-sensitive dependency may exist. | Keep the first-attempt failure and investigate timing, state, and worker context; do not count the retry as a root-cause fix. |
| CI failures but local runs pass | Different timing, resource limits, configuration, or external dependencies. | Compare environment and configuration, retain failure diagnostics, and reproduce under similar load where possible. |
| Suite is slow despite few failures | Too many browser-level checks, excessive fixed waits, or excessive concurrency contention. | Move checks that do not need a browser lower in the stack, replace sleeps with condition waits, and tune worker count deliberately. |
8. Screenshot a page for visual debugging
A browser screenshot can preserve what the user-visible page looked like when a visual or rendering issue occurred. It is one diagnostic artifact; it does not replace logs, traces, application state, or a test assertion. For a local Playwright run, add a screenshot at the relevant point:
import { test } from '@playwright/test';
test('capture page state', async ({ page }) => {
await page.goto('https://example.com');
await page.screenshot({ path: 'artifacts/page.png', fullPage: true });
});
Use unique artifact paths in parallel runs so workers do not overwrite one another. Keep artifacts for failures when your CI setup supports that, and avoid capturing sensitive page content into broadly accessible logs.
Or skip the browser setup
If you need a clean screenshot of a live page while documenting or diagnosing a browser issue, ScreenshotNeo provides a website screenshot API and MCP server for developers. The DIY Playwright flow above is appropriate when the test must exercise your own browser session and application state. For a standalone page capture, one GET request can return an image or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
- Cookie banners are accepted and removed before capture; newsletter popups and chat widgets are removed too.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers identify the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs.
- The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
9. Performance, reliability, and cost
- Runtime: Browser startup, navigation, and fixed delays add time. Use browser tests where they provide browser-specific confidence and wait on relevant conditions.
- Reliability: Deterministic setup, independent tests, isolated parallel data, and failure context make failures easier to reproduce.
- CI capacity: More workers can reduce elapsed time until resource contention or shared dependencies become the bottleneck. Tune against your environment rather than adopting a universal number.
- Maintenance cost: Keep each test focused on setup, one meaningful action, and an observable result. Prefer lower-level checks for behavior that does not require a browser.
- Screenshot capture cost: A local browser capture consumes your test infrastructure. ScreenshotNeo bills only clean shots; its plans are Free for 1,000 per month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.
10. A practical checklist
- Can this requirement be verified without a real browser?
- Does every test create or deliberately fixture its own prerequisites?
- Does each wait target a meaningful condition and have a timeout?
- Can tests run in a different order without changing results?
- Are data and output files isolated across workers and shards?
- Do first-attempt failures retain enough state and timing evidence?
- Are retries exposing intermittent failures rather than hiding them?
- Is parallelism tuned to the app, external services, and CI resources?
Frequently asked questions
Why are my automated tests flaky?
Look first for timing assumptions, asynchronous races, shared state, order dependencies, and environment differences. A retry that passes identifies an intermittent symptom but not its cause.
Should I increase every test timeout?
No. Use a timeout around a meaningful condition. Raising limits globally can slow feedback and conceal an underlying delay or failure.
When should I use browser automation instead of unit tests?
Use browser automation when real browser behavior or a critical user interaction is part of the requirement. Prefer lower-level checks when they can answer the question with less infrastructure and clearer diagnosis.
Does a separate Playwright worker guarantee isolated tests?
No. Worker processes do not automatically isolate shared backend records, accounts, external services, or output paths.


