ScreenshotNeo

BlogEngineering

Common Challenges in Automated Website Testing and How to Solve Them

Learn why browser tests become brittle and how to improve selectors, synchronization, isolation, third-party coverage, device testing, and accessibility checks.

By the ScreenshotNeo team4 October 202610 min read

Automated website tests become brittle when they depend on implementation details, race the interface, share mutable state, or rely on services outside your control. Make them more trustworthy by checking user-visible behavior, isolating each test, controlling data and external responses, and waiting for meaningful UI conditions. Browser tests complement component, API, exploratory, and accessibility work; no framework removes flakiness by itself.

This guide shows how to diagnose common failures and apply practical fixes with Playwright. The same principles apply to other browser automation tools, including Cypress.

1. Start with the right job for a browser test

Use end-to-end browser tests to verify important user journeys through the running application: for example, that a user can sign in, find an item, and complete a purchase. They exercise the application as a user experiences it, but are relatively dependent on the browser, application state, and network.

  • Use component tests for isolated UI behavior and rendering variations.
  • Use API tests for service rules, validation, and data contracts that do not require a browser.
  • Use browser tests for a small set of high-value journeys and interactions across application boundaries.
  • Use exploratory testing to investigate unexpected behavior and interactions that scripted paths may miss.
  • Use accessibility checks and manual assessment for accessibility; an automated scan alone cannot establish conformance or a usable assistive-technology experience.

Keep browser coverage focused on behavior users depend on. A test that passes through every internal implementation detail can be costly to maintain and may not tell you whether the experience still works.

2. Make selectors resilient to UI changes

Tests often break after harmless refactoring because they target generated classes, deeply nested DOM structure, or positional selectors such as .panel > div:nth-child(2) button. Those details may change while the user-facing behavior stays the same.

Prefer locators that express what a user can identify: a role and accessible name, a label, or visible text. Use a test ID when you need an explicit automation contract that has no suitable user-facing locator. A test ID can stabilize a selector, but it does not demonstrate that the element is accessible.

import { test, expect } from '@playwright/test';

test('user can save profile changes', async ({ page }) => {
  await page.goto('/settings/profile');

  await page.getByLabel('Display name').fill('Avery Chen');
  await page.getByRole('button', { name: 'Save changes' }).click();

  await expect(page.getByRole('status')).toHaveText('Profile saved');
});

This test identifies the input by its label, the action by its button name, and the result by a status message. If the application has no meaningful label or status, that may be a product accessibility or usability issue worth fixing rather than bypassing with a brittle selector.

When a test ID is appropriate

Use a stable test ID for controls that cannot be reliably identified by role, label, or text, such as a complex visualization or a repeated item with no distinct accessible name. Treat the ID as a deliberate contract: document it where useful and change it intentionally when the tested contract changes. Avoid turning every element into a test ID; user-oriented locators also check that the interface presents useful names and labels.

3. Replace timing guesses with condition-based waits

A fixed sleep such as waitForTimeout(3000) guesses how long the page needs. It wastes time when the page is fast and can still fail when the page is slower. Prefer waiting for the action or state that matters.

Playwright locators wait for actionability before actions, and web-first assertions retry while waiting for the expected state. These mechanisms improve synchronization, but cannot fix an application defect or uncontrolled test data.

import { test, expect } from '@playwright/test';

test('search results appear after a query', async ({ page }) => {
  await page.goto('/search');
  await page.getByRole('searchbox', { name: 'Search products' }).fill('camera');
  await page.getByRole('button', { name: 'Search' }).click();

  await expect(page.getByRole('heading', { name: 'Search results' })).toBeVisible();
  await expect(page.getByRole('list', { name: 'Products' })).toContainText('camera');
});

Wait for the meaningful outcome, not a generic signal that may not match your application. For example, a page can reach a network-idle state while a delayed UI update is still pending; assert the rendered state the user needs.

Common synchronization choices

  • After a click or fill: assert the resulting UI state with a retrying assertion.
  • For navigation: wait for the destination URL or a stable element on the destination page.
  • For a loading indicator: assert that it disappears and that the expected content appears.
  • For a specific response: wait for the relevant request or response only when that network event is part of the behavior being verified.
  • For a genuine time-based feature: control or advance time using an appropriate test mechanism instead of sleeping for an arbitrary interval.

4. Isolate tests and control their data

A test that passes alone but fails in a suite may depend on cookies, local or session storage, database rows, or setup performed by another test. Parallel execution makes shared state problems easier to expose. A failure can then cascade: one test changes state that a later test assumes is still pristine.

Make each test establish its own prerequisites and use data it controls. Playwright recommends independent tests with their own storage, cookies, and data, and a stable staging environment when tests need a database.

  1. Create a known user or record through a test setup path or fixture.
  2. Give each test unique data when concurrent runs might collide.
  3. Reset or clean up mutable records in a reliable teardown path.
  4. Keep authentication setup explicit; do not assume a preceding test logged in.
  5. Run a failing test by itself and in the suite to detect order dependence.

Avoid tests that rely on production data or mutate a shared account. If setup and cleanup are expensive, use fixtures or a controlled staging environment while preserving the rule that each test can run independently.

5. Control third-party services in application tests

External websites and services can change independently, slow down, fail, show consent banners, or place overlays over your page. If an end-to-end test is checking your own checkout behavior, making it depend on a live payment provider’s page adds an unrelated failure source.

For application behavior that consumes a third-party response, return a controlled response in the test. Keep a separate, deliberate integration check when the external integration itself is the behavior under test. Test only what your team controls in the main application suite.

import { test, expect } from '@playwright/test';

test('shows shipping estimate from a controlled response', async ({ page }) => {
  await page.route('**/api/shipping-estimate', async route => {
    await route.fulfill({
      status: 200,
      contentType: 'application/json',
      body: JSON.stringify({ days: 3, price: 7.5 })
    });
  });

  await page.goto('/checkout');
  await page.getByLabel('Postal code').fill('10001');
  await page.getByRole('button', { name: 'Get estimate' }).click();

  await expect(page.getByText('3 days')).toBeVisible();
});

The route pattern and response should reflect the endpoint and contract your application actually uses. Add deliberate integration coverage separately if you need to detect a provider-side contract change.

6. Choose browser and device coverage deliberately

Run tests in the browsers and environments that matter to your users. A team’s default browser is not automatically representative. Consider supported browser engines, operating systems, viewport sizes, and the device-specific behaviors your product relies on.

Browser emulation can help check responsive layouts and some device settings, but it does not reproduce every behavior of a physical device. Cloudflare also documents that browser automation frameworks are unsupported for solving production challenges, and that emulation can differ from physical-device behavior.

  • Prioritize coverage based on your supported browsers and audience.
  • Use viewport and device emulation for layout and responsive behavior checks.
  • Use real devices when the behavior depends on device hardware or physical-browser characteristics.
  • Use a provider’s supported test keys or test mode for anti-bot challenges; do not design tests to bypass production security challenges.

7. Treat automated accessibility results as one input

Automated accessibility scans can find some common, machine-detectable problems, but they cannot prove WCAG conformance or establish that someone can use the experience with assistive technology. Some questions require human judgment about whether semantics match meaning. Third-party content can add further challenges.

Scan important UI states after interacting with the page, so the scan includes dialogs, menus, validation messages, and other dynamically revealed content. Review findings in context, perform manual assessment, and include people with relevant access needs in user testing where possible.

Cypress’s documentation reports that its Cypress Accessibility product can catch “up to 57% of issues that would appear in a manual audit.” This is a vendor-published, product-specific figure; it is not a general estimate for all automated accessibility tools.

8. Diagnose failures systematically

When a test fails, identify whether the cause is a product defect, an unstable test, shared state, a changed external dependency, or an environment mismatch. Do not automatically retry every failure until it passes; retries can hide a real reliability problem.

Symptom Likely cause Practical fix
Element not found Selector depends on CSS or DOM structure, or the UI did not reach the expected state. Use a role, label, or accessible name; assert the state that reveals the element.
Click times out Element is covered, disabled, moving, or otherwise not actionable. Inspect the page state and overlay; wait for the intended state and fix genuine UI issues.
Passes alone, fails in suite Shared cookies, storage, data, or order-dependent setup. Give the test independent setup and data; check cleanup and parallel collisions.
Intermittent timeout Fixed timing assumption, slow dependency, overloaded environment, or genuine hang. Replace sleeps with meaningful assertions; control dependencies and inspect traces or logs.
Unexpected banner or overlay Consent UI or external content appeared in the run. Control the external response for application tests; test the banner itself in a dedicated scenario.
Challenge blocks automation Production anti-bot protection does not support automated solving. Use the provider’s documented test mechanism, such as test keys, in the appropriate environment.
Accessibility scan is clean but users struggle Automated checks cannot assess every semantic or interaction issue. Manually review the experience and include assistive-technology and inclusive user testing.

9. Improve debugging, runtime, and reliability

Keep the browser suite small enough to give fast feedback and meaningful enough to cover critical journeys. Move checks that do not require a browser to component or API tests. Run independent tests in parallel only after their state and data are isolated.

  • Make failures observable: retain the relevant error, screenshot, trace, or browser logs in CI when available.
  • Separate environment failures from product failures: inspect whether the browser, test runner, network, or application produced the error.
  • Use retries as a diagnostic signal: a test that passes only after retry is still evidence of instability; track and fix the underlying cause.
  • Keep external checks deliberate: test your integration contract with controlled responses in the main suite, and run live checks separately when needed.
  • Watch test data cleanup: stale records and collisions can turn reliable tests into order-dependent ones.

More browser coverage costs time and compute, and an uncontrolled live dependency can make both runtime and results less predictable. There is no universal ideal number of end-to-end tests: prioritize risk, supported environments, and the cost of a missed defect.

10. Or skip the browser setup

If your task is to capture a page for visual review, documentation, or a test artifact, ScreenshotNeo provides a screenshot API and MCP server. It can return a PNG, JPEG, WebP, or PDF from one GET request, without setting up a browser for that capture.

Install the Python dependency with pip install requests, then set YOUR_API_KEY to your ScreenshotNeo key:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

See the ScreenshotNeo API documentation for request options. Equivalent cURL and Node.js requests are:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo removes known cookie and consent banners, newsletter popups, and chat widgets before the capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Sign up for free and get 1,000 screenshots a month with no card.

FAQ

Should every user journey have an end-to-end test?

No. Use browser tests for high-value behavior that needs a real browser journey; cover lower-level rules with component or API tests where those are a better fit.

Do test IDs make tests accessible?

No. They provide a selector contract for automation. Accessible names, labels, semantics, and interaction still need to be checked.

Can a green automated accessibility scan establish WCAG conformance?

No. Automated checks identify some detectable issues. Human assessment and user testing are needed to evaluate the broader experience.

Should I automate solving a production CAPTCHA?

No. Use the service provider’s documented test mode or test keys in the intended environment.

Primary sources