ScreenshotNeo

BlogGuides

How to Make Web Automation Reliable in Production

Build dependable browser automation with resilient locators, state-based waits, isolated runs, safe retries, and useful failure evidence.

By the ScreenshotNeo team4 October 20269 min read

Reliable web automation comes from proving the outcome of each step, synchronizing on the state the workflow needs, isolating runs, and collecting enough evidence to explain failures. Retries can help with transient problems, but a test that passes only on retry is a flakiness signal. Before retrying an action with business side effects, check whether the first attempt already succeeded.

This guide focuses on browser end-to-end tests and authorized workflows. The right reliability target depends on your application, its dependencies, and the runtime environment.

1. Define what success means

Start with the result a user or downstream system should observe. A click returning without an exception does not prove that an application accepted the change. Choose an outcome that demonstrates success, such as a confirmation message, a changed status, or a newly created record identifier.

  1. Write down the expected visible or business outcome.
  2. Identify an observable signal that proves it happened.
  3. For important writes, decide how to check whether the action committed if the browser loses the response.

This last step matters for submissions that create purchases, publish content, send messages, or otherwise cause side effects. If the result is ambiguous, inspect the current application state before repeating the action. Microsoft’s Playwright Workspaces guidance likewise says to observe the page and determine whether the expected state change occurred before retrying an action in relevant recovery cases (Playwright Workspaces remote automation guidance).

2. Use resilient locators and wait for the needed state

Prefer locators based on user-facing attributes such as roles and accessible names, or explicit test contracts. Locators tied to incidental DOM structure or styling classes tend to break when the page is reorganized. Playwright recommends user-facing locators and notes that its locators automatically wait and retry (Playwright best practices).

A Playwright click also checks that its locator resolves to one element and that the element is visible, stable, enabled, and able to receive events (Playwright actionability checks). Follow actions with an assertion about the expected result, rather than treating the action itself as proof.

import { test, expect } from '@playwright/test';

test('user can save profile changes', async ({ page }) => {
  await page.goto('https://example.com/profile');

  await page.getByRole('textbox', { name: 'Display name' }).fill('Ada Lovelace');
  await page.getByRole('button', { name: 'Save changes' }).click();

  await expect(page.getByRole('status')).toHaveText('Profile saved');
});

Replace the example URL and expected message with signals your application actually exposes. If the page has no useful accessible name or status, add a stable test contract or improve its accessible markup.

Why fixed sleeps are fragile

A fixed sleep waits even when the page is ready early, yet can still be too short on a slow run. Browser navigation reaching a document readiness state also does not guarantee that a JavaScript application has rendered the control or result your workflow needs. Selenium documents this race between navigation and later JavaScript-driven changes (Selenium waiting strategies).

Use an assertion or explicit wait for the actual condition. In Selenium, for example:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

with webdriver.Chrome() as driver:
    driver.get('https://example.com/profile')
    save = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.ACCESSIBILITY_ID, 'Save changes'))
    )
    save.click()
    WebDriverWait(driver, 10).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, '[role="status"]'))
    )

Selenium locator strategies vary by binding and application; use a locator supported by your chosen binding and stable in your page. The timeout is a limit for waiting on the condition, not a fixed delay before every action.

3. Make runs independent

Give each test a fresh browser context and independent application data wherever feasible. Playwright test pages use isolated Browser Contexts, equivalent to fresh browser profiles (Playwright writing tests). Selenium’s guidance also recommends test independence, avoiding shared state, and fresh browsers as useful practices, while emphasizing that no single approach fits every situation (Selenium encouraged behaviors).

  • Do not let one test depend on another test’s cookies, local storage, or setup order.
  • Use distinct accounts or uniquely identified test data when concurrent runs could collide.
  • Clean up created data where practical, and make cleanup safe to repeat.
  • Keep authentication setup explicit and refreshable; stale sessions often look like locator or navigation failures.

Isolation makes failures easier to reproduce and reduces the chance that a retry inherits damaged state. Some workflows necessarily share a controlled account or environment; document that dependency and serialize only the parts that cannot safely run independently.

4. Use retries as a signal, not a repair

Playwright Test does not retry by default. When retries are configured, a test that fails and then passes is reported as flaky (Playwright retries). Track retry-pass cases and investigate their cause: timing, shared state, external dependencies, authentication, or resource pressure can all be involved.

Retries are most straightforward for read-only steps or operations designed to be idempotent. For a write whose result is unclear, first look for evidence that the initial attempt committed. If it did, continue from the new state or clean it up; do not blindly replay a purchase, publish, or message action.

Set retry limits based on the disruption a transient failure causes and the cost of delayed feedback. A high retry count can hide recurring defects and lengthen CI runs. Report both the initial failure and retry result so the signal is not lost.

5. Capture diagnostics and protect them

A failure is much easier to fix when the report includes the state around it. Playwright Trace Viewer can show a test timeline, DOM snapshots, and network requests. The Playwright guide describes collecting traces on the first retry as a CI option and warns that tracing every test can add overhead (Playwright best practices).

For a Playwright project, a practical starting point in playwright.config.ts is:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: process.env.CI ? 1 : 0,
  use: {
    trace: 'on-first-retry',
    screenshot: 'only-on-failure',
    video: 'retain-on-failure'
  }
});

Choose artifact settings that fit your data policy and storage budget. Traces, screenshots, videos, page URLs, and network evidence can include credentials, personal information, or business data. Restrict access and retention to the debugging need, and avoid publishing artifacts in broadly accessible CI logs.

6. Prove the workflow in its target environment

A workflow that passes on a developer laptop may behave differently in CI because browser versions, dependencies, authentication, network routes, CPU and memory limits, or concurrency differ. Begin with one high-value workflow and reproduce failures in the same environment that will run it. Expand browser coverage and parallelism after the workflow is stable and you have measured its behavior.

  1. Pin or otherwise control browser and runtime dependencies used by the run.
  2. Run with the same CI image, authentication setup, and network path intended for production.
  3. Record failure category and evidence: locator mismatch, unmet state, expired authentication, dependency outage, or resource limit.
  4. Increase concurrency gradually and watch for collisions, timeouts, and shared-resource pressure.
  5. Review retry-pass cases and recurring failures as reliability work, not merely dashboard noise.

There is no universal production architecture or success percentage established by the cited guidance. Choose targets that reflect the workflow’s impact and dependencies, then validate them in your own operating environment.

7. Choose an execution setup that fits the team

Compare frameworks and execution options against the requirements you actually have:

Question What to evaluate
Language and skills Whether the team can maintain the tests and fixtures over time.
Browser coverage The browsers and device conditions your product must support.
Synchronization Whether actions wait for actionability and assertions wait for expected state.
Isolation How contexts, sessions, and application data remain independent.
CI compatibility Browser installation, system dependencies, authentication, networking, and resource limits.
Diagnostics Whether failures preserve enough timeline, page, and request evidence to debug.
Execution ownership Whether operating browser infrastructure yourself or using managed execution suits your constraints.

Playwright provides actionability waits, asynchronous assertions, contexts, retries, and traces. Selenium’s official guidance emphasizes environment-specific recommendations rather than one universally correct design. Microsoft documents Playwright Workspaces as one hosted remote-browser option; that establishes a category to evaluate, not a requirement for every team (Playwright Workspaces).

8. Troubleshoot common failures

Symptom Likely cause What to do
Element not found Locator depends on changed DOM structure, wrong accessible name, or page state is not ready. Inspect the trace and DOM snapshot; prefer a role and name or explicit test contract, then wait for the expected state.
Click times out Element is hidden, moving, disabled, covered, or the locator matches multiple elements. Check actionability details and current page state; make the locator unique and fix the underlying UI condition instead of forcing a click.
Works locally, fails in CI Different browser/dependencies, stale auth, network behavior, or constrained resources. Reproduce with the CI image and credentials; compare browser versions and inspect console/network evidence.
Passes only on retry Timing race, shared state, external instability, or resource contention. Keep the retry record, classify the cause, and repair it; do not count the retry as proof of stability.
Duplicate record or repeated side effect A retry replayed a write that the first attempt had already committed. Check for the intended record or status before replay; use application-supported idempotency where available.
Trace or screenshot missing Artifact policy does not collect it for that outcome, or the CI job does not retain the output. Review runner configuration and artifact upload paths; collect selectively and confirm retention while respecting data access controls.
Long and inconsistent run times Fixed sleeps, excessive retries, overloaded workers, or unnecessary serialization. Replace sleeps with condition waits, inspect retry rates, and raise concurrency only after checking shared data and resource capacity.

9. Keep performance, reliability, and cost in balance

  • Wait on conditions: remove broad fixed delays so ready pages proceed promptly while slower pages still have a bounded chance to complete.
  • Keep diagnostics targeted: traces and videos help explain failures but consume runtime and storage; collect them according to a failure policy.
  • Scale carefully: concurrency can shorten a suite, but shared accounts, test data, external services, and worker limits can make results less reliable.
  • Make recovery explicit: retries cost time and may repeat side effects. Prefer state checks and safe recovery paths.
  • Measure your own baseline: track duration, failure categories, retry-pass cases, and artifact volume in the target environment. The cited sources do not provide a universal benchmark.

10. Or skip the browser setup

If the job is to capture a page image or PDF rather than exercise an interactive workflow, ScreenshotNeo provides a website screenshot API and MCP server. Its one-call API can capture a URL without you setting up a browser runner:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. The response identifies the page verdict and billing status in headers.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

11. FAQ

Should every test get its own browser?

Use independent contexts and data wherever practical. The exact boundary depends on how the application and account model work; document and control unavoidable shared state.

How many retries should a production suite have?

There is no universal number. Set a small limit appropriate to the workflow, report retry-pass cases, and investigate them rather than treating retries as a stability guarantee.

Can screenshots prove an automation workflow succeeded?

A screenshot can show visible page state, but a workflow should assert the domain outcome that matters. For consequential writes, confirm the committed state through an appropriate application signal.

Do hosted browsers remove the need to debug flaky tests?

No. Hosted execution changes where browsers run; locators, state synchronization, isolation, dependencies, and recovery behavior still need to be reliable.