ScreenshotNeo

BlogHow-to

How to Use AI for Automated Website Testing

Use AI to plan and draft browser tests, then validate them with Playwright, trace failures, and combine automated accessibility scans with human review.

By the ScreenshotNeo team4 October 20268 min read

Use AI to help plan test cases, draft browser interactions, and inspect failures. Run the resulting tests with a repeatable browser test runner such as Playwright, and review every generated scenario against the user outcome it is meant to verify. AI can speed up drafting and browser exploration; it does not establish that a test is correct. Automated accessibility scans also catch only some issues, so combine them with manual assessment and inclusive user testing.

1. Define the behavior before asking AI

Choose one real user journey, such as account creation, search, checkout, or form submission. Write down its starting state, the user action, and the observable result that means success. Include important failure cases, such as invalid input or an unavailable result.

  • Starting state: What data and page state must exist?
  • Action: What does the user do?
  • Expected outcome: What should the user be able to observe?
  • Failure behavior: What should happen for invalid or missing input?

Give the AI these requirements and ask for a small set of scenarios, including expected outcomes and any assumptions. Ask it to mark assumptions instead of inventing product behavior. This keeps the test grounded in the feature contract rather than in whatever happens to be visible in the page.

2. Generate a first draft with Playwright

Playwright offers two useful starting points: its code generator records interactions as you perform them, and its documented agent workflows support planning, generating, and healing tests. Treat either output as a draft. Check that each step and assertion represents the intended user behavior. See the Playwright test generation guide and release notes for current workflows.

To record a flow locally, install Playwright and start codegen:

npm init playwright@latest
npx playwright codegen https://your-site.example

Use the opened browser to perform the journey. Inspect the generated code, remove irrelevant steps, add assertions for the outcomes you wrote down, and save it in your test suite. Codegen recommends locators based on page content, with role, text, and test-ID locators among its preferred choices.

3. Write a reviewable test

Here is a complete example for a search page. Replace the URL, accessible labels, and expected result with behavior from your own application. The test assumes the search field has an accessible name of “Search” and submitting a query displays a heading containing the query.

import { test, expect } from '@playwright/test';

test('search shows results for the submitted query', async ({ page }) => {
  await page.goto('https://your-site.example');
  await page.getByRole('searchbox', { name: 'Search' }).fill('blue shoes');
  await page.getByRole('button', { name: 'Search' }).click();
  await expect(page.getByRole('heading', { name: /blue shoes/i })).toBeVisible();
});

Save it as tests/search.spec.ts in a Playwright project and run npx playwright test. Playwright’s runner auto-waits for actionability and retries assertions. Prefer locators that describe the interface as users encounter it:

  • getByRole() for buttons, links, headings, and other semantic elements.
  • getByLabel() for labeled form controls.
  • getByPlaceholder() when placeholder text is the appropriate identifier.
  • getByText() for visible text where a role-based locator does not fit.
  • getByTestId() for a stable test contract your team intentionally maintains.

Avoid treating a generated CSS path or a successful click as proof of correct behavior. Add assertions that check the actual result, and keep selectors aligned with the application’s accessible semantics where possible. Playwright documents its test runner, locator, waiting, assertion, and isolation features; these help with reliable execution but cannot validate whether the scenario captures the right requirement.

4. Choose browser and device coverage

Playwright supports Chromium, Firefox, and WebKit, plus branded browsers and device emulation. Select projects based on the browsers and devices your users rely on and the risk of the journey. A critical checkout flow may warrant broader coverage than a low-risk internal page. Keep Playwright current so tests run against recent browser versions. See Playwright browser documentation.

Start with the smallest useful matrix: the primary browser and viewport, then add engines or device profiles for known audience needs and risk. Emulation is useful for checking layouts and browser behavior, but it is not the same as validating every real device configuration.

5. Use traces to understand failures

When a test fails, first determine whether it exposed a product defect, a test that no longer matches the intended behavior, or an environment problem. Playwright trace artifacts can include an execution timeline, DOM snapshots, network requests, console logs, and screenshots. Review that evidence before accepting an AI-suggested repair; a repair that makes the test green can still weaken its assertion or change what it verifies.

npx playwright test --trace on

Use the trace viewer and test output to identify the first meaningful divergence: a request failed, a page did not reach the expected state, a locator matched the wrong element, or the assertion contradicts the current product contract. Fix the cause, then rerun the relevant test.

6. Add accessibility checks, with human review

Automated accessibility scans can flag some detectable issues, including low contrast, unlabeled controls, and duplicate IDs. Playwright documents use of @axe-core/playwright with axe-core. A scan is one check in an accessibility process, not a certification: many issues need manual testing. Combine automated checks with manual assessment and inclusive user testing, as described in the Playwright accessibility testing guide.

import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';

test('home page has no detected accessibility violations', async ({ page }) => {
  await page.goto('https://your-site.example');
  const results = await new AxeBuilder({ page }).analyze();
  expect(results.violations).toEqual([]);
});

Install the integration with npm install --save-dev @axe-core/playwright. Adapt the scan to your pages and review any exclusions deliberately. Add keyboard and assistive-technology assessment to cover needs a rule-based scan cannot judge.

7. Keep AI-generated tests maintainable

  • Keep each test focused on a user outcome and isolate its setup so it does not depend on another test’s run order.
  • Review generated assumptions, selectors, and assertions before committing the test.
  • Use semantic locators and explicit assertions. Do not add arbitrary sleeps to cover timing uncertainty when the runner can wait for the relevant condition.
  • Keep test data and environment requirements understandable to the team.
  • When a test fails, use execution evidence to distinguish application behavior from test or environment problems.
  • Revisit browser coverage when your supported audience or risk changes.

AI can propose scenarios, code, or a repair, but a reviewer must decide whether the proposed test still checks the user promise. The documented browser and runner capabilities support repeatable testing; they do not guarantee correct test intent.

8. Troubleshooting common problems

Symptom Likely cause What to do
Locator times out The element is absent, its accessible name differs, or the page is in an unexpected state. Inspect the DOM snapshot and page state; verify the user-facing name and prerequisite navigation. Prefer a semantic locator that matches the actual control.
Test passes locally but fails in CI Different browser versions, environment data, configuration, or timing expose a dependency. Inspect the trace and CI logs, compare environments, and make setup explicit. Avoid masking the issue with a longer fixed sleep.
Generated test clicks the wrong element Several elements share similar text or the generated locator is too broad. Use a role locator with a more specific accessible name or scope it to the relevant region; assert the resulting state.
Test is green but misses a regression The assertion checks an incidental detail or only that an action completed. Restate the expected user outcome and assert that result directly. Review the test against the feature requirement.
Accessibility scan reports no violations, but users still encounter barriers The issue requires human judgment or is outside the automated rules being run. Perform manual assessment and include people with relevant access needs in testing. Treat scan output as partial evidence.
AI “healing” changes a locator or assertion The page changed, or the proposed repair weakens the intended check. Review the diff and trace against the user expectation. Accept only a change that preserves the requirement.

9. Performance, reliability, and cost considerations

The research sources establish Playwright capabilities, but do not provide comparative speed, reliability rates, or pricing figures for AI testing tools. Avoid using unsourced benchmark claims to choose a workflow. For practical control, begin with a focused journey, run the relevant test during development, and choose broader browser coverage according to user impact and risk. Keep tests isolated and use runner waiting and retrying assertions instead of arbitrary delays. This makes failures easier to interpret; it does not eliminate environment variability.

Account for the resources your setup actually consumes: browser processes, test environments, CI execution, and any AI service you choose. Record the provider’s current pricing and limits separately, since they depend on the service and plan. The workflow can also start with Playwright’s documented code generation and locally run tests without assuming that a paid AI service is required.

Or skip the browser setup

For rendered screenshots as visual evidence in a testing workflow, ScreenshotNeo is a website screenshot API and MCP server. A screenshot can help inspect a page state, but it does not replace assertions that verify behavior.

One GET request returns an image or PDF. Full API options are in the ScreenshotNeo documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before capture; more than 60 known consent platforms, newsletter popups, and chat widgets are removed. Each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 shots a month with no card. Paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, no card required.

FAQ

Can AI replace a QA engineer?

No. It can help draft scenarios and interact with browser tooling, while people remain responsible for deciding what should be tested and whether results meet user needs.

Does a passing accessibility scan mean a site is accessible?

No. Automated scans find some issues; manual assessment and inclusive user testing are also needed.

Which browsers can Playwright test?

Its supported browser engines include Chromium, Firefox, and WebKit. It also documents branded browser and device emulation options; choose coverage for your audience and risk.

Should I accept AI-generated test repairs automatically?

No. Check that the change preserves the expected user behavior and does not weaken the assertion.