ScreenshotNeo

BlogGuides

AI-Driven Testing: How to Automate Website QA with AI

Use AI to draft website tests, then run and review them with Playwright. This guide covers setup, runnable examples, accessibility checks, debugging, and screenshot capture.

By the ScreenshotNeo team4 October 202611 min read

AI can help draft website tests from a user journey or recorded browser actions. A browser test runner still executes the checks, and a person should review that every action and assertion matches a real requirement. A practical stack is Playwright for browser automation, its Codegen recorder or an AI assistant for the first draft, and trace files for diagnosing failures.

This guide builds a runnable Playwright test, shows how to use AI responsibly in the workflow, and covers accessibility checks, reliability, performance, troubleshooting, and screenshot capture.

1. What AI-driven website QA means

“AI-driven testing” can mean several different things. In this workflow, AI helps author or inspect tests; Playwright drives the browser and decides whether explicit checks pass. Keeping those roles distinct makes the result easier to review.

Part What it does What it does not establish
AI assistant Drafts a scenario from a written requirement, helps inspect a page, or adapts recorded code to project conventions. That the scenario is correct, complete, or aligned with the business rule.
Playwright Runs browser actions and assertions across Chromium, Firefox, and WebKit, with auto-waiting, retries, isolated contexts, parallel execution, and traces. That a passing test covers every important user outcome.
Reviewer Checks test intent, selectors, data, assertions, and failure evidence before committing. That automated checks replace exploratory, accessibility, or user testing.

There is no quantified accuracy or productivity result established by the cited documentation. Treat generated code as a draft and judge it by the same standards as hand-written tests.

2. Install Playwright and write a first test

The example below uses TypeScript and Playwright Test. It checks a simple form journey on a local application. Replace the sample route, labels, and expected confirmation with the behavior your product requires.

npm init playwright@latest

Follow the prompts to choose TypeScript and install the browsers. Add a test such as tests/contact.spec.ts:

import { test, expect } from '@playwright/test';

test('visitor can submit the contact form', async ({ page }) => {
  await page.goto('http://127.0.0.1:3000/contact');
  await page.getByLabel('Name').fill('Riley Example');
  await page.getByLabel('Email').fill('riley@example.com');
  await page.getByLabel('Message').fill('Please send product details.');
  await page.getByRole('button', { name: 'Send message' }).click();
  await expect(page.getByRole('status')).toContainText('Message sent');
});

Run it with npx playwright test. The app must be running at the URL in the test. If the generated starter project has a configured webServer, Playwright can start the app as part of the test run.

Why these locators and assertions?

getByLabel and getByRole describe the controls in terms of how a user encounters them. The final assertion verifies an outcome, rather than merely checking that a click did not throw an error. Playwright recommends user-facing locators such as role, text, and labels, while test IDs can be useful when they provide a stable contract. Avoid selectors coupled to incidental DOM nesting when a semantic locator is available. Playwright locator documentation.

3. Use Codegen or AI to create a draft

Playwright Codegen records browser interactions and generates starter code. It can add assertions for visibility, text, and values, and its locator generation prioritizes role, text, and test IDs. Start your app, then run:

npx playwright codegen http://127.0.0.1:3000

Use the opened browser to perform one short, high-value journey. Save the output into a test file, then review every step. Codegen is useful for capturing selectors and interaction sequences; it cannot decide whether the journey represents the acceptance criteria.

An AI assistant can also draft a test from a requirement. Give it the project’s test conventions and a precise outcome, for example:

Write a Playwright Test in TypeScript for the contact form.
Use the existing test conventions in this repository.
Verify that a visitor can submit valid name, email, and message values,
and assert that the page announces the successful submission.
Use role and label locators where possible. Do not invent selectors;
if a label or success message is unclear, list the missing information.

If an AI assistant is connected to a live browser through an MCP integration, it may inspect the rendered page and help author a test from observed controls. Microsoft’s documented workflow combines a running browser, Playwright MCP, natural-language instructions, project conventions, and Codegen, then calls for reviewing and committing the result. Microsoft Learn: AI-assisted test authoring.

Review checklist for generated tests

  • Does each action belong to the stated user journey?
  • Does each assertion express an acceptance criterion, including the final outcome?
  • Are locators meaningful and stable, and do they target the intended element?
  • Are test data and environment assumptions safe and repeatable?
  • Does the test fail when the requirement is broken, rather than only when the page is unavailable?
  • Can a teammate understand and maintain the test without the original AI prompt?

4. Configure browsers, retries, and test execution

Playwright supports Chromium, Firefox, and WebKit. A generated configuration can include browser projects; a concise example is:

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  forbidOnly: Boolean(process.env.CI),
  retries: process.env.CI ? 2 : 0,
  reporter: 'html',
  use: {
    baseURL: 'http://127.0.0.1:3000',
    trace: 'on-first-retry',
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
    { name: 'firefox', use: { ...devices['Desktop Firefox'] } },
    { name: 'webkit', use: { ...devices['Desktop Safari'] } },
  ],
});

Then use relative paths, for example await page.goto('/contact'). The exact device presets and configuration options can change; consult the Playwright configuration reference for current details.

Retries are a diagnostic and resilience mechanism for intermittent failures, not evidence that a flaky test is healthy. When a retry passes, inspect why the first attempt failed and fix the underlying timing, isolation, or environment issue. Parallel runs can shorten elapsed time, but tests must not depend on shared mutable accounts or data.

5. Add accessibility checks without overclaiming

Automated accessibility rules can identify some issues, including contrast problems, missing accessible labels, and duplicate IDs. They do not find every accessibility problem. Playwright’s accessibility guidance recommends combining automation with manual assessments and inclusive user testing. Playwright accessibility testing.

A common integration uses axe-core’s Playwright package. Install it with:

npm install --save-dev @axe-core/playwright

Then add a scan to a meaningful page state:

import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';

test('contact page has no automatically detected serious accessibility violations', async ({ page }) => {
  await page.goto('/contact');
  const results = await new AxeBuilder({ page })
    .withTags(['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa'])
    .analyze();
  expect(results.violations).toEqual([]);
});

Use the tags and severity policy that fit your team’s standard, and review the findings rather than blindly suppressing them. A passing scan is not proof that a page is accessible or WCAG-conformant. Keyboard use, focus order, understandable instructions, and real user experience need human assessment.

6. Diagnose failures with traces and screenshots

Enable traces on retry or failure so a test report carries evidence. A Playwright trace can show a timeline with DOM snapshots, network requests, console logs, and screenshots. Open a trace with:

npx playwright show-trace path/to/trace.zip

Classify a failure before changing the test:

  1. Application regression: the expected control or result is genuinely broken.
  2. Test defect: the locator, assertion, data, or flow does not represent the requirement.
  3. Environment problem: the app, dependency, network, or test data was unavailable or inconsistent.
  4. Timing or race: the test assumes readiness that has not occurred. Prefer Playwright’s locator and assertion waiting behavior to fixed sleeps.

Use screenshots as supporting evidence when a visual state matters, but keep semantic assertions for functional requirements. A screenshot can show what rendered; it does not by itself prove that a form submitted or a business rule was enforced.

7. Make the test suite reliable and fast

  • Keep cases focused: a test should make one user outcome clear. Smaller failures are easier to diagnose.
  • Use isolated state: Playwright creates isolated browser contexts for tests. Avoid shared accounts or records that parallel tests can overwrite.
  • Wait on outcomes: use locator actions and web-first assertions that wait for the expected state. Fixed delays make tests slow and can still be too short.
  • Use deterministic data: seed or create known test data and clean it up where appropriate. Do not rely on a changing production page for a regression test.
  • Choose parallelism deliberately: parallel workers increase resource use and can expose shared-state problems. Start with a level the CI environment can support.
  • Keep traces useful: retaining traces on retry or failure gives evidence without collecting them for every successful run.
  • Test browser coverage by risk: run all supported engines for critical journeys and decide whether a smaller smoke set is sufficient for every commit.
  • Review maintenance cost: generated tests still need ownership as product flows, copy, and selectors change.

Playwright’s test runner includes auto-waiting, retrying assertions, isolation, parallel execution, and tracing. These features help structure browser tests; they do not guarantee reliable test design. See the Playwright documentation.

8. Capture a clean page screenshot for QA evidence

A browser-runner screenshot is useful for reviewing a specific test state. For a standalone capture—such as documenting how a public page renders—ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. Its clean-shot workflow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. See ScreenshotNeo and its API documentation.

For test assertions tied to authenticated state, dynamic data, or a precise interaction, use the browser test itself. A standalone URL screenshot is complementary evidence, not a replacement for the test runner.

Screenshot options relevant to QA

ScreenshotNeo offers full-page captures with lazy images loaded, CSS-selector element captures, dark mode, 12 device presets or a custom viewport, and retina scale. It also supports custom CSS and JavaScript, clicks before capture, hiding selectors, waiting for a selector, delay, or network idle, blocking ads, trackers, requests or resource types, and custom headers, cookies, user agent, and Authorization. Other available options include timezone, geolocation, transparent background, image resizing, a caller-selected cache TTL, signed links for public image tags, async jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI spec. PDF options include paper size, margins, landscape, and page ranges. HTML/CSS can also be rendered to an image.

9. Or skip the browser setup

For a standalone website screenshot, call the API directly. Replace the target URL as needed and keep your API key private. See the ScreenshotNeo API docs for parameters and formats.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; response headers say the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.

10. Troubleshooting

Symptom Likely cause What to do
Browser executable is missing The package is installed but its browser binaries are not. Run npx playwright install or install the required browser listed by the project setup.
Navigation fails with connection refused The local server is not running, or the test URL or port is wrong. Start the app, verify the URL in a browser, and align baseURL and webServer configuration.
Strict mode reports multiple matches The locator matches more than one element. Make the locator more specific using its accessible name, a containing region, or a stable test ID. Confirm the intended target before narrowing it.
Element is not actionable or test times out The page may not be ready, an overlay may block the control, or the locator may be wrong. Inspect the trace and DOM snapshot. Wait for the actual expected state, handle the overlay if it is part of the flow, or correct the locator. Avoid arbitrary sleeps.
Test passes locally but fails in CI Different browser dependencies, timing, resources, environment variables, or shared data. Compare browser versions and configuration, inspect CI traces, stabilize test data, and reduce unsafe parallel sharing.
Retries pass but failures keep appearing A flaky test or unstable environment is being masked by retries. Use the first failed attempt’s trace to identify the race or dependency, then fix the cause instead of increasing retries alone.
Accessibility assertion reports violations The scan found rule violations, or its scope/tags need review. Inspect each finding in context, fix genuine issues, and document justified exceptions. Add manual checks for issues automated rules cannot detect.
Screenshot differs between runs Dynamic content, animation, viewport, fonts, or remote assets changed. Control the test state and viewport, disable or await animations where appropriate, and compare the relevant region. Do not hide meaningful regressions.

11. Cost and operational considerations

Playwright is an open-source browser automation framework, but running a suite still consumes CI compute and engineering time for test data, maintenance, and failure investigation. The researched sources do not establish a universal cost or time saving from AI-assisted test authoring. Estimate using your own run frequency, browser matrix, worker count, runtime, and maintenance needs.

ScreenshotNeo pricing is Free for 1,000 shots per month with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. Only clean shots are billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. These API captures are useful for repeatable public-page evidence; they do not replace assertions in the browser test suite.

12. Frequently asked questions

Can AI write Playwright tests?

Yes. It can draft code from a requirement or help adapt recorded browser actions. Review the test’s intent, selectors, data, and expected result, then run it and inspect failures before committing.

Can browser-action recording create a complete regression test?

It creates a starting sequence and may add basic assertions. A developer still needs to verify that the assertions cover the actual acceptance criteria and that the test behaves usefully when the requirement breaks.

Does an automated accessibility scan prove a site is accessible?

No. Automated scans find some detectable issues. Manual assessment and inclusive user testing remain necessary.

Which languages can Playwright tests use?

Playwright documents language bindings for TypeScript, Python, .NET, and Java. The runnable test in this article uses TypeScript.

Sources