Can AI Test My App? How AI Helps with Software Testing
AI can draft and help maintain browser tests, but people still need to check that tests assert the right behavior. Here’s a practical Playwright workflow.
Yes. AI can help test a web app by drafting browser tests from recorded interactions or plain-language scenarios, and some workflows can run tests and help diagnose or repair failures. It does not establish that an app is fully tested or bug-free. Treat generated tests as drafts: check that their steps and assertions match the behavior you intend to protect, then run and review them.
This guide uses Playwright, whose Codegen tool records browser actions and generates test code. It also covers an AI-assisted workflow, practical safeguards, common problems, and how to decide where AI fits in your testing process.
1. What AI can and cannot do in software testing
For browser-based apps, AI can help turn a user flow into test code, explore an app to propose scenarios, and—in some documented workflows—run tests and attempt to repair failures. Playwright describes both interaction recording through Codegen and a planner, generator, and healer agent workflow. The plan and any generated changes still need review.
A useful test checks an outcome that matters. For example, a checkout test should verify the expected order confirmation or state change, not merely that someone clicked the checkout button. A generated test can encode a mistaken expectation and still pass consistently.
The cited workflows concern browser automation and test-code generation. They do not establish universal coverage for native mobile, performance, security, accessibility, or every other quality dimension. Keep the checks those areas require.
2. A practical first workflow: record a Playwright test
If your web app already uses Playwright, start with one important user flow. Codegen opens a browser and an inspector; perform the flow, then review and adapt the generated code.
Install and record
In a Node.js project, install Playwright Test if it is not already part of your test setup:
npm init playwright@latest
Then record a flow against your local app. Replace the URL with the route and app you want to test:
npx playwright codegen http://localhost:3000
In the opened browser, perform a representative scenario—for example, sign in with a test account and open a page. The inspector generates code as you interact. Save the useful steps in a test file such as tests/account.spec.ts, then add or refine assertions for the outcomes your app is meant to guarantee.
Example runnable test
This example checks that a user can open an account page and sees a heading. Change the route, accessible name, and expected heading to match your app. It assumes the app is available at the configured base URL and has a test account or a page accessible without sign-in.
import { test, expect } from '@playwright/test';
test('account page shows its heading', async ({ page }) => {
await page.goto('/account');
await expect(
page.getByRole('heading', { name: 'Account' })
).toBeVisible();
});
Run it with:
npx playwright test tests/account.spec.ts
Playwright’s locator guidance favors role, text, and test ID locators. Codegen attempts to make ambiguous matches unique. Prefer selectors tied to the user-visible interface or an intentional test ID, and review generated selectors when the page has repeated labels or dynamic content. See the Playwright Codegen documentation.
3. Add AI assistance without handing it the final say
A browser-connected assistant can use app context to draft a scenario or adapt a recorded flow to project conventions. Microsoft’s documented example workflow runs the app, connects an assistant to a browser through Playwright MCP, asks for a scenario, then reviews and commits the generated test.
- Start the app with the same configuration and test data used by your team.
- Describe one specific user scenario, including the starting state and expected result.
- Let the assistant inspect the browser or adapt a recorded happy path when the scenario is complex.
- Review the generated test against the app’s requirements, test data, and established code conventions.
- Run the test, inspect failures, and commit only changes that a reviewer understands.
For more agent involvement, Playwright documents a planner that explores an app and creates a test plan, a generator that turns the plan into test files, and a healer that executes the suite and attempts to repair failing tests. The documentation labels this as next-version material, so check current stable-version support and setup requirements before adopting it. A repair attempt is a proposed change, not proof that the test or app is correct. See Playwright Test Agents and Microsoft’s Playwright MCP testing workflow.
4. Review generated tests before relying on them
- Check the scenario: Does it cover a real user task and a meaningful starting state?
- Check each assertion: Does it verify the required outcome, rather than just an action or incidental page detail?
- Check data and access: Are credentials, fixtures, and test records safe and reproducible?
- Check selectors: Will they identify the intended control when labels repeat or content changes?
- Check failure behavior: Does a failure explain what broke, and does the test avoid hiding a real product regression?
- Review every generated code change: Generated code can be inaccurate or insecure. GitHub recommends carefully reviewing and testing Copilot output; apply the same care to generated tests.
Microsoft’s sample workflow includes reviewing generated tests before committing them. Make that a normal part of the process: a passing test only provides evidence for the behavior it actually checks. See Microsoft’s workflow guidance and GitHub’s guidance on generated code.
5. Troubleshooting AI-assisted browser tests
| Symptom | Likely cause | What to do |
|---|---|---|
| Codegen opens the wrong page or cannot reach the app | The app is not running, or the URL or port is wrong. | Start the app, confirm the local URL in a browser, then rerun Codegen with that URL. |
| A generated locator matches multiple elements | The page contains repeated text, roles, or labels. | Use a more specific user-facing locator or add an intentional test ID; avoid selecting an arbitrary match without checking the page. |
| A test passes locally but fails in a clean run | It may depend on existing session state, data, or timing. | Make setup and test data explicit, and assert the expected state rather than adding an unexplained fixed delay. |
| A test fails after a visual or copy change | The test may rely on incidental wording or structure. | Decide whether the changed text is part of the contract. Update the test if behavior is still correct; otherwise investigate the regression. |
| An AI repair makes the test pass but changes its meaning | The repair may have weakened or removed a meaningful assertion. | Compare the change to the intended behavior, restore important checks, and rerun the scenario. |
| An AI-generated test is syntactically valid but checks the wrong thing | The prompt or inferred expected behavior was incomplete or mistaken. | Rewrite the scenario with a clear starting state and expected outcome, then verify each assertion against requirements and app behavior. |
6. Performance, reliability, and cost considerations
The cited sources describe capabilities and workflows; they do not establish measured time savings, defect detection rates, or coverage improvements. Avoid assuming AI will reduce a particular amount of testing time. Measure your own workflow if those outcomes matter.
Reliability depends on clear scenarios, controlled test data, useful assertions, and review of generated changes. Keep generated tests readable so failures can be diagnosed and maintenance decisions are deliberate. For external AI or browser-connected tools, check access and data-handling requirements for the app and test environment before sharing context.
Tool costs and setup vary, and the research for this guide does not establish a neutral vendor pricing comparison. Choose based on fit with your existing language and test suite, required browser and platform coverage, quality of failure evidence, review controls, setup, and data handling.
7. Or skip the browser setup
If you need a screenshot of a page as part of a visual check, ScreenshotNeo is a website screenshot API and MCP server. A request can return a PNG, JPEG, WebP, or PDF. It captures a page; it does not replace a browser test or verify that your app behaves correctly.
Example cURL request, with the API details in the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
8. Frequently asked questions
Can AI test my app without existing tests?
It can help draft an initial browser test from recorded actions or a described scenario. You still need to define expected behavior and review the result.
Does a passing AI-generated test mean my app is bug-free?
No. It only provides evidence for the paths and outcomes the test checks. It does not establish complete coverage or correctness across other quality areas.
Should I use AI for every test?
Start with a valuable, clearly specified browser flow. Keep the approaches that produce readable tests your team can review and maintain.
Can ScreenshotNeo replace Playwright?
No. ScreenshotNeo captures web pages, while Playwright automates browser interactions and assertions. A screenshot can support visual review, but it does not prove an interaction or expected application outcome.


