How to Use AI for Test Automation
Use AI to draft and debug tests with a workflow that grounds suggestions in requirements and live application evidence, then verifies every change.
Use AI to draft or adapt a bounded test, suggest browser locators, and interpret real failures. Give it the requirement, project conventions, trusted documentation, and evidence from the running application. Then verify the behavior and locators, run the test repeatedly, and review the code and dependencies before merging. AI can speed up parts of test work, but generated tests still need human review; a passing test does not establish complete coverage.
1. Choose a bounded test task
Ask for a specific contribution tied to a requirement or code change. Useful tasks include drafting a unit test for a named function, listing edge cases for a validation rule, writing an API test for a documented response, or proposing a browser scenario for a user journey.
Provide the relevant requirement, source files, test framework and version, nearby tests, project conventions, and trusted project documentation. Include constraints such as supported language version, test data rules, and whether network access is allowed. GitHub recommends checking generated output against the project’s requirements, purpose, architecture, and patterns, and supplying trusted project documents as context. GitHub Copilot best practices
Example prompt for a unit test
Write tests for the function validateCoupon in src/coupons.ts.
Requirement: reject expired coupons, accept an active coupon, and reject a coupon whose minimum cart total is not met.
Use the test framework and conventions shown in test/coupons.test.ts. Do not change production code or add dependencies.
Return the proposed test and identify any assumptions that the requirement does not settle.
The prompt asks for a small change, points to local conventions, and requires assumptions to be surfaced. Confirm that the examples and requirement are current before relying on the resulting tests.
2. Ground browser tests in the running application
For browser automation, give the model access to the application state or supply concrete observations from a running instance. Ask it to propose locators and explain what each one identifies, then verify those locators against the live page before adopting them. Selenium’s AI-agent guidance recommends checking the application instead of guessing at its state and locators. Selenium documentation
Prefer locators tied to accessible names, labels, or stable test attributes where the application provides them. Treat any suggested selector as a hypothesis: confirm it matches the intended element, remains unique in the relevant state, and works after the page has loaded. Do not assume a selector is correct because it looks plausible in generated code.
3. Supply real failure evidence
When a test fails, share the actual exception, relevant stack trace, test code, logs, and the expected behavior. For a browser failure, a screenshot or a description of the visible page at failure time can help distinguish a wrong locator from a navigation, timing, or application problem. Remove secrets and sensitive user data before sharing diagnostic material.
Ask the model to explain what the evidence shows, list plausible causes, and propose the smallest change that addresses the requirement. Selenium’s guidance describes using concrete exception details and screenshots to diagnose failures rather than relying on a guessed cause. Selenium documentation
4. Run the test and repeat it
Run the individual test first, inspect its result, fix issues, and run it several times before treating it as stable. A single passing run may miss a race condition or timing-sensitive failure. Once the focused test is reliable, run the relevant suite and the checks required by the project.
- Run the new or changed test by itself.
- Inspect failures against the requirement and the actual application state.
- Repeat the run to look for intermittent behavior.
- Run related tests and the project’s normal suite.
- Run static analysis and any dependency or license checks used by the project.
Do not let an AI suggestion turn a failing test into a skipped or deleted test without understanding why it failed and whether the requirement changed.
5. Review the generated change before merging
Read the test as code you are responsible for maintaining. Check that it asserts the intended behavior, exercises meaningful cases, uses the right fixtures, and does not rely on accidental details of the current implementation. Verify API names and dependencies against trusted documentation. Generated code can contain hallucinated APIs, incorrect logic, ignored constraints, or unnecessary dependencies.
- Compare each assertion with the stated requirement.
- Check edge cases and failure paths relevant to the change.
- Confirm locators and interactions against the running application.
- Review added dependencies, their legitimacy, and applicable licenses.
- Inspect removed, disabled, or skipped tests and understand the reason.
- Run static analysis and the relevant test suite.
- Keep human review as a gate, especially for generated code or outputs used in practice.
GitHub’s guidance specifically calls out hallucinated APIs, incorrect logic, ignored constraints, and tests that are deleted or skipped instead of fixed. GitHub Copilot best practices
6. Test AI-enabled products as AI systems
If the product under test uses a model, ordinary application tests are only part of the work. OWASP’s AI Testing Guide Version 1.0 groups assessment into AI application, AI model, AI data, and AI infrastructure testing, with a repeatable sequence: define the objective, execute the test, interpret the response, and recommend remediation. OWASP AI Testing Guide
Define representative inputs and adversarial inputs that could expose failure modes, including prompt injection where relevant. Evaluate outputs across a range of inputs, because model performance may vary by case. Review outputs with a person wherever possible, especially when the system generates code or other consequential content. These are safeguards, not proof that a model or testing tool is safe. OpenAI safety resources
7. Capture browser evidence for test review
A screenshot can make a browser failure easier to inspect and share with a reviewer. In a DIY workflow, capture the browser state at the point of failure, preserve the relevant logs, and make sure the image corresponds to the same run as the exception. If you need a repeatable URL-based capture as part of debugging or documentation, ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API returns a screenshot or PDF, and its MCP tools include take_screenshot, get_page_info, and capture_pdf.
Or skip the browser setup
Make one request to capture a page as an image. Replace the example target URL with the page you need, and use your ScreenshotNeo access key. See the ScreenshotNeo API documentation for the request details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month, with no card.
Common errors and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| Generated test calls a method that does not exist | The model inferred an API or used documentation for another version. | Check the installed version and trusted project or framework documentation; correct the call and add no dependency unless it is needed and approved. |
| Browser locator finds nothing | The page state differs from the prompt, the selector is wrong, or the element has not appeared yet. | Inspect the live page and failure screenshot, verify the selector against the intended element, and use the framework’s appropriate wait for the actual state. |
| Test passes once and fails on another run | A race, timing assumption, shared state, or unstable test data may be involved. | Repeat the test, inspect logs and state, isolate shared data, and wait for a meaningful condition instead of adding an arbitrary delay without evidence. |
| AI proposes skipping or deleting the failing test | The suggestion may hide the failure instead of addressing it. | Keep the test until the cause is understood; compare the expected behavior with the requirement and fix the underlying issue. |
| Test passes but does not catch the regression | The assertions may not cover the required behavior or relevant edge case. | Map assertions to the requirement, add focused cases for missing paths, and review whether the test would fail if the behavior were broken. |
| Generated change adds an unexpected package | The model may have invented or unnecessarily selected a dependency. | Verify the package, source, license, and necessity; prefer the existing project stack when it can solve the task. |
| AI feature behaves differently across inputs | Model outputs can vary across cases, and a narrow test set may miss failure modes. | Evaluate representative and adversarial inputs, inspect the output, and include AI application, model, data, and infrastructure concerns in scope. |
Performance, reliability, and cost
Keep AI tasks narrow enough that a reviewer can verify the result. Use relevant context instead of sending an entire repository when a focused set of files, tests, and requirements will do. Run generated tests in the same environment and CI path as the project tests so local success is not mistaken for reliable behavior. Repeat tests that may be timing-sensitive and investigate flaky outcomes rather than masking them with retries alone.
No general productivity, quality, or cost-saving figure is established by the sources used for this guide. A 2024 study reports reviewing 55 AI-based test automation tools and empirically assessing two selected tools on two open-source projects; those study scope figures are not a general performance result or proof of vendor superiority. Study record
Frequently asked questions
Can AI decide whether my test suite has complete coverage?
No. Use it to suggest cases, then compare them with requirements, risks, and the behavior your team needs to protect. The cited guidance does not establish that AI independently finds every important case.
Should I use AI-generated tests without reviewing them?
No. Verify behavior, framework calls, dependencies, and test results before merging generated changes.
What should I give an AI when a browser test fails?
Share the actual exception and relevant logs, the test and requirement, and where useful a screenshot of the failure state. Remove credentials and sensitive data first.
Do I need a separate testing approach for an AI feature?
Include the application, model, data, and infrastructure in the assessment, and exercise representative and adversarial inputs with human review.
Further reading
Software Testing with Generative AI is an optional learning resource; the cited material includes a chapter on AI-assisted testing for developers and examples involving GitHub Copilot and ChatGPT. Check current availability and edition details before purchasing. Referenced resource record


