ScreenshotNeo

BlogGuides

Automatic Test Creation: Common Questions and Answers

Learn what automatic test creation produces, how to prepare useful inputs, and how to review generated tests before relying on them.

By the ScreenshotNeo team4 October 20268 min read

Automatic test creation uses requirements, manual test cases, or natural-language descriptions to draft test cases, manual steps, or executable automation. Start by deciding which output you need: these workflows have different inputs and produce different artifacts. Treat generated content as a draft: review it, run it in the target environment, and refine it before relying on it.

1. What does automatic test creation mean?

The phrase covers several distinct workflows. A tool may propose test scenarios from requirements, expand a test case into manual steps, or write automation code from a saved case or prompt. Some products generate tests inside a specific platform or framework. “Automatic” describes how the draft is produced; it does not mean the result is automatically correct or ready for production.

Starting material Typical output What to check
Requirements or acceptance criteria Candidate test cases, and sometimes detailed steps Coverage, assumptions, expected results, traceability to requirements
A saved manual test case Automation code or framework-specific steps Selectors, setup, assertions, dependencies, and whether it runs in your project
A natural-language description Platform-specific tests or a draft test design Product and UI support, ambiguity, and missing context
Existing code or project examples Tests shaped to the supplied patterns Whether examples are current, representative, and safe to reuse

2. What information should I give a test generator?

Provide the clearest source material your chosen tool accepts. Good inputs reduce avoidable ambiguity, but they cannot guarantee correct or complete output.

  • State the behavior: describe the user action and the expected result, including important conditions and boundary cases.
  • Use explicit steps: name the action, target, and expected outcome. Avoid vague phrases such as “works correctly.”
  • Use consistent terminology: keep feature, field, and role names consistent across the requirement and examples.
  • Add implementation context when relevant: selectors, existing code, fixtures, configuration, coding conventions, and representative tests can help when generating executable automation.
  • Identify scope: specify the requirement section or feature to cover when the tool supports prompt-scoped generation.
  • Call out constraints: include supported browsers, roles, data conditions, integrations, or other environment assumptions the test must respect.

For example, “Verify checkout” is underspecified. A more useful requirement says which item is in the cart, what payment or validation condition to exercise, what the user does, and what visible result counts as success or failure.

3. What do current documented tools generate?

These examples show why it is important to check the specific product’s supported inputs and outputs. They are vendor-documentation snapshots, not a universal feature list; verify current availability and terms before adopting a tool.

Requirements to cases and steps

Katalon documents generating test cases from requirements, then generating steps using the case name, description, preconditions, and linked requirements. Its documented workflow requires AI features to be enabled and an ALM integration such as Jira or Azure DevOps. The page says the workflow retrieves summaries and descriptions from those ALM requirements; image attachments are supported there, while other attachment formats are not currently supported. Katalon recommends reviewing generated content before approval. See Katalon’s documentation.

Saved manual cases to automation

TestRail documents generating automation from one saved test case at a time. The listed choices are Java or Python with Selenium or Playwright; BDD-style cases map to Cucumber for Java or Behave for Python. The AI uses text fields from the case, not attachments or structured metadata. Project files such as selectors, examples, configuration, and coding conventions can provide additional context. See TestRail’s getting-started documentation.

Natural language to platform tests

ServiceNow’s Yokohama release documentation describes Test generation as accepting natural-language requirements and building on the Automated Test Framework. In that documented release, it is available only to Next Experience UI users. This is a version-specific product example, not a general requirement for test generation. See ServiceNow’s Yokohama documentation.

Scoping generated cases

BrowserStack’s FAQ says a prompt can name a section or line to scope test-case generation. It also documents ordering for a single input document and settings that can or cannot be changed during later iterations. See the BrowserStack test-case generator FAQ.

4. How do I review automatically generated tests?

Use a review process that checks both what the test says and what happens when it runs. TestRail says generated automation is intended to be reviewed, tested, and refined by a human; its best-practices guidance warns that code can look correct but fail in practice. Katalon also warns that AI-generated results may contain errors.

  1. Check source coverage. Trace each case to a requirement or intended behavior. Look for important paths and constraints that were omitted.
  2. Check assertions. Confirm each test has a meaningful expected result. A sequence of actions without an assertion may not verify the behavior.
  3. Check assumptions and data. Look for invented roles, values, states, or preconditions. Make test data explicit and safe for the target environment.
  4. Inspect code and selectors. Confirm selectors identify the intended elements and that setup, cleanup, waits, and error handling fit the project.
  5. Run in the target environment. Execute the test against the supported application version and configuration. A plausible draft is not execution evidence.
  6. Refine and record ownership. Edit or discard weak output, keep traceability to the source, and decide who will update the test when behavior changes.

Good prompting can improve relevance, but it does not prove correctness. A 2026 survey of 21 primary studies reports that its reviewed approaches did not satisfy all six quality dimensions it considered: automation, ambiguity handling, domain applicability, traceability, evaluation thoroughness, and hallucination control. This is the survey’s finding, not a universal benchmark or accuracy rate. See the survey abstract.

5. How should I choose a test-generation tool?

Decision area Questions to ask
Starting material Can it use your actual source: prompts, requirements, manual cases, existing code, selectors, images, or linked ALM records?
Output Does it produce candidate cases, manual steps, executable code, or tests tied to a particular platform?
Compatibility Which product tier, UI, languages, frameworks, ALM integrations, and attachment types are supported?
Control and ownership Can you inspect, edit, discard, export, run, and trace output in your own workflow?
Review and lifecycle Who approves the draft, checks execution evidence, and maintains the test when requirements change?
Data handling Where are inputs processed? What are retention, model-improvement, access-control, and opt-out terms?

Check current product documentation and contractual terms, since features and policies can change by product version, deployment, and plan.

6. What should I know about sensitive data?

Before entering proprietary requirements, credentials, customer information, or other sensitive content, inspect the current terms and configuration for the exact tool and deployment. Do not assume one vendor’s handling rules apply to another.

For example, ServiceNow’s Yokohama documentation says that its documented feature transfers data from customer instances to a centralized ServiceNow environment and potentially to third-party cloud infrastructure. It also describes using inputs, outputs, and edits to improve its technologies, with an opt-out for future data collection. Those statements apply to that documented ServiceNow feature; consult its current documentation and your configuration before use. See ServiceNow’s data-handling information.

7. Or skip the browser setup

If your work also needs website screenshots as test evidence or visual input, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It complements test generation; it does not generate or validate software tests. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, and failed loads are never billed, and responses identify the page verdict and billing status in headers. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

8. Troubleshooting generated tests

Symptom Likely cause What to try
Cases miss part of the requirement The input is broad, ambiguous, or contains multiple unrelated sections Clarify expected outcomes and scope generation to a named section when supported; compare every result to the source.
Generated steps lack expected results The source case has actions but no explicit assertions or nonempty expected results Update the source with observable outcomes, then generate or edit the steps again.
Automation uses wrong selectors or project patterns Relevant selectors, examples, or conventions were not included or were stale Provide current project files where supported and inspect every selector and fixture.
Output uses an unsupported language, framework, UI, or integration The product only supports specific combinations, tiers, or product interfaces Check the current compatibility documentation before generation; choose a supported target or another workflow.
Generated code looks plausible but fails Missing context, invalid assumptions, timing, test data, or code defects Run it, examine the failure, correct setup and assertions, and retain human review before merging.
Attachments do not inform the result The tool may ignore attachments or accept only specific attachment types Check documented input handling; put essential information in supported text fields or a supported file format.
Privacy review blocks use Data processing, retention, or model-improvement terms are unclear or unsuitable Consult current vendor terms and admin controls; omit sensitive material until approved.

9. Performance, reliability, and cost

There is no universal generation speed, accuracy rate, or cost figure supported across these products. Measure the workflow you plan to use: preparation and review time, number of drafts that need substantial edits, execution pass rate in your environment, maintenance effort, and tool or plan costs. A large batch of weak cases can cost more to review than a smaller, well-scoped generation.

For reliability, retain the source requirement, generated draft, human edits, and execution result together where your process permits. Re-run affected tests when requirements, application behavior, selectors, or dependencies change. Treat generation as an aid to drafting and coverage exploration, not a substitute for executable checks or review ownership.

10. Frequently asked questions

Can automatic test creation replace QA engineers?

No. It can help draft cases or code, but people still need to define expected behavior, assess coverage, validate output, and maintain tests.

Can I generate tests for only part of a requirement?

Some tools support scoped prompts. BrowserStack’s documented FAQ says users can name a section or line to scope generation. Confirm the current behavior in the product you use.

Should generated tests be committed immediately?

Only after the team’s normal code review and execution checks. Keep the draft editable and traceable to its source.

Does adding more context guarantee better tests?

No. Relevant, accurate context can make output more specific, but the generated result still needs review and execution.