ScreenshotNeo

BlogEngineering

How to Generate Test Data with Generative AI

Generate test data with AI by defining a schema, choosing a generation method, and validating quality, coverage, and privacy before use.

By the ScreenshotNeo team4 October 202613 min read

To generate test data with generative AI, first define the behavior you need to test, the target schema, and the business rules the data must satisfy. Then choose whether you need a few values, a reusable generator, a synthesized dataset based on source tables, or inputs for generated test cases. Parse and validate every output before using it. AI-generated or synthetic data is not automatically private, representative, or correct.

This guide covers a repeatable workflow, runnable examples, validation, privacy controls, tool-specific options, troubleshooting, and operational tradeoffs.

1. Define the test objective before generating data

Start with the behavior under test and the outcomes you expect. A prompt such as “make realistic users” leaves the model to guess which cases matter. Instead, describe concrete scenarios and acceptance criteria.

  • Behavior: What feature, API, validation rule, or failure path are you testing?
  • Scenarios: Include ordinary, boundary, invalid, and rare combinations that matter to the test.
  • Fields: List field names, types, nullability, formats, ranges, allowed values, and uniqueness requirements.
  • Relationships: Specify foreign keys, shared identifiers, ordering, and cross-field invariants.
  • Expected outcomes: State what the application should accept, reject, calculate, or display for each scenario.
  • Quantity and repeatability: Decide how many records you need and whether the same seed or input must reproduce them.

Keep the test objective distinct from realism. A plausible address or transaction is not useful if it fails to exercise the intended edge case. For AI systems, keep test data separate from training, validation, and evaluation data unless the evaluation design explicitly calls for otherwise.

2. Choose the right generation shape

Generative AI can produce individual values, a complete data file, or code that generates data. Research on LLM-based test-data generation describes prompting for raw data, generator code, and code that uses faker libraries as distinct approaches. Structured-data platforms and test-case tools address different workflows.

Approach Good fit Watch for
Prompted values A small fixture or a handful of isolated examples. Output formatting drift, invalid fields, and weak repeatability.
Model-generated generator code A reusable script for a defined set of cases. Review and execute generated code in a controlled environment; enforce constraints in code rather than trusting comments.
Faker-backed generator Many varied records with familiar types such as names, dates, and addresses. Faker output does not know your hidden business rules; add deterministic constraints and scenario-specific cases.
Warehouse-native synthesis Artificial rows shaped from existing tables, where column types and relationships matter. Check edition requirements, null handling, output behavior, and privacy controls for the particular platform.
Test-case tool population Populating inputs associated with generated or captured test cases. This may be a test-case workflow, not a general-purpose dataset generator; check its environment settings and defaults.

Choose by input basis, output shape, schema fidelity, relational consistency, privacy controls, repeatability, integration, data volume, and operational dependencies. The available sources do not establish one universally best method or provide an independent head-to-head benchmark.

3. Ask for structured output and enforce it in code

For a small fixture, specify a schema and request JSON only. Use non-sensitive examples. Treat the response as untrusted input: parse it, validate fields, and reject or repair violations before tests consume it.

Example prompt

Generate 8 JSON objects for testing a subscription signup endpoint.
Return a JSON array only, with no prose or markdown.

Schema:
- email: string, syntactically valid, unique within this batch
- plan: one of "free", "team", "enterprise"
- seats: integer from 1 through 500
- trial_days: integer from 0 through 30
- country: two-letter uppercase code

Include these scenarios at least once:
- ordinary free signup
- minimum seats and zero trial
- maximum seats and maximum trial
- invalid business combination: enterprise with 1 seat (expected to be rejected)
- missing optional country (use null)

Do not use real people's details. Do not include extra keys.

There is a deliberate tension in this example: the schema says country is a two-letter code, while one requested scenario uses null. In production, make nullability explicit in the schema, for example country: string or null. This illustrates why prompts should be reviewed for contradictions and why validation must encode the actual contract.

Python: request and validate JSON

This example uses an OpenAI-compatible chat completions endpoint as an illustration. Set the endpoint and model to the provider you use; the provider’s documentation defines its request format, authentication, retention, and structured-output support. No real service endpoint or model behavior is assumed here.

import json
import os
import requests

endpoint = os.environ["LLM_CHAT_COMPLETIONS_URL"]
api_key = os.environ["LLM_API_KEY"]
model = os.environ["LLM_MODEL"]

prompt = """Return a JSON array only, containing 3 test records.
Each object must have:
- email: unique string
- plan: one of free, team, enterprise
- seats: integer 1..500
- trial_days: integer 0..30
Include an ordinary case, a boundary case, and an invalid business
combination (enterprise with 1 seat). Do not use real people's details."""

response = requests.post(
    endpoint,
    headers={"Authorization": f"Bearer {api_key}"},
    json={
        "model": model,
        "messages": [{"role": "user", "content": prompt}],
        "temperature": 0,
    },
    timeout=60,
)
response.raise_for_status()
content = response.json()["choices"][0]["message"]["content"]
records = json.loads(content)

if not isinstance(records, list):
    raise ValueError("Expected a JSON array")

seen_emails = set()
for index, record in enumerate(records):
    if set(record) != {"email", "plan", "seats", "trial_days"}:
        raise ValueError(f"Record {index}: unexpected or missing keys")
    if not isinstance(record["email"], str) or not record["email"]:
        raise ValueError(f"Record {index}: email must be a non-empty string")
    if record["email"] in seen_emails:
        raise ValueError(f"Record {index}: duplicate email")
    seen_emails.add(record["email"])
    if record["plan"] not in {"free", "team", "enterprise"}:
        raise ValueError(f"Record {index}: invalid plan")
    if type(record["seats"]) is not int or not 1 <= record["seats"] <= 500:
        raise ValueError(f"Record {index}: seats out of range")
    if type(record["trial_days"]) is not int or not 0 <= record["trial_days"] <= 30:
        raise ValueError(f"Record {index}: trial_days out of range")

print(json.dumps(records, indent=2))

For production use, add a JSON Schema or equivalent validator and validate business rules separately. A model response can be syntactically valid while still violating your application’s contract.

Example: deterministic generator with Faker

For bulk fixtures, use a generator program with deterministic rules. A faker library can supply varied base values, but the program should control IDs, ranges, edge cases, and relationships. Install the dependency with python -m pip install Faker.

from faker import Faker
import json
import random

fake = Faker()
Faker.seed(2026)
rng = random.Random(2026)

plans = ["free", "team", "enterprise"]
records = []
for index in range(20):
    plan = plans[index % len(plans)]
    records.append({
        "id": f"user-{index + 1:04d}",
        "email": f"test-user-{index + 1:04d}@example.invalid",
        "plan": plan,
        "seats": 1 if index == 0 else rng.randint(2, 50),
        "trial_days": 0 if index % 4 == 0 else rng.randint(1, 30),
    })

# Explicit edge case, not left to chance.
records.append({
    "id": "user-enterprise-min-seats",
    "email": "enterprise-min-seats@example.invalid",
    "plan": "enterprise",
    "seats": 1,
    "trial_days": 30,
})

print(json.dumps(records, indent=2))

The code uses a reserved invalid domain for test email addresses and a fixed seed to make its pseudo-random choices repeatable. A fixed seed does not make a dataset representative, and library changes may alter generated values. For stable snapshots across environments, pin dependencies and keep explicit fixtures for important cases.

4. Generate relational data without breaking joins

When tests span tables, define the relationship model before generation. For example, a customer, order, and order-line dataset needs stable customer IDs, order-to-customer references, and order-line-to-order references. Generate parent rows first or create a shared key map; then validate every foreign key.

  • Specify primary-key uniqueness and foreign-key existence.
  • Describe cross-column rules, such as order totals matching line totals.
  • Decide whether the same entity must retain the same key across tables or repeated runs.
  • Include orphan, duplicate, and cyclic-reference cases only when those are intentional test scenarios.

Snowflake documents a warehouse-native option, GENERATE_SYNTHETIC_DATA, for generating a table using source column names and types, with artificial values described as statistically similar. Its documentation describes treatment for statistical fields, categorical strings, and non-categorical strings; non-categorical strings are redacted unless a replacement output format is specified. Join-key handling and a consistency secret can support consistent keys across runs or tables. These are documented behaviors, not a guarantee that the result is correct for your tests or private.

-- Illustrative call shape; verify argument names and options against
-- the current Snowflake procedure reference for your account.
CALL GENERATE_SYNTHETIC_DATA(
  'TEST_DB.TEST_SCHEMA.SOURCE_TABLE',
  'TEST_DB.TEST_SCHEMA.SYNTHETIC_TABLE'
);

Snowflake documents this procedure as requiring Enterprise Edition or higher. Its optional similarity filter removes rows according to documented nearest-neighbor distance measures; it is not a complete privacy guarantee. Snowflake warns that enabling the filter fails when non-string columns contain nulls. Review the current Snowflake synthetic data guide and procedure reference before using it, since product behavior and syntax can change.

5. Populate generated test cases carefully

Some tools populate data as part of test-case generation or captured workflows. Katalon TrueTest documents Disabled, Raw, Raw with PII mocked values, and Synthetic modes. Its documentation says modes are configured by tracking environment, Disabled is the default, and changing modes requires contacting TrueTest support. It describes Synthetic mode as using an AI-based model to generate realistic values based on captured patterns. This is a product-specific test-case workflow, not a general-purpose dataset synthesizer. Check the current Katalon TrueTest test-data documentation for current behavior.

Enterprise test-data-management services may combine privacy assessment, masking, mining and provisioning, synthetic generation, or database virtualization. For example, Infosys describes such a service. IRI describes RowGen for referentially correct test data in production-like formats, but the cited material does not substantiate a generative AI feature. Treat vendor descriptions as capability claims rather than independent comparisons: Infosys test data management and IRI RowGen.

6. Validate data before a test suite consumes it

Use automated checks for properties that can be checked mechanically, and human review for scenario intent and unexpected sensitive-looking output. Validation should fail closed: invalid generated records should not silently enter a shared test environment.

  1. Parse and type-check: reject malformed JSON, missing fields, wrong types, extra fields, and invalid encodings.
  2. Enforce schema: validate nullability, formats, ranges, enums, uniqueness, and length limits.
  3. Check business invariants: evaluate conditional rules and calculated fields in application code or a domain validator.
  4. Check relational integrity: verify primary and foreign keys, shared identifiers, and intentional orphan cases.
  5. Check scenario coverage: prove each required ordinary, boundary, invalid, and rare case exists; do not infer coverage from a record count.
  6. Check distribution fit: where production-like patterns matter, compare relevant aggregate properties against approved, non-sensitive targets. Plausibility alone is not a distribution check.
  7. Review privacy risk: inspect for sensitive-record matches and consider whether auxiliary information could identify people.
  8. Check repeatability: regenerate when needed and compare stable invariants or snapshots; record prompt, model/tool version, seed, and schema version where policy permits.
  9. Run tests against the data: confirm the fixtures trigger the expected application outcomes, not merely that they pass a validator.

AWS’s generative AI testing guidance discusses practices such as holdout data, human evaluation, and adversarial testing. The UK government Data and AI Ethics Framework recommends testing through build and after launch and using anonymised or synthetic data where possible. These are evaluation practices, not a single validated score for test-data quality: AWS testing guidance and the UK Data and AI Ethics Framework.

7. Protect privacy and control data handling

Do not treat “synthetic” as synonymous with anonymous. Risk depends on what data informed the generation, what was sent to an external service, how outputs compare with real records, who can access them, and what auxiliary information an attacker could use. UK government guidance warns that AI can re-identify people thought to be anonymised by linking information. ISTQB training material also notes that an LLM could generate values matching sensitive data; it does not establish a measured probability.

  • Prefer schemas and invented examples over real personal data in prompts.
  • Review provider terms, retention, access, and training settings before sending data; do not assume defaults.
  • Keep prompts, outputs, credentials, and generated datasets within approved access controls and retention periods.
  • Check for plausible real-person matches and consider auxiliary-data re-identification risk.
  • Use any similarity filter as one control in a broader threat model, not as proof of privacy.
  • Reassess when the source data, model, prompt, environment, or downstream use changes.

For AI-system evaluation, keep test material separate from training and validation data according to the evaluation design. The Australian Government AI Technical Standard discusses this separation and the use of synthetic data to supplement dataset completeness.

8. Troubleshooting common failures

Symptom Likely cause Fix
Response contains prose or markdown around JSON The prompt requests JSON but the endpoint does not enforce structured output. Use the provider’s structured-output or JSON mode if available; otherwise extract cautiously and reject parse failures. Validate the parsed shape.
Missing fields, extra keys, or wrong types The schema was vague, long, or contradictory, or the model did not follow it. Simplify the schema, make nullability explicit, split large requests, and enforce a schema validator after generation.
Records look valid but tests miss edge cases Coverage was left to random generation or described without a measurable requirement. Require named scenarios and assert their presence. Create deterministic fixtures for rare and boundary combinations.
Duplicate IDs or broken joins Independent prompts generated related tables without shared key rules. Generate identifiers in code or from a shared key map; validate uniqueness and every foreign key.
Generated values are too similar to source records The workflow learned from or was given source data, and outputs were not screened for similarity or re-identification risk. Reduce sensitive inputs, review access and retention, run privacy checks, and assess the threat model. A similarity filter is not a blanket guarantee.
Snowflake similarity-filter call fails on nulls Snowflake documents failure when the optional filter is enabled and non-string columns contain nulls. Review the current procedure guidance; handle nulls according to the documented constraints or adjust the workflow before enabling the filter.
Snowflake procedure unavailable The account does not meet the documented edition requirement. Snowflake documents Enterprise Edition or higher as required; confirm account edition and current feature availability.
Katalon data mode remains Disabled Disabled is documented as the default, and the page says mode changes require TrueTest support. Check the tracking environment and follow the current product documentation’s support process.
Fixtures change after a model or dependency update Generation is nondeterministic or tool/library behavior changed. Pin versions where feasible, use seeds for pseudo-random generators, save critical fixtures, and compare semantic invariants rather than assuming identical output.
Provider request times out or is rate-limited Large output, service load, request limits, or network instability. Generate smaller batches, use bounded retries with backoff for transient failures, and log request identifiers without logging sensitive prompts or outputs.

9. Performance, reliability, and cost considerations

There is no source-backed universal claim that generative AI test-data workflows save time, lower costs, or reduce defects. Measure the workflow in your environment, including generation, validation, human review, repair, and maintenance effort.

  • Batch size: Large prompts can be truncated or produce inconsistent records. Start with a small batch, validate it, then scale in bounded batches.
  • Determinism: Model settings such as temperature do not guarantee identical results across providers or model revisions. For reproducible CI, keep critical fixtures under version control or generate them with deterministic code.
  • Failure handling: Separate transient provider errors from invalid output. Retry transient errors with limits; do not blindly retry malformed output indefinitely.
  • Cost accounting: Include model or platform charges, warehouse compute where applicable, engineering time for constraints and validation, storage, and review. Compare against the alternative workflow using your own measurements.
  • Data volume: Use code or platform-native generation for large batches where practical; use LLMs for scenario design or constrained examples only when that fits the system and policies.
  • Change control: Version schemas, prompts, generator code, and test expectations. Revalidate after model, source-data, dependency, or downstream-system changes.

10. A practical checklist

  • Test objective and expected outcomes are written down.
  • Schema, nullability, relationships, and business constraints are explicit.
  • Ordinary, boundary, invalid, and rare cases are named and counted.
  • Generation method matches output size, structure, and repeatability needs.
  • Sensitive input, provider handling, output access, and retention are reviewed.
  • Schema, business rules, joins, scenario coverage, and privacy checks pass.
  • Important fixtures are reproducible and versioned.
  • Generated data has been exercised against the system under test.
  • Validation is repeated after material changes to prompts, models, data, or use.

Or skip the browser setup

If part of your test-data workflow is collecting reference screenshots of pages or checking rendered test fixtures, ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns an image or PDF. It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. AI agents can take screenshots through its MCP server. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write("shot.webp", res);

Use https://screenshotneo.com to learn about the service and the docs for request options. Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

Can generative AI produce test data without a schema?

It can produce examples from a description, but without explicit fields, types, relationships, and rules there is no reliable contract to validate against. Define those first for repeatable use.

Is synthetic test data automatically anonymous?

No. Generated output can resemble sensitive records or become identifying when combined with other information. Assess inputs, outputs, access, and re-identification risk.

Should I use an LLM or a faker library?

Use the method that fits the data shape and controls you need. A faker-backed program is useful for repeatable varied values; an LLM can help create constrained examples or generator code. Both require validation and explicit edge cases.

Can I use generated data to evaluate an AI model?

It can supplement evaluation data, but maintain separation from training and validation data as required by the evaluation design, and verify that generated cases represent the behaviors being measured.

Sources