How to Test Large Language Models at Scale
Build a repeatable LLM evaluation program: define what you need to learn, test representative cases, measure uncertainty, and inspect failures.
To test large language models at scale, define the decision your evaluation will inform, build a representative set of real tasks, lock the model and scoring protocol, automate repeatable runs, and report uncertainty alongside scores and failures. A benchmark score answers a bounded question about a particular test setup; it does not establish broad production quality on its own.
For application evaluations, pair established benchmarks with cases from your own workflows. For AI agents, evaluate the full trace—including tool calls, handoffs, and guardrails—not only the final answer. Treat evaluation as a continuing measurement program: keep a stable regression set, add newly observed cases, and rerun after changes to models, prompts, tools, or retrieval.
1. Define the decision and claim
Begin by writing down what you will decide from the result. Examples include choosing between two models for a support workflow, checking whether a model follows a structured-output constraint, or assessing whether an agent handles a risky request safely. State the exact claim under test and what evidence would support or weaken it.
Specify the intended users, tasks, operating context, and risks. If comparing models, decide in advance which conditions must be equivalent. If testing a safeguard, define the behavior class and how success and failure will be scored. NIST’s draft guidance for automated benchmark evaluations starts with evaluation objectives and benchmark selection, and notes that automated benchmarks do not meet every evaluation objective. Treat it as draft guidance, not a finalized standard. NIST AI 800-2 announcement.
2. Build a representative evaluation set
Use public benchmarks to establish a common reference point, then add application-specific examples that reflect the work your system must do. Define the sampling frame: relevant users, task types, languages, input lengths, data conditions, and edge cases. A result only generalizes to a population that your examples reasonably represent.
- Keep a stable regression set. Use it to detect changes across model, prompt, or application versions.
- Maintain a separate refresh set. Add new cases over time so repeated development does not tune exclusively to a fixed, visible test set.
- Mine real examples carefully. Production logs can reveal useful cases and failure patterns. Apply privacy, access, and governance controls before using them.
- Record exclusions. Explain which users, languages, task types, or unusual inputs are outside the evaluation’s scope.
OpenAI’s evaluation guidance recommends task-specific tests that reflect real-world distributions, logging during development, continuous evaluation, and human calibration of automated scoring. OpenAI evaluation best practices.
Coverage should be complementary, not a claim that one suite is exhaustive. HELM is an example of shared scenarios and metrics; NIST’s ARIA program describes model testing, red-teaming, and field testing; NIST GenAI supports measurement and benchmark development across generative AI activities. Choose methods that fit the question and operating risk. HELM paper, NIST ARIA, and NIST GenAI.
3. Lock the evaluation protocol
The setup is part of the result. Version and record enough detail for another team to understand what ran and, where possible, repeat it. For each run, capture:
- Model identifier and version, provider or local runtime, and inference settings.
- System and user prompts, templates, and any prompt transformations.
- Dataset version, sampling method, split, and the cases actually run.
- Retrieval context, tools, tool descriptions, guardrails, and agent harness configuration.
- Sampling and retry behavior, output limits, timeouts, and stopping rules.
- Scorer version, rubric, aggregation method, and any judge model and prompt.
- Runtime environment and relevant execution conditions.
For model comparisons, use equivalent conditions where possible; describe unavoidable differences. If outputs are stochastic and the decision depends on repeatability, run selected cases more than once and preserve each result instead of silently keeping the best one. The lm-evaluation-harness paper discusses sensitivity to evaluation setup and reproducibility problems. lm-evaluation-harness paper.
4. Choose metrics and graders that match the claim
Use the simplest grading method that validly measures the target. Report metric definitions and aggregation rules, not only a composite score.
| Evaluation target | Useful grading approach | Check |
|---|---|---|
| Exact format, required fields, or a hard constraint | Deterministic assertions or schema validation | Test malformed, missing, and boundary values |
| Code or structured actions with objective outcomes | Executable tests, sandboxed checks, or outcome validation | Control the environment and record test versions |
| Quality with multiple acceptable answers | Explicit rubric and sampled human review | Calibrate graders and inspect disagreements |
| Broad subjective comparison | Pairwise comparison or rubric-based model grading, with human calibration | Document judge, prompt, order effects, and known failure modes |
An LLM judge can increase grading throughput, but its score is not self-validating. Compare a sample of its judgments with human judgments, inspect disagreements, and track judge changes just as you track model changes. OpenAI’s guide recommends human calibration and notes that comparison, classification, or rubric scoring can fit model strengths better than unconstrained generation. OpenAI evaluation best practices.
5. Automate runs without hiding failures
Scale by automating the same explicit protocol across cases and versions. Persist raw inputs, outputs, scores, errors, timestamps, and configuration alongside summaries. Parallelize or batch only within the provider’s rate limits and your own capacity; record concurrency, retries, backoff, and timeout policies because they affect execution and cost.
- Load a versioned dataset and evaluation configuration.
- Run each case under the recorded model and application setup.
- Store the raw response and execution metadata before scoring.
- Apply deterministic checks and configured graders.
- Collect errors and grader disagreements as results, rather than dropping them.
- Review aggregate changes and a sample of outputs, including failures.
Do not equate throughput with validity. A fast run can still use unrepresentative examples or a misaligned grader. Retry transient errors according to a declared policy, but preserve the first attempt and retry history. Avoid silently retrying a poor model answer as if it were a transport failure.
6. Evaluate agents at the workflow level
When the product is an agent, include the steps between request and final response. A correct-looking final answer can hide a wrong tool choice, unsafe action, failed handoff, or guardrail regression.
- Capture traces of model calls, tool calls and results, guardrails, and human handoffs.
- Grade tool choice, arguments, sequence, policy compliance, handoff behavior, and end-to-end outcome where relevant.
- Debug representative traces first; turn recurring cases into a versioned dataset for repeatable runs.
- Evaluate with the same tool access, instructions, and budgets the product will use, or disclose differences.
OpenAI’s agent evaluation guidance recommends moving from trace debugging to datasets and repeatable runs for larger comparisons over time. OpenAI agent evaluation guide.
7. Quantify uncertainty and generalization
Before computing an interval or declaring a winner, name the quantity you are estimating. Benchmark accuracy is performance on the exact included questions. Generalized accuracy estimates performance over a broader universe of similar questions. These targets can differ and require different estimation approaches.
NIST AI 800-3 discusses this distinction, emphasizes explicit statistical assumptions, and illustrates generalized-accuracy analysis with generalized linear mixed models (GLMMs). Its example analyzes 22 frontier LLMs on GPQA-Diamond, BIG-Bench Hard, and Global-MMLU Lite; those figures describe that report’s illustration, not a universal ranking. Item selection adds uncertainty when the intended claim extends beyond the tested set. NIST AI 800-3 announcement.
For a practical evaluation, report the sample size, the uncertainty method and assumptions, and whether uncertainty covers only run-to-run variation, item sampling, or both. Do not claim a meaningful ranking if the uncertainty does not support one. If cases are grouped by user, task, or source, use an analysis that respects that grouping rather than treating every response as independent by default.
8. Include risks and operating context
Accuracy is only one dimension. Depending on deployment, evaluate robustness, adversarial inputs, privacy and policy behavior, modality-specific cases, and whether the system fails safely. NIST ARIA describes model testing, red-teaming, and field testing as distinct levels; NIST GenAI covers work across modalities, adversarial evaluation, benchmark creation, and prompting effects. These examples support selecting complementary methods for the context; they do not imply every project needs the same test battery. NIST ARIA and NIST GenAI.
9. Publish an interpretable report
A report should let a reader tell what the result establishes and what it does not. Include:
- The decision, claim, and intended population of tasks.
- System, model version, prompts, tools, harness, and execution conditions.
- Dataset sources, versions, sampling, splits, sample size, and exclusions.
- Metric definitions, graders, aggregation, run budget, retries, and concurrency.
- Scores with uncertainty and the assumptions behind it.
- Failure analysis, grader disagreements, and known validity limitations.
- Raw artifacts or enough detail to audit them when sharing is safe and appropriate.
NIST AI 800-2 centers analysis and reporting in its draft workflow; NIST AI 800-3 stresses disclosing assumptions. HELM’s release of prompts and completions is an example of a transparency practice. NIST AI 800-2, NIST AI 800-3, and HELM paper.
Choosing evaluation tooling
Choose tools against the actual workflow rather than a generic feature checklist. Useful comparison questions include:
- Can it run your hosted APIs and local models?
- Can you define custom tasks as well as use established benchmark suites?
- Does it preserve dataset versions, configurations, and repeatable runs?
- Can it combine deterministic checks, human review, and model-based grading?
- For agents, does it capture traces and support tool and handoff grading?
- Can you control concurrency, retries, observability, and cost accounting?
- Can it report uncertainty and export raw results?
- Do privacy, access control, deployment, audit, and portability meet your needs?
These are selection criteria inferred from the measurement needs described by NIST, OpenAI, and the lm-evaluation-harness paper; they are not a head-to-head product comparison.
Performance, reliability, and cost
Evaluation cost grows with the number of cases, repeat runs, input and output size, tools invoked, and grader calls. Estimate a pilot before a full run: count expected model calls, retries, agent steps, and judge calls, then measure actual usage. Keep the run budget in the report. Prioritize representative cases and decision-relevant comparisons; do not remove hard cases simply to make a run cheaper.
Parallel execution can reduce elapsed time but may increase rate-limit errors or make failures harder to diagnose. Use bounded concurrency, explicit timeouts, and retry policies for transient transport failures. Cache only when the cache key includes all inputs and configuration that affect the result, and distinguish cached results from fresh model outputs. Preserve raw artifacts so a scorer change can be applied without rerunning the model when appropriate.
Troubleshooting common evaluation problems
| Symptom | Likely cause | Fix |
|---|---|---|
| Two teams get different scores | Different model versions, prompts, sampling, data, graders, or harness settings | Compare run manifests and lock the full protocol; disclose unavoidable differences. |
| A benchmark win does not improve the product | The benchmark distribution does not match production tasks | Add representative workflow cases and check the sampling frame against intended use. |
| A score changes sharply after grader updates | The scoring rule or judge changed, not necessarily the model capability | Version graders, rescore saved outputs, and report the change separately. |
| Results vary across runs | Stochastic generation, changing service versions, or unstable execution conditions | Record settings and versions, repeat cases when needed, and report run-to-run variation. |
| Agent score looks good despite user-visible failures | Only final answers were graded; tool calls, handoffs, or guardrails were missed | Capture and grade end-to-end traces and inspect representative failures. |
| Many cases time out or hit rate limits | Concurrency exceeds limits, budgets are too small, or timeouts do not fit task length | Bound concurrency, set task-appropriate timeouts, apply recorded backoff, and retain error records. |
| A model appears to overfit the evaluation | Repeated tuning against a fixed visible set | Keep a hidden or held-back regression set and refresh a separate case pool. |
| Confidence intervals appear overly narrow | Analysis treats related items or repeated runs as independent | State the estimand and assumptions; use a method that reflects item and run structure. |
Or skip the browser setup
For evaluations that include rendered web pages—for example, checking whether an agent selected the right page or whether a visual workflow changed—you can capture a URL with one ScreenshotNeo API request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
FAQ
How often should I rerun an LLM evaluation?
Rerun it when a relevant part of the system changes, such as the model, prompt, retrieval data, tools, or grader. A continuous evaluation process can also rerun key checks during development and before releases.
Should I use a public benchmark or my own test set?
Use both when they serve different purposes: public benchmarks provide a shared reference, while application-specific cases test whether the system handles your intended workflows.
Does a higher score mean a model is safer or better in production?
Only for the measured tasks, conditions, and scoring rules. Broader claims need representative evidence, appropriate risk testing, and a clear account of uncertainty and limitations.
Can an LLM grade another LLM?
It can help scale subjective scoring when paired with a defined rubric, human calibration, and checks for disagreement and judge failure modes.
What is the difference between benchmark and generalized accuracy?
Benchmark accuracy concerns the exact tested items. Generalized accuracy concerns a broader population of similar items and must account for uncertainty introduced by selecting a sample.


