ScreenshotNeo

BlogAI agents

How to Improve Reliability in Agentic Software Development

Build a reliability workflow for agentic software: define success, evaluate complete tool-using tasks, add safeguards, and monitor real use.

By the ScreenshotNeo team4 October 20268 min read

To improve reliability in agentic software development, evaluate the complete task the agent performs, make trials repeatable, limit how untrusted input can affect tools, and monitor real use. Define success before implementation, test realistic multi-step workflows in isolated environments, inspect traces as well as outcomes, and turn production failures into new evaluations. No single benchmark or guardrail proves an agent is reliable for every task.

1. Decide whether an agent fits the task

An agent uses a model to manage a workflow and tools to interact with external systems. That flexibility is useful when work involves complex decisions, hard-to-maintain rules, or unstructured data. For a routine with clear inputs and deterministic rules, a conventional program may be simpler to reason about and operate. Start by validating that the task needs the agent’s ability to choose steps or interpret variable inputs. OpenAI’s practical guide to building agents describes these fit considerations.

2. Define reliability in terms of user outcomes

Before choosing a model or building a workflow, write down what a correct result means for representative user tasks. Include the expected final state, acceptable variations, prohibited actions, and important failure conditions. A useful evaluation is tied to the task and the distribution of cases the system will actually encounter, rather than a generic score.

  • Outcome: What should be true when the task is done?
  • Constraints: Which actions, data access, or side effects are disallowed?
  • Failure conditions: What would make an apparently successful response unsafe or incomplete?
  • Regression cases: Which past user-impacting failures must continue to be caught?

OpenAI recommends defining objectives, data, metrics, comparisons, and an iteration process, and calibrating automated graders against human judgment. Log behavior so real failures can become evaluation cases. See OpenAI’s evaluation best practices. The cited page includes a deprecation notice for its Evals platform; check the live notice before choosing implementation tooling.

3. Evaluate the whole multi-step workflow

Many agent errors only become visible across multiple turns: a weak interpretation leads to a poor tool choice, which changes state and causes later steps to fail. Run evaluations through the actual loop of model, tools, instructions, and environment. Grade both the final task state and the steps taken to get there.

For coding agents, tests can verify behavior, but passing tests alone may miss poor tool choices, instruction violations, or risky actions. Review traces alongside outcomes. OpenAI’s agent evaluation documentation describes trace grading for debugging and repeatable datasets and evaluation runs for comparing behavior over time.

  1. Choose realistic tasks, including edge cases and prior failures.
  2. Run the agent with the same tools and meaningful constraints it will have in use.
  3. Check final state against task-specific criteria.
  4. Inspect traces for unnecessary actions, bad tool selection, instruction handling, and uncertainty.
  5. Record results and compare changes against a fixed set of cases.

4. Make evaluation trials repeatable

Start each trial from a clean, isolated environment. Shared files, caches, leftover processes, resource exhaustion, or other state can make trials dependent on one another. This can create correlated failures or make results look better than they are. Keep the environment stable and close enough to production to represent the system users will encounter.

When a trial fails, preserve the inputs, environment details, tool results, and trace needed to reproduce it. Separate an agent failure from an infrastructure failure, such as an unavailable dependency or exhausted resource. Anthropic’s guide to agent evaluations discusses isolated trials and combining automated evaluation with production monitoring and human review.

5. Put boundaries around inputs and tool actions

Retrieved pages, files, and tool outputs are untrusted input. Prompt injection is text that attempts to override the agent’s instructions. Avoid letting raw untrusted content directly determine an action. Where possible, extract and validate specific structured fields, constrain tools to the authority needed for the task, and require confirmation for consequential operations.

  • Validate inputs at the boundary and prefer structured fields to free-form instructions.
  • Limit which tools the agent can call and what data or systems each tool can reach.
  • Use approval steps for consequential MCP operations.
  • Keep guardrails layered: a guardrail node alone is not foolproof.
  • Review traces and evaluate adversarial cases to learn where controls fail.

Structured outputs and isolation can reduce risk, but do not eliminate it. See OpenAI’s agent safety guidance for input handling, tool approvals, and layered controls.

6. Monitor deployment and feed failures back into evaluations

Offline evaluations help teams iterate before release; production monitoring reveals distribution changes and failures that the test set missed. Combine automated evaluations with monitoring, user feedback, transcript review, and periodic human assessment. Use controlled comparisons such as A/B tests when they fit the rollout and its risk.

Define what to monitor for the particular application: task completion, failed or repeated tool calls, policy violations, unexpected state changes, and user-reported problems. Review incidents and add representative cases to the regression suite. OpenAI’s report on monitoring internal coding agents describes categories such as restriction circumvention, concealed uncertainty, reward hacking, unauthorized data transfer, destructive actions, and prompt injection. These are examples of monitored behaviors in that report, not estimates of how often they occur across the industry, and monitoring does not guarantee every action is blocked before it happens.

7. Treat benchmark scores as evidence with limits

A benchmark measures performance on its tasks and graders; it does not establish reliability across all real workflows. Audit the problem statements and tests. OpenAI’s July 2026 SWE-Bench Pro audit identified issues including tests stricter than the prompt, underspecified prompts, low-coverage tests, and misleading prompts. Its headline estimate was approximately 30% broken tasks. The report separately says an automated pipeline flagged 200 of 731 public-split tasks (27.4%), while human annotation identified 249 of 731 (34.1%); these are different methods and should not be conflated. It also reported that a frontier model’s pass rate on that public split rose from 23.3% to 80.3% over eight months. That is a result on this benchmark split, not a stable general measure of coding-agent reliability.

When assessing a benchmark result, ask whether the task matches your workflow, whether tests reflect stated requirements, whether tests cover likely incomplete fixes, and whether the environment and grader are reproducible. Use benchmark results as one input alongside your own task evaluations and production evidence.

8. Choose evaluation and observability tooling by workflow

Tooling should fit how you run trials, inspect traces, compare versions, and meet hosting or data-residency needs. Anthropic describes Harbor as oriented to containerized trials, Braintrust as combining offline evaluation and production observability, LangSmith as integrated with the LangChain ecosystem, and Langfuse as a self-hosted open-source alternative. These are descriptions from Anthropic’s article, not a current independent feature audit; verify present capabilities and fit before adopting a tool.

For any option, check isolated trial support, task and grader definition, trace capture, offline evaluation, production monitoring, experiment tracking, hosting requirements, and integration with your development stack. OpenAI’s evaluation guidance distinguishes trace grading for debugging from repeatable dataset runs for ongoing comparison.

9. A practical improvement loop

  1. Pick a bounded workflow. State why an agent is useful for it and identify the systems it may change.
  2. Write success and failure criteria. Include realistic inputs, edge cases, forbidden actions, and user-visible outcomes.
  3. Build a repeatable evaluation. Use isolated starting state, stable dependencies, the real tool loop, and task-specific graders.
  4. Inspect traces. Look at intermediate decisions and tool calls, not just final answers.
  5. Add layered controls. Validate untrusted input, constrain access, and require approval where the consequences warrant it.
  6. Release with monitoring. Track failures and feedback in the deployed distribution.
  7. Update the suite. Convert meaningful incidents into regression cases and rerun them after changes.
  8. Audit the graders. Confirm that prompts and tests reward the intended result and that benchmark tasks are sound.

Or skip the browser setup

If your agent workflow needs website screenshots as an input, ScreenshotNeo provides a screenshot API and MCP server. One GET request can return an image or PDF; the API supports PNG, JPEG, and WebP output.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo API documentation

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each removal step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses report page verdict and billing headers.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card.

Troubleshooting reliability work

Symptom Likely cause What to do
Results vary between identical trials Shared state, changing dependencies, nondeterministic external services, or resource pressure Reset to an isolated starting state, pin dependencies where practical, record environment details, and distinguish infrastructure failures from agent failures.
All tests pass but users still report failures The evaluation cases do not match the real task distribution, or tests check only the final response Review traces and user incidents, add representative tasks, and grade the required final state and important actions.
The agent follows instructions found in retrieved content Untrusted text is influencing behavior or directly driving tool calls Treat retrieved text as data, extract validated fields, constrain available actions, and evaluate prompt-injection cases.
A benchmark score seems implausible or changes sharply Task or grader defects, environment changes, or split-specific effects Audit prompts and tests, reproduce the run, and report the benchmark and split with the score.
Monitoring finds an issue after damage occurs The signal is asynchronous or controls did not require approval before the action For consequential tools, put authorization and action constraints before execution; use monitoring as detection and learning, not the only control.
Evaluation failures cannot be reproduced Inputs, tool responses, environment state, or version information were not retained Capture the trace and relevant environment and dependency details for each run, while handling sensitive data according to your application’s requirements.

Performance, reliability, and cost considerations

More evaluation cases and human trace reviews require time and compute, so prioritize cases by user impact and likelihood, then retain a representative regression set. Run inexpensive automated checks frequently and reserve deeper human review for ambiguous, high-impact, or newly observed behavior. Keep trials isolated without making the evaluation environment so artificial that it stops representing production.

Reliability depends on the full system: model, prompt, tools, external services, state, grader, and operational controls. Record versions and relevant run conditions so a score can be interpreted and compared. The sources provide workflow guidance, not a universal reliability guarantee or cost benchmark; estimate costs from your task volume, tool usage, evaluation frequency, and review process.

FAQ

Is a coding test suite enough to evaluate an agent?

No. Tests can verify code outcomes, while trace review can reveal poor tool choices, instruction handling, or risky behavior that the final test result misses.

Do guardrails prevent prompt injection?

No single guardrail guarantees prevention. Treat untrusted content as data, constrain actions, validate structured fields, and test the workflow with adversarial inputs.

Should every task use an agent?

No. For well-specified routines with stable rules, a deterministic implementation may be easier to maintain and evaluate.

Can a benchmark score predict production reliability?

Only partly. It reflects performance on a particular set of tasks and graders. Validate it against your workflow and audit task and test quality.