How to Measure Software Quality in Agile Teams
Measure product quality and delivery performance together. Choose context-specific indicators, review trends, and use them to improve outcomes without reducing quality to one score.
How do you measure software quality in an Agile team? Define quality around the product’s users and risks, then track a small set of product outcome measures alongside delivery performance. ISO/IEC 25010:2023 can help organize product quality requirements; DORA’s current five delivery performance measures help teams understand throughput and instability. Neither is a universal quality score, so interpret trends in the context of the application or service.
This guide explains how to select measures, define them so the data is meaningful, review them with the team, and avoid targets that distort decisions.
1. Define what quality means for this product
Start with a user need, stakeholder expectation, or product risk—not a metric dashboard. “High quality” could mean accurate calculations for a finance product, accessible workflows for a public service, or dependable performance for a time-sensitive application. The relevant measures depend on what users need and what can go wrong.
ISO/IEC 25010:2023 provides a product quality model with nine characteristics to help specify, measure, and evaluate software quality. Use it as a prompt to check that your requirements cover relevant dimensions. It does not choose your requirements or measures for you; those need to fit your product and stakeholders. See the ISO/IEC 25010:2023 product quality model.
For each quality objective, write down:
- Who cares: the user or stakeholder affected.
- What matters: the behavior or outcome they need.
- What could fail: the risk, including its likely impact.
- What evidence would help: an observable indicator that can inform a decision.
For example: “A customer can complete checkout without losing their basket” is an objective. A team might examine checkout completion, errors during checkout, and support reports about lost baskets. Those are possible indicators, not universal measures; define them against the product’s actual behavior and available data.
2. Measure product quality and delivery performance
Product measures describe outcomes or properties of the software for its users. Delivery measures describe how changes move through the team’s delivery process and how often delivery is disrupted. Together they offer a broader view; delivery speed alone does not establish whether the product meets users’ needs.
| Measurement area | Question it helps answer | Possible evidence | Blind spot to consider |
|---|---|---|---|
| Product quality | Does the software behave as users and stakeholders need? | Task outcomes, observed defects, accessibility findings, reliability signals, or product-specific measures tied to requirements | A measure can miss user groups, contexts, or quality characteristics that were not included in the objective. |
| Delivery performance | How quickly do changes reach users, and how often do deployments need recovery or rework? | DORA’s five software delivery performance measures | Delivery metrics do not fully describe product quality or user value. |
| Team or development experience | What organisational or product-development goal should be understood? | A framework such as SPACE, DevEx, or H.E.A.R.T., chosen to fit the goal | Frameworks answer different questions and should not be treated as interchangeable scorecards. |
DORA currently describes five software delivery performance measures. It groups them into throughput and instability and advises interpreting them in context. DORA says these measures are best suited to one application or service at a time. Consult the DORA software delivery performance metrics guide for the current definitions.
| DORA measure | What it helps describe | Question for the team |
|---|---|---|
| Change lead time | How long a change takes to reach production | Where does work wait between being ready and reaching users? |
| Deployment frequency | How often changes are deployed | Can changes be released in a way that fits the service and its users? |
| Failed deployment recovery time | How long recovery takes after a failed deployment | How quickly can the service recover when a deployment causes a problem? |
| Change fail rate | How often deployments require intervention or result in an incident | What deployment problems are recurring, and what conditions precede them? |
| Deployment rework rate | How much deployment activity is rework | How much delivery effort is being redirected to address deployment-related problems? |
Use DORA’s framework-selection guidance when deciding whether delivery metrics alone answer your question or should sit alongside another framework. The choice should follow the organisation’s goals; frameworks such as SPACE, DevEx, and H.E.A.R.T. address different measurement needs. See DORA’s measurement framework guidance.
3. Turn an objective into a usable measure
A number is useful only when the team understands what it represents and how it was produced. Before adding a measure to a review, document its operational definition.
- Name the indicator. Use a name that describes the observed event or outcome, not an interpretation.
- Define the calculation. State the numerator, denominator, inclusion rules, exclusions, and unit where relevant.
- Identify the data source. Record the system or process that supplies the data and who can check its quality.
- Set the review period. Pick a period that matches the question and produces enough observations to interpret. Explain any known gaps.
- Assign an owner. Name who maintains the definition and brings context to the review.
- Connect it to a decision. Agree what the team might investigate or change if the measure shifts.
For instance, a team tracking production defects should define what counts as a defect, which releases and incidents are included, how duplicates are handled, and how severity is recorded. Without those rules, changes in counting practice can look like changes in software quality.
ISO/IEC 25020:2019 is a separate quality measurement framework for designing and evaluating measurement models. It may help teams formalize a measurement approach; consult the ISO/IEC 25020:2019 listing for its scope.
4. Establish a baseline and review trends
Use an initial period to understand current behavior and data limitations. A baseline describes what the measure looks like under current conditions; it is not automatically a target or evidence that the current state is acceptable.
- Choose the application or service being measured and keep its boundary consistent.
- Check that the data source captures the events your definition requires.
- Record the baseline and any context that could affect interpretation, such as a release, incident, traffic change, or instrumentation change.
- Review the direction and pattern over time instead of reacting to a single point without context.
- Choose one improvement action, name an owner, and decide when to review its effect.
Keep the review focused on learning: “What changed, what evidence could explain it, and what should we investigate?” If a change in the measure has several plausible explanations, gather evidence before attributing it to a team or a particular practice.
5. Avoid misleading targets and single-score thinking
Do not treat one proxy as the definition of quality. A deployment rate, test count, defect count, or user rating can each illuminate part of a problem, but none represents every product-quality characteristic and delivery outcome. The research sources do not establish universal quality scores or target thresholds.
- Pair measures that reveal different effects. Review delivery throughput with instability measures, and connect both to relevant product outcomes.
- Check the incentive before setting a target. Ask what people might do to improve the number and whether that behavior could hurt users, reliability, or maintainability.
- Keep definitions stable enough to compare. If the calculation changes, note when and why; do not silently join incompatible data.
- Use context, not league tables. Compare a service with its own history and operating conditions before comparing it with a different system.
- Treat measures as prompts for investigation. A signal can tell you where to look; it does not necessarily explain the cause.
This is practical measurement advice, not a claim that a particular target always causes a particular behavior. DORA’s guidance to interpret measures in context and to focus them on an application or service supports cautious use rather than universal thresholds.
6. A practical Agile quality-measurement routine
Keep measurement work close to planning and delivery so it can shape improvement rather than becoming a separate reporting exercise.
- During product or backlog refinement: identify the user need, quality characteristic, and risk for the work.
- When defining acceptance: state how the team will recognize the expected outcome, including relevant non-functional requirements.
- During delivery: collect evidence from the agreed sources and note data-quality problems or changes in instrumentation.
- After release: review product outcomes and delivery signals for the relevant service.
- In a team improvement discussion: choose a small, actionable change and revisit the evidence after it has had time to affect the system.
Keep a compact measurement record that the team can revisit:
Objective: Customers can complete checkout without losing their basket.
Product indicator: Checkout completion and reports of lost baskets.
Delivery indicators: DORA measures for the checkout service.
Definition and source: Document calculation, inclusion rules, and data source.
Baseline: Record the period and known data gaps.
Review: Inspect trend and context; agree one follow-up action and owner.
This is a template, not a prescribed metric set. Select measures that support a concrete decision for your product.
7. Capture screenshots as one source of visual evidence
When the quality question concerns a rendered page or user-facing visual change, screenshots can help a team inspect the same state across a release or environment. They are evidence for visual review, not a substitute for user outcomes, accessibility checks, reliability signals, or the other measures the product needs.
A developer can capture a page with a browser automation library, save the image, and compare it in a review workflow. The exact setup depends on the chosen browser and tooling; ensure the capture uses a representative viewport, waits for the page state being evaluated, and avoids exposing private user data or secrets in stored images.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request with a URL returns a PNG, JPEG, WebP, or PDF. For a page-review workflow, the request can provide a consistent capture without setting up browser automation in your project.
See the ScreenshotNeo API documentation for request options. This runnable cURL example saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Cookie banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
8. Performance, reliability, and cost considerations
Measurement has a cost: teams spend time collecting, maintaining, and interpreting evidence. Keep the set small enough to review and useful enough to guide action. Automate collection where practical, but periodically check that the instrumentation still measures what the definition says.
- Performance: Prefer measures that can be collected without adding avoidable work to the production request path. For user-perceived performance, define the relevant user journey and observation conditions rather than relying on an unrelated aggregate.
- Reliability: Interpret incidents, recovery, failed changes, and product behavior with service context. Missing or inconsistent incident and deployment records weaken conclusions.
- Cost: Include the effort to maintain dashboards, instrumentation, and reviews. Add a new measure only when it helps answer a decision the team actually faces.
- Privacy: Limit collected data to what the objective requires. Visual evidence and event records can expose user information; use appropriate access and retention practices.
These are implementation considerations for teams; the cited standards and DORA material do not prescribe a universal instrumentation architecture or cost threshold.
9. Troubleshooting measurement problems
| Problem | Likely cause | What to do |
|---|---|---|
| The metric changes abruptly, but the product seems unchanged. | Instrumentation, event definitions, or data pipelines may have changed. | Check source data, releases, and definition history before interpreting the change. |
| Teams disagree about what counts. | The operational definition leaves edge cases open. | Document inclusions, exclusions, duplicates, and ownership; annotate the date any definition changes. |
| Delivery looks faster while users report more problems. | Delivery measures are being read without product outcomes or instability context. | Review relevant product signals and DORA instability measures together; investigate the service conditions. |
| A dashboard produces many numbers but no action. | Measures are not connected to a decision or review owner. | Remove or defer measures without a clear use; assign an owner and an improvement question to the remaining set. |
| A target improves while workarounds increase. | The target may reward a proxy rather than the outcome. | Revisit the incentive, inspect user and reliability evidence, and treat the metric as a signal rather than a score. |
| Comparisons across teams look unfair. | Services, user needs, definitions, and operating conditions differ. | Use service-level trends and context. Avoid treating different systems as directly comparable without a defensible basis. |
10. Frequently asked questions
How many quality metrics should an Agile team track?
There is no universally correct count. Start with the smallest set that covers the team’s important product objectives and delivery questions, then remove measures that do not inform a decision.
Should quality metrics be used to evaluate individual developers?
The measures described here are intended to help understand product and service outcomes and guide team improvement. They do not, by themselves, establish an individual’s contribution or explain why an outcome occurred.
Are DORA metrics a measure of software quality?
They describe software delivery performance and instability. Use them alongside product-quality measures when the question concerns whether the software meets user and stakeholder needs.
Which framework should a team adopt first?
Start from the decision or goal you need to support. ISO/IEC 25010 can organize product-quality characteristics; DORA measures delivery performance; SPACE, DevEx, and H.E.A.R.T. may fit other organisational or product-development questions. DORA recommends choosing frameworks to fit organisational goals.


