How to Monitor Automated Tests
Build a reliable CI feedback loop with per-test results, runtime trends, actionable alerts, and a process for investigating flaky tests.
To monitor automated tests, run fast, relevant checks on each code change; retain machine-readable results, durations, logs, and useful failure artifacts; review individual test and suite trends; and route actionable failures to an owner. Put slower integration, full-suite, and performance checks later in the pipeline or on a schedule when they do not belong on every commit. Track flaky tests explicitly and investigate their causes instead of treating retries or quarantine as fixes.
1. Choose when each test suite runs
Start with a fast set of tests that gives useful feedback on every relevant change. Add integration and regression tests in later pipeline stages, and schedule long-running suites where that fits your release risk and runtime. A broad test pack that makes feedback too slow can lead people to stop running or trusting it. HMRC’s engineering standard recommends running automated tests regularly and ideally on every change. Microsoft’s testing guidance likewise recommends choosing test practices in light of the workload and risk.
| Suite | Typical trigger | Monitoring focus |
|---|---|---|
| Unit and fast component tests | Every relevant commit or pull request | Failures, duration changes, and whether tests are isolated |
| Integration and service tests | Pull request, merge, or a later pipeline stage | Dependency health, environment failures, and repeated failure clusters |
| Full regression suite | Scheduled run, release candidate, or later stage | Coverage of important risk areas, full runtime, and regressions |
| Load or performance suite | Scheduled run or pre-production stage | Changes against a consistent baseline and test environment |
Adjust this mix to your release process, risk, dependencies, and how long each suite takes. Run checks often enough to detect regressions, but keep the fast feedback path usable.
2. Keep per-test results and failure evidence
A green or red job summary cannot tell you which test changed, whether it is getting slower, or whether a failure came from code, test state, or the environment. Preserve test-level results and run history. Where your tooling supports them, retain the commit, branch, test identifier, environment, error, stack trace, and reproduction details. For UI failures, screenshots or video can help show what the test saw.
Use a machine-readable report format supported by your framework and CI service, and make failed-test output easy to open from the run. For example, pytest can write JUnit XML and GitHub Actions can retain that report as an artifact:
name: Tests
on:
pull_request:
push:
branches: [main]
jobs:
unit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install -r requirements.txt
- name: Run tests and write a machine-readable report
run: python -m pytest --junitxml=test-results/junit.xml
- name: Retain test report
if: always()
uses: actions/upload-artifact@v4
with:
name: junit-test-results
path: test-results/junit.xml
if-no-files-found: ignore
This workflow assumes the repository has a requirements.txt that installs pytest and its test dependencies. The if: always() condition keeps the artifact upload step eligible after test failure; if-no-files-found: ignore avoids a second error if the report was never created. Configure your CI platform’s native test-result ingestion if you want its test-level view and analytics; artifact storage alone is not the same as an analytics dashboard.
3. Watch a small set of useful signals
- Outcomes: job-level and individual-test pass/fail history.
- Duration: per-test and per-stage timings, including changes over time.
- Flakiness: tests that fail intermittently or have a persistently low success rate.
- Failure clusters: repeated errors that may share an environment, dependency, or code cause.
- Artifacts and context: logs, stack traces, screenshots where useful, and the run’s commit and environment.
- Coverage: evidence about untested paths, especially important risk areas.
- Ownership and follow-up: severity, assigned owner, and age of tracked failures.
Use coverage to find important gaps, not as a score to maximize without regard to assertion quality or maintenance cost. A stable suite can still miss critical behavior. Likewise, a pass-rate number can hide irrelevant tests or recurring failures in a small but high-risk area. Review trends and connect investigations to user and business risk.
4. Make failure notifications actionable
Send alerts to the team or owner able to investigate the affected code. Include the test name, commit or build link, environment, error and stack trace, available artifacts, and a direct link to the run. Consider routing repeated or high-impact patterns differently from one-off failures, while preserving the individual results that explain the pattern. A notification that says only “pipeline failed” adds work instead of helping diagnosis.
Agree on who triages new failures and how a failure becomes a tracked defect. Record the first failing run and whether a retry passed. That history helps distinguish a persistent regression from intermittent behavior without erasing the original evidence.
5. Diagnose flaky tests instead of normalizing them
A flaky test passes and fails against the same code under conditions that should be equivalent. Common causes include uncontrolled system state, shared or uncleared test data, test-order assumptions, parallel interference, thread-safety issues, unstable infrastructure, and timing-sensitive assertions. pytest’s guidance on flaky tests discusses these failure mechanisms.
- Confirm the pattern using test-level history, including first-attempt results and retries.
- Compare the failing and passing runs: commit, environment, order, parallel worker, logs, and dependencies.
- Check for leaked state, shared data, reliance on execution order, race conditions, and tight timing thresholds.
- Fix the underlying test, application, or environment cause where possible.
- If temporary quarantine is necessary, keep the test visible, assign an owner, and set a review or expiry point.
Retries can help reveal intermittent behavior, but a retry that turns a job green can conceal the initial failure. Keep retry counts and first-attempt outcomes in reports when your platform permits it. CircleCI documents that auto-reruns can suppress an earlier failure when a later attempt passes; recurring failures should remain available for investigation. Quarantine also needs follow-up: permanent exclusion can allow regressions through.
6. Add visual evidence when UI tests fail
For browser-based tests, a screenshot can preserve the page state associated with a failure. Capture at the point of failure and attach the result to the same run alongside the test name, commit, and error. Keep the capture useful: the relevant viewport, page state, and environment matter more than collecting images without a question to answer.
For a page your test is authorized to access, this minimal Playwright Python example captures the rendered page after navigation. Install Playwright and its Chromium browser in the job before running the script:
from pathlib import Path
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page(viewport={"width": 1440, "height": 900})
try:
response = page.goto("https://example.com", wait_until="networkidle", timeout=30_000)
if response is None or not response.ok:
status = response.status if response else "no response"
raise RuntimeError(f"Navigation failed: {status}")
page.screenshot(path="artifacts/page.png", full_page=True)
finally:
browser.close()
This is a basic evidence-capture example, not a test assertion or a full visual-regression system. In a test suite, put capture logic in a failure hook or fixture so it runs when a test fails, and upload the artifact even when the test step fails. Create the output directory before capture if it may not exist, and avoid network-idle waits on pages that continuously poll or stream; a specific readiness selector or bounded wait can be more reliable.
7. Decide when you need additional monitoring tools
First check whether your framework and CI platform already retain individual results, timing, history, artifacts, and notifications. Add a dedicated test analytics service when you need cross-project aggregation, longer history, flaky-test triage, or scheduled API checks beyond the built-in workflow. Compare report-format support, test-level history, artifact retention, parallel-run analysis, alert routing, setup effort, and plan limits.
CircleCI documents CI test-result storage, failed-test output, run timing, and insights into flaky, slow, and low-success tests. Postman Monitors are a fit for scheduled or CLI-triggered API checks with run history and notifications; its documentation describes plan and feature limitations, so verify current terms for your use case. These are examples of documented workflows, not a universal ranking of vendors.
For API monitors, also compare schedule frequency, execution region, private-network access, and authentication support. For a practical triage process, GitLab’s handbook describes internal automation that analyzes failure data, identifies high-impact flaky files, creates issues, and routes them to owners. Treat that as one operational example rather than a required threshold model.
8. Troubleshoot common monitoring failures
| Symptom | Likely cause | What to do |
|---|---|---|
| The job is red, but there is no test-level detail | The report was not generated, retained, or ingested by the CI test-results feature. | Check the report path and test command; retain reports on failed runs and configure native ingestion if available. |
| The report artifact is missing after a failure | The upload step ran only on success, the path differs, or the test process stopped before writing it. | Run upload steps with an always-run condition, verify the path, and inspect whether the framework wrote a partial report. |
| A test fails only in CI or only in parallel | Environment differences, shared state, order dependence, races, or resource contention. | Compare environment and worker details, isolate test data, and reproduce with controlled parallelism. |
| Retries turn failures green and the team stops investigating | The dashboard emphasizes the final attempt and hides the initial failure. | Preserve attempt-level outcomes and retry counts; treat repeated first-attempt failures as work to triage. |
| Alerts are frequent but lack useful context | Notifications omit test identity, commit, error details, artifacts, or ownership. | Add those fields and route alerts to the responsible team; tune routing without deleting the underlying results. |
| Runtime grows while the suite remains green | Slow tests or pipeline stages are hidden by aggregate status. | Track per-test and per-stage durations over time; investigate regressions before they make feedback unusable. |
| Browser capture times out waiting for network idle | The page keeps making requests or never becomes idle. | Wait for a meaningful selector or use a bounded, appropriate readiness condition instead. |
9. Keep the monitoring system reliable and affordable
- Keep the commit path fast: run the smallest high-value checks on each change and place slower suites at a later stage or schedule.
- Retain what supports diagnosis: preserve reports and failure artifacts long enough for your team’s triage cycle; set retention to match storage and investigation needs.
- Control artifact volume: capture screenshots or video where they answer a debugging question, and avoid retaining redundant large artifacts indefinitely.
- Make failures reproducible: record environment and run context so infrastructure noise can be separated from product regressions.
- Review alert noise: route by ownership and impact, but keep source results accessible so filtering does not erase evidence.
- Measure feedback cost: monitor total pipeline time and scheduled-run usage; optimize the suite based on duration and risk rather than removing valuable assertions blindly.
There is no single useful test cadence or retention period for every team. Choose them from the cost of a missed regression, test runtime, release process, and the time needed to investigate failures.
Or skip the browser setup
For standalone page captures used as visual evidence or scheduled page checks, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF; see the API documentation for options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Replace the example URL with a page you are authorized to capture and keep the API key out of source control. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
FAQ
Should every test run on every commit?
Run the fast, relevant checks on each change. Put slower or broader suites in later stages or scheduled runs when their runtime and risk make that a better feedback tradeoff.
Is test coverage a good health metric?
It is useful evidence for finding untested paths, especially in high-risk areas. It does not show whether assertions check important behavior, so do not optimize coverage as an isolated target.
Should a flaky test be quarantined?
Only as a temporary, visible measure with an owner and review point. Preserve its failures and investigate the cause so quarantine does not become permanent silence.
Do I need a separate test observability service?
Not necessarily. Start with framework and CI features that retain test-level results, timing, history, artifacts, and notifications. Add a separate service when those features no longer meet your aggregation, history, or triage needs.


