Test Observability: How to Monitor and Debug Automated Tests
Learn how to connect CI test results with traces, logs, commits, and run history so your team can find failures, slow tests, and flaky behavior.
Test observability means collecting test-level outcomes and the execution context needed to explain them, then connecting that information to CI runs, application telemetry, and history. Start by recording each test’s identity, result, duration, error, commit, branch, pipeline run, and environment. When a test fails, link its record to the relevant trace and logs; over multiple runs, use the history to find recurring failures, slowdowns, and flaky behavior.
There is no single required backend. You can instrument tests with OpenTelemetry and export signals to infrastructure you already operate, extend a general observability platform to cover CI, or use a test-focused analytics service. Choose based on your CI provider, test framework, required test-level detail, data policies, maintenance effort, and total cost.
1. What test observability should show
A CI summary that says “tests failed” is a starting point, not a useful diagnosis. A developer should be able to move from a failed pipeline to the specific test, its error and stack trace, the code revision, and the relevant requests or service spans. A team should also be able to compare outcomes and durations across runs.
| Question | Useful information |
|---|---|
| Which test failed? | Test and suite names, framework, result, assertion or error, and stack trace. |
| Where did it fail? | Repository, revision, branch, CI provider, pipeline and job identifiers, and environment where available. |
| What happened during it? | Run duration and related traces, logs, requests, and service spans. |
| Is this new or recurring? | Outcome and duration history, linked to runs and revisions. |
| Is the suite getting slower? | Test-level and suite-level duration trends, plus enough run context to compare changes. |
Capture only context your systems can reliably provide. CI providers and instrumentation differ, and semantic convention support is not universal. OpenTelemetry’s CI/CD conventions are intended to give telemetry a more consistent vocabulary, but check the current specification and your integrations before depending on particular attribute names.
2. Build the data path
- Instrument the test runner and pipeline. Capture one result per test when possible, and record pipeline and job boundaries.
- Preserve correlation fields. Include stable test and suite identity, result, duration, repository revision, branch, run, and environment where available.
- Export telemetry. Send traces, metrics, and logs to a collector or backend your team can operate and query.
- Connect failures to execution context. Make it possible to move from a failed test to the relevant pipeline execution and related telemetry.
- Retain history and review it. Compare test outcomes and duration over time; investigate repeated errors, newly slow tests, and changing failure patterns.
OpenTelemetry describes a vendor-neutral framework for instrumenting, generating, collecting, and exporting traces, metrics, and logs. Its CI/CD semantic conventions provide shared attributes, including a test namespace, to make telemetry more consistently interpretable. These conventions are foundational; do not assume every item is stable or implemented by every CI provider. OpenTelemetry documentation
The OpenTelemetry demo provides one example architecture: a containerized pytest suite queries Jaeger for traces, Prometheus for metrics, and OpenSearch for logs to check whether services emit expected signals. Those products are examples, not required components. OpenTelemetry demo
3. A practical OpenTelemetry example
The following Python example creates a span around a test-like operation and records a failure. It demonstrates the instrumentation shape; it does not, by itself, integrate with a test runner, collect every test result, or configure a production exporter. For a real suite, use the framework’s instrumentation or a test-runner integration where available, and add stable CI identifiers from your environment.
from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode
tracer = trace.get_tracer("ci.tests")
def run_test(test_name, suite_name, commit_sha, branch, run_id):
attributes = {
"test.name": test_name,
"test.suite": suite_name,
"vcs.revision": commit_sha,
"vcs.ref": branch,
"ci.run.id": run_id,
}
with tracer.start_as_current_span("test", attributes=attributes) as span:
try:
# Replace with the actual test operation.
assert 2 + 2 == 4
span.set_attribute("test.result", "passed")
except Exception as exc:
span.set_attribute("test.result", "failed")
span.record_exception(exc)
span.set_status(Status(StatusCode.ERROR, str(exc)))
raise
if __name__ == "__main__":
run_test(
test_name="test_addition",
suite_name="math",
commit_sha="REPLACE_WITH_COMMIT_SHA",
branch="REPLACE_WITH_BRANCH",
run_id="REPLACE_WITH_CI_RUN_ID",
)
Install and configure the OpenTelemetry SDK and an exporter appropriate to your backend before expecting spans to appear there. Keep attribute names aligned with the current semantic conventions and the fields your CI integration actually supplies. Avoid putting secrets, credentials, or sensitive user data in span attributes or logs.
Use the runner’s test identity
Wrap individual tests when you need test-level diagnosis, not just one span around the entire test command. Record suite and test identity consistently. If the test framework or an integration already creates test spans, avoid creating duplicate spans without a clear reason. Add a link or shared correlation identifier between the test span and pipeline run so a failure can be followed in both directions.
Test the instrumentation itself
Telemetry can silently stop being useful if instrumentation changes or exporters are misconfigured. OpenTelemetry’s Java SDK testing utilities include in-memory exporters and readers, plus JUnit extensions for inspecting emitted spans, metrics, and logs without sending them to a backend. This checks instrumentation behavior; it is distinct from monitoring an entire test suite in CI. OpenTelemetry Java testing documentation
4. Find failures, slow tests, and flaky behavior
Debug one failed test
- Open the failing test record and confirm its test and suite identity.
- Read the assertion or error and stack trace; note the duration and environment.
- Confirm the repository revision, branch, CI run, and job so you can reproduce the same code and execution conditions.
- Follow links or shared identifiers to relevant traces and logs. Look for failed requests, unexpected service behavior, or missing telemetry around the test’s time window.
- Compare the failure with prior runs and nearby changes. A single failure can be a regression, an environmental issue, or nondeterministic behavior; the record alone does not prove which.
Useful failure context includes the assertion or error, stack trace, test and suite identity, run duration, revision and branch, and related requests or service spans. Elastic describes tracing pipeline executions and drilling into build errors and details. Datadog’s product documentation describes test errors and stack traces alongside branch, commit, and author information. These are vendor-described capabilities, not independent performance findings. Elastic Observability · Datadog CI Test Visibility documentation
Find slowdowns
Track duration at both test and suite level. Compare repeated runs and connect changes to revisions or pipeline configuration changes. A suite getting slower may reflect one newly expensive test, broader environment contention, or a changed dependency; inspect the underlying test and run context before treating a duration increase as a code regression. Use history to identify candidates for investigation rather than assuming a single run establishes a trend.
Investigate flaky tests
A flaky test can pass on one run and fail on another even when the code under test has not changed. Preserve repeated outcomes and their run context, then investigate nondeterministic dependencies and environmental conditions. A rerun can show that behavior varies, but it does not identify the cause or repair the test.
A 2022 multivocal review examined 651 items: 560 academic articles and 91 grey-literature articles. That is the size and composition of the review corpus, not an industry flaky-test rate. The review also summarizes estimates from earlier work, but those figures have distinct populations and dates and should not be treated as current universal prevalence. 2022 multivocal review of flaky tests
5. Choose an implementation approach
| Approach | Good fit when | Questions to check |
|---|---|---|
| OpenTelemetry with an existing backend | You want vendor-neutral instrumentation and can reuse production observability skills and infrastructure. | How much instrumentation and collector operation is needed? Is test context consistent? What will telemetry volume and retention cost? |
| General observability platform extended to CI/CD | You want pipeline traces, dashboards, alerts, and performance views alongside service telemetry. | Which CI systems are supported? How much is automatic? Does the view reach the individual test case detail developers need? |
| Test-focused analytics or visibility service | You prioritize test history, test-level errors, flakiness analysis, or suite exploration. | Does it support your framework and workflow? What are the data handling, retention, access-control, plan, and cost terms? |
Elastic documents pipeline summaries with duration and failure-rate history, as well as CI pipeline tracing and a pytest plugin example. Currents describes test execution history, flakiness, regression analytics, and suite exploration. Datadog describes test errors and stack traces with commit context. These vendor materials help identify capabilities to evaluate; they do not establish a neutral ranking or current prices. Verify current integrations, plan limits, retention, and terms with each provider. Elastic CI/CD pipeline monitoring · Currents documentation · Datadog CI Test Visibility documentation
Compare candidates against your actual stack using these criteria:
- CI provider and test-framework coverage.
- Test-level failure context and trace/log correlation.
- History, flaky-test analysis, and duration or bottleneck views.
- Alerts and fit with the developer workflow.
- Setup and ongoing maintenance.
- Data residency, retention, access controls, and sensitive-data handling.
- Total cost, including telemetry volume, storage, and any required infrastructure.
6. Keep the system reliable and economical
Reliability
- Keep test execution useful if telemetry export is unavailable. Decide whether export errors should fail CI; for many teams, telemetry delivery should not hide the test’s actual result.
- Preserve a local or CI-accessible test report so a backend outage does not erase the immediate failure details.
- Use consistent identifiers across test records, pipeline runs, traces, and logs. Missing or unstable correlation fields make cross-linking unreliable.
- Check instrumentation with focused tests, and alert on missing telemetry only when the signal is meaningful for your setup.
- Control retries carefully. A retry can provide evidence about nondeterminism, but report the initial outcome and retry outcome separately.
Performance and cost
Instrumentation and export add work to CI. Measure the impact in your own pipeline, especially for large suites. Keep attributes useful and bounded; avoid attaching large payloads or sensitive data. Sampling can reduce trace volume, but may omit the exact failing execution, so choose a policy that preserves diagnostic value for test failures. Retention, ingestion, and query costs vary by backend and configuration; the research materials do not establish comparable current prices. Estimate cost from your expected run volume, telemetry volume, retention period, and provider terms.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No test spans appear | The SDK or exporter is not initialized, the test path is not instrumented, or export is misconfigured. | Confirm initialization occurs in the CI process, verify the exporter endpoint and credentials, and inspect local exporter diagnostics. Add an instrumentation test where practical. |
| Spans appear but have no test details | The instrumentation wraps only the whole test command, or test and CI attributes are not populated. | Instrument at test granularity and pass stable test, suite, revision, branch, and run fields from the runner and CI environment. |
| A test cannot be connected to its pipeline | Correlation identifiers are absent, inconsistent, or differ between systems. | Define a small set of shared identifiers and populate them consistently in test telemetry and pipeline records. |
| Trace or log links are empty | Related services do not emit telemetry, timestamps or context do not align, or trace/log correlation is not configured. | Check service instrumentation, clocks and time windows, and the backend’s correlation setup. Treat missing telemetry as missing evidence, not proof that nothing happened. |
| Failures appear as flaky but are hard to reproduce | Run context is incomplete or environmental dependencies vary between attempts. | Retain environment and run identifiers, compare failure and pass records, and investigate external dependencies, concurrency, and timing assumptions. |
| Telemetry makes CI slow or expensive | Excessive volume, payload size, or export work; retention may also be too long for the diagnostic need. | Reduce unnecessary attributes and payloads, review sampling and retention, and measure export overhead. Preserve enough failure context for diagnosis. |
| Instrumentation tests pass but CI data is missing | In-memory checks validate emitted signals but not the CI exporter, credentials, network path, or backend ingestion. | Keep the unit-level instrumentation check and add a small end-to-end pipeline verification for delivery and correlation. |
8. Or skip the browser setup
When the test or CI investigation also needs a screenshot of a page, you can capture it with one GET request instead of installing and operating a browser capture stack. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers say which page verdict applied and whether the request was billed.
- An MCP server lets AI agents, including Claude and Cursor, use screenshot tools.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Sign up for 1,000 free screenshots a month, with no card required.
9. Frequently asked questions
Is test observability the same as test reporting?
Test reporting summarizes outcomes. Test observability connects those outcomes to execution context and related telemetry, then retains enough history to investigate patterns.
Do I need OpenTelemetry to monitor CI tests?
No. It is one vendor-neutral option. A platform or test analytics service may provide integrations that fit your framework and CI workflow better.
Does rerunning a failed test prove it is flaky?
A different result on a rerun is evidence of variable behavior. It does not establish the cause, and one matching result does not rule out flakiness.
What should we instrument first?
Start with test identity, result, duration, error details, revision, branch, and CI run identifiers. Add trace and log correlation where it helps answer concrete debugging questions.


