ScreenshotNeo

BlogEngineering

How to Implement Test Observability to Improve Software Quality

Connect test outcomes to logs, metrics, and traces so CI failures are easier to diagnose and telemetry is trustworthy.

By the ScreenshotNeo team4 October 202610 min read

Test observability connects a test’s pass or fail result to what the application did and what telemetry it emitted during that run. Implement it by defining the questions your team needs answered, instrumenting the test and application boundaries, correlating test runs with telemetry, asserting signals locally, validating the full delivery path, and tracking failures over time.

A passing assertion alone does not show why a request failed across services, whether a dependency was called, or whether telemetry reached its backend. A useful test can verify both the application result and the trace produced by the operation. OpenTelemetry’s demo illustrates checks across traces, metrics, and logs; its trace-based tests check an operation’s result and its emitted trace.

1. Decide what a test must help you learn

Start with concrete questions, rather than collecting telemetry without a debugging purpose. For a critical request or workflow, ask:

  • Which test, run, service, and operation failed?
  • Which services did the operation call, and where did it spend time?
  • Did the application emit the expected error or other log record?
  • Did expected metrics and traces reach the configured backends?
  • Does the same test produce different outcomes with unchanged code?

These questions imply different evidence. Logs provide detailed context such as errors and stack traces. Traces show how services interact during an operation. Metrics help surface abnormal behavior over time. Use all three when they answer distinct questions; do not require every signal for every test. Google Cloud’s OpenTelemetry guidance describes OpenTelemetry as a vendor-neutral way to collect application telemetry and send it to a destination.

2. Instrument the application and test boundaries

Instrument the path the test exercises, including relevant service boundaries. Use the instrumentation supported by your actual language and framework. Ensure trace context propagates through the system under test; otherwise, a request may appear as unrelated spans or have no trace that the test can retrieve.

Keep a stable test identity and a distinct run identity. A practical correlation record includes the test name, run or CI job ID, build revision, environment, and trace ID when available. Pass trace context through the same request path the test is validating. Avoid placing secrets or sensitive personal data in telemetry attributes.

Choose attributes that help narrow a search, such as service name, operation, test identity, and environment. Keep high-cardinality values such as unique run IDs out of metric labels when they would create an excessive number of time series; attach them to traces or logs where your backend supports that use.

3. Correlate the test result with its telemetry

Make the relationship explicit: the test triggers an operation, records or receives the relevant trace context, checks the operation’s output, and then checks the telemetry associated with that operation. This follows the pattern in OpenTelemetry’s trace-based test example.

Do not rely only on searching a shared backend by a broad service name or approximate timestamp. Concurrent CI jobs can overlap. Prefer a run-specific correlation value or trace ID, and include it in assertion failures. If the test cannot obtain the trace ID directly, use a unique test-run attribute and a bounded time window as a fallback; keep the query narrow enough to avoid matching another run.

4. Add focused local telemetry assertions

Use in-memory exporters or readers for focused tests of instrumentation code. These tests can assert that an expected span, metric, or log record is created without starting a collector or backend. OpenTelemetry documents Java SDK testing utilities and in-memory assertions in its Java SDK documentation.

A framework-neutral test shape looks like this:

// Pseudocode: adapt exporter and query APIs to your language and SDK.
function test_checkout_emits_telemetry() {
  const recorder = createInMemoryTelemetryRecorder();
  const app = createApp({ telemetryExporter: recorder });

  const response = app.handle(checkoutRequest({
    testRunId: "checkout-test-42"
  }));

  assertEqual(response.status, 200);
  assertSpanExists(recorder, {
    name: "checkout",
    attributes: { "test.run_id": "checkout-test-42" }
  });
  assertMetricRecorded(recorder, "checkout.requests", { valueAtLeast: 1 });
}

This is deliberately pseudocode, not a runnable SDK-specific example: exporter setup, metric readers, and log assertions differ by language and SDK version. Use your framework’s actual in-memory testing API. Keep assertions focused on the contract you need, such as an operation span and a request count, rather than pinning every internal span name or attribute. Overly detailed assertions can make harmless instrumentation changes break product tests.

5. Validate the complete telemetry path

In-memory tests do not exercise exporter configuration, collector routing, credentials, network delivery, backend ingestion, or query visibility. Add a small telemetry sanity suite that runs against the real telemetry path in a representative test environment.

  1. Trigger one known operation for each component or service in scope.
  2. Declare which signals that component is expected to emit: logs, metrics, traces, or a specific subset.
  3. Query the configured backends using the run ID, trace ID, or another narrow correlation key.
  4. Assert that expected signals are visible and that the trace contains the required service boundaries.
  5. Set bounded waits and actionable timeouts for eventual ingestion; report the signal and query that were missing.

The OpenTelemetry demo uses separate trace, metric, and log backends and defines expected signals per service. That is a useful model: verify what each component should deliver, not merely that the test process exited successfully.

Run this suite at a cadence that fits the environment. A fast local or per-change in-memory suite catches instrumentation regressions quickly. A smaller backend-path suite can run in CI or on a schedule to catch broken exports, routing, credentials, and backend visibility. Avoid making every unit test depend on a shared backend; backend availability can obscure whether the product behavior itself is correct.

6. Make failures actionable

A telemetry assertion should tell the engineer what was expected, what was found, and how to locate the related run. OpenTelemetry’s testing guidance says: “When a test fails, the output should make it obvious what was being checked and show a clear diff between actual and expected values, without long hand-written messages.” See its testing guidance.

Expected trace for test=checkout_test run=checkout-test-42
  operation: checkout
  required services: web, payment
  observed services: web
Missing: payment span
Trace query: test.run_id="checkout-test-42"
Build: 8f31a2c

Keep failure output concise. Include a link or query to the trace or backend result when your environment can generate one. Do not dump credentials, full request headers, or sensitive payloads into CI logs.

7. Track test quality and investigate flakiness

Store test outcomes with enough history to compare repeated runs of the same test and code. Flakiness means the same code can produce both passing and failing outcomes. A failure that varies across retries or runs may point to timing, shared state, external dependencies, or infrastructure, though a flaky symptom can still reveal a real product race.

Useful team-defined indicators include:

  • Test duration and changes in its duration over time.
  • Failure rate by test, service, and environment.
  • Pass/fail variation for unchanged code.
  • Missing expected telemetry, separated by signal and component.
  • Time needed to find the relevant trace or error context.

These are suggestions for operational tracking, not a universal standard or published benchmark. Google’s historical article reported that about 1.5% of test runs in its own corpus had a flaky result and almost 16% of its tests had some level of flakiness; it also observed that about 84% of pass-to-fail transitions in its post-submit system involved a flaky test. Those 2016 figures describe Google’s systems and should not be treated as current industry rates. See Google’s article on flaky tests.

If you quarantine a flaky test to unblock the critical path, record an owner, reason, review date, and repair plan. Google’s article notes that quarantine can remove a test from the critical path but may hide a race or another real bug. Monitor the quarantined test and return it to the normal suite after addressing its instability.

8. Account for privacy, retention, and operating cost

Telemetry has volume and access implications. Decide which attributes are safe to collect, who can query them, and how long each signal needs to be retained. Avoid logging secrets and unnecessary personal data. Sample high-volume traces where appropriate, but make sure sampling does not silently remove the evidence your telemetry sanity checks require. Keep a deterministic unsampled path or a test-specific sampling policy if those checks depend on complete traces.

Measure ingestion and query costs in your own environment; the cited sources do not establish universal cost figures. Keep backend-path checks small, avoid querying broad time ranges, and use retention appropriate to how long engineers need to investigate failures.

9. Troubleshoot common failures

Symptom Likely cause What to check or change
The application test passes, but no trace is found. Instrumentation is disabled, trace context did not propagate, export failed, or the backend query is too narrow. First assert the span with an in-memory exporter. Then inspect exporter and collector errors, confirm the trace ID or run ID, and widen the time window slightly.
Spans exist, but a downstream service is missing. The downstream call was not instrumented, context was lost at the boundary, or the test path did not invoke that dependency. Verify the dependency was called, check context propagation across the client/server boundary, and assert the expected child span locally.
Telemetry appears intermittently in CI. Asynchronous export or backend ingestion has not completed before the query, or shared CI load affects timing. Use bounded polling with a clear deadline, inspect exporter queue/drop errors, and keep the test query scoped to its unique run.
A metric assertion finds duplicate or unrelated data. The query matches other concurrent tests, or a cumulative reader is being interpreted as a per-test value. Use a test-specific dimension where appropriate, reset or isolate the reader, and assert a delta or scoped measurement.
The test is flaky only when telemetry checks are enabled. The assertion depends on wall-clock timing, ordering, eventual ingestion, or an unstable internal span detail. Wait for the specific signal with a bounded timeout, assert stable instrumentation contracts, and report actual versus expected data.
Backend checks fail while local in-memory checks pass. Exporter configuration, collector routing, credentials, network access, or backend indexing is broken. Inspect each hop in the export path and run a small known operation through the same environment and credentials.
Queries return too much data or become slow. Correlation keys are missing, time ranges are broad, or high-cardinality labels have increased backend load. Add a run or trace correlation key, narrow the time window, and move unique identifiers out of metric labels.
Quarantined tests stay ignored. No owner or review date was assigned, or the underlying cause is difficult to reproduce. Assign an owner and deadline, preserve run history and telemetry, and track the repair as work with an explicit completion condition.

10. A practical rollout checklist

  1. Pick one high-value user operation and list the questions its failures should answer.
  2. Confirm the application and test framework can emit and capture the required signals.
  3. Propagate a run identity and trace context through the tested operation.
  4. Add focused in-memory assertions for stable instrumentation contracts.
  5. Add a small backend-path check for expected signals and services.
  6. Make failures show the test identity, missing expectation, and a way to find telemetry.
  7. Record test history and review flaky failures; give quarantined tests owners and dates.
  8. Review telemetry volume, retention, access, and sensitive attributes.

Or skip the browser setup

If a test or CI diagnostic needs a screenshot of a page, a browser setup is another dependency to install and maintain. ScreenshotNeo provides a one-request website screenshot API. See the ScreenshotNeo API documentation for its parameters and formats.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. ScreenshotNeo is made by Yorker Media. Sign up for 1,000 free screenshots a month, with no card.

FAQ

Does every test need to query a telemetry backend?

No. Use in-memory checks for fast instrumentation feedback and reserve backend queries for checks that need to validate export, routing, ingestion, or visibility.

Should telemetry assertions check exact span trees?

Only when the tree structure is part of the behavior you need to protect. Prefer stable contracts, such as a required service call or operation span, so internal instrumentation changes do not create needless failures.

How do I distinguish a product regression from a flaky test?

Compare repeated outcomes for the same code and inspect correlated traces, logs, environment, and dependency behavior. A changing result suggests flakiness, but investigate the underlying cause because a race can be a real product defect.

Can screenshots replace logs, metrics, or traces?

No. Screenshots can preserve visual state for a page, while logs, metrics, and traces explain application events and service behavior. Use screenshots as supporting evidence when the rendered page itself matters.