ScreenshotNeo

BlogEngineering

Test Observability: What It Is and How to Use It

Test observability adds traces, metrics, and logs to pass/fail results so teams can understand test behavior and diagnose unexpected outcomes.

By the ScreenshotNeo team4 October 20268 min read

Test observability is the practice of collecting and using telemetry from test runs and the systems they exercise. It gives a team evidence to understand behavior, troubleshoot unexpected outcomes, and check relevant execution details beyond a binary pass or fail.

The core signals are traces, metrics, and logs. OpenTelemetry provides vendor-neutral APIs, SDKs, instrumentation, and collection components to generate, collect, and export those signals. It does not provide the storage and visualization backend; teams choose that separately.

1. What test observability means

A test result tells you whether an assertion passed. Telemetry helps explain what happened while the test ran: which operation was executed, which services handled it, how long steps took, and what errors or measurements were emitted. Observability is useful when it lets developers ask questions about system behavior without needing to predict every failure in advance.

Signal What it helps answer Typical test use
Traces What path did this operation take, including across service boundaries? Find where a multi-service workflow slowed down or diverged.
Metrics How much, how often, or how long? Inspect counts, durations, or other measurements associated with a run.
Logs What events or details were recorded? Review errors and contextual events around a failure.

Observability depends on useful emitted data. If the application or test environment does not expose relevant telemetry, a backend cannot reconstruct the internal behavior from a pass/fail result alone.

2. When to use it

  • Distributed flows: A test crosses service boundaries, and the final result does not identify which part behaved unexpectedly.
  • Intermittent failures: A failure is hard to reproduce locally, so evidence from the original run matters.
  • Unexpected latency or measurements: A test needs to inspect duration or emitted measurements as part of diagnosis.
  • Execution-path requirements: A test should validate that the intended services or operations participated, in addition to validating the output.

For a small, deterministic unit test, ordinary assertions may be enough. Add telemetry where it can answer a real failure question; instrumentation without a diagnostic purpose adds maintenance work.

3. A practical workflow

  1. Write down the questions. For a failing run, should a developer be able to tell which operation ran, where it went, which dependency responded unexpectedly, or which measurement changed?
  2. Instrument the code path. Use code-based instrumentation when application-specific details or custom signals are needed. Use zero-code instrumentation to get started or when changing application code is impractical. These approaches can complement each other.
  3. Keep the test run correlated with its telemetry. For distributed work, preserve the trace associated with the tested operation so the test result can lead to the relevant request path.
  4. Choose where to inspect it. Use in-memory capture for self-contained assertions when the language SDK supports it. Export to a backend when teams need broader inspection, storage, or visualization.
  5. Assert behavior and relevant evidence. Check the expected result and the trace evidence that matters. Avoid making incidental span names or internal details the only contract unless those details are intentionally part of the test.

4. Trace-based testing

A distributed trace records the path of an operation across services. Trace-based testing runs an operation, captures its trace, and validates the trace alongside the operation’s output. The OpenTelemetry Demo trace-based testing guide describes this pattern for a shopping flow that uses multiple services.

A useful trace assertion checks a meaningful property of the execution, such as whether an expected operation or service appears. Pair it with the normal output assertion: a trace can show that a path ran, but does not by itself prove that the user-visible result is correct.

Keep assertions stable

  • Prefer expected operations, service participation, or meaningful error status over incidental internal implementation details.
  • Be deliberate about which parts of a trace are contractual. Internal service changes can legitimately alter spans.
  • When a trace crosses services, ensure context is propagated so spans belong to the same operation.
  • Retain enough run context to find the trace from the test result, while avoiding sensitive values in telemetry.

5. Choose an instrumentation and inspection approach

Approach Useful when Tradeoffs
Code-based instrumentation You need application-level detail or precise custom signals. Requires code and instrumentation maintenance; offers control over detail.
Zero-code instrumentation You need a quick start or cannot change the application. Setup constraints and available application-specific context may limit what you can inspect.
In-memory telemetry assertions A test should validate emitted telemetry without sending it to a backend. Support and utilities vary by language SDK and version.
Backend-centered trace analysis You need to inspect telemetry beyond one test process or across services. Requires backend integration, storage, visualization, and correlation setup.

These patterns can be combined. OpenTelemetry is designed to work with different backends, including open-source and commercial offerings. Its APIs and SDKs create or instrument telemetry, and its Collector can receive, process, and export it; storage and visualization are provided separately.

6. Example: assert telemetry in a Java test

The OpenTelemetry Java SDK testing artifact documents in-memory exporters and assertion utilities so tests can inspect telemetry without exporting it to a backend. The exact setup and APIs depend on the Java SDK version. Check the current official Java documentation and testing artifact instructions before copying a version-specific dependency or API call into a project.

A version-neutral test outline is:

// Pseudocode: use the in-memory exporter and assertion utilities
// documented for the OpenTelemetry Java SDK version in your project.

@Test
void checkoutEmitsExpectedTelemetry() {
    var exporter = createInMemoryExporter();
    var telemetry = createTestTelemetry(exporter);

    var result = checkout(telemetry);  // exercise the operation

    assertEquals(EXPECTED_RESULT, result);
    var spans = exporter.getFinishedSpanItems();
    assertTrue(spansContainExpectedCheckoutOperation(spans));
}

The helper names above describe the test structure; they are placeholders, not Java SDK APIs. Use the actual exporter, SDK setup, and assertion utilities documented for the project’s language and SDK version. The cited in-memory testing guidance is specific to Java and should not be assumed to apply unchanged to other languages.

7. Screenshot evidence for browser-driven tests

For a browser test, a screenshot can complement traces, metrics, and logs by showing the rendered page at the point of failure. A screenshot is visual evidence, not telemetry: it will not explain a distributed request path or replace instrumentation. Capture it at a deliberate checkpoint, such as after navigation or when an assertion fails, and associate it with the same test run.

For a do-it-yourself capture, use the browser automation framework already running your test. For example, with Playwright and Node.js:

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();
try {
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  await page.screenshot({ path: 'test-failure.png', fullPage: true });
} finally {
  await browser.close();
}

Install Playwright and its browser runtime using the project’s standard setup. Replace the sample URL with the page under test, and capture only the relevant state if full-page images are too large. Keep credentials and sensitive page content out of artifacts and apply the test system’s access and retention rules.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A single request returns an image or PDF; its API parameters also support names used by other screenshot APIs. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie banners are accepted like a visitor would accept them, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

9. Performance, reliability, and cost

Performance

  • Emit and retain signals that answer the failure questions; unnecessary high-volume detail increases processing and storage needs.
  • In-memory assertions can keep a self-contained telemetry check local to a test, where supported by the language SDK.
  • Exporting telemetry to a backend adds integration and network considerations. Keep test assertions focused so diagnostic infrastructure does not obscure the product behavior under test.
  • For browser screenshots, choose a wait condition that matches the page and capture point. Waiting for every network connection to become idle can be unsuitable on pages with ongoing requests; a specific selector or page event may be more appropriate.

Reliability

  • Make sure trace context reaches the services involved, or a distributed operation may appear as disconnected traces.
  • Decide what the test should do if telemetry export or the backend is unavailable. A product behavior test can often report telemetry separately rather than fail solely because an external analysis system is down; a telemetry pipeline test should assert its own delivery behavior.
  • Keep a path from the test identifier to the emitted trace and artifacts. Without correlation, telemetry may exist but be difficult to locate.
  • Do not treat screenshot evidence as a substitute for logs, spans, or assertions. Each reveals a different part of the failure.

Cost

OpenTelemetry is a framework and toolkit, not the storage and visualization service. Estimate cost in the chosen backend based on the telemetry retained and the team’s needs; the research sources do not provide a universal price or cost benchmark. In-memory test assertions avoid sending that test telemetry to a backend, but only cover the local checks they support.

10. Troubleshooting

Symptom Likely cause What to check
No spans or measurements appear The code path is not instrumented, the SDK is not configured, or the exporter is not connected. Confirm instrumentation and SDK setup, then verify that the test executes the instrumented operation.
Trace ends at a service boundary Trace context is not being propagated, or the receiving service is not instrumented. Inspect context propagation and instrumentation on both sides of the request.
Telemetry assertion fails intermittently The test reads before spans finish, or concurrent activity makes the observed data nondeterministic. Wait for the operation and relevant spans to finish; isolate the test’s telemetry scope where possible.
Telemetry is visible locally but missing in the backend Export configuration, Collector processing, or backend integration may be incorrect. Check the exporter destination and Collector path; OpenTelemetry itself is not the backend.
Assertions break after an internal refactor The test may depend on incidental span names or implementation structure. Assert stable behavior and intentional execution requirements instead.
Browser screenshot is blank or incomplete Capture ran before the relevant page state rendered, navigation failed, or the target needs a later checkpoint. Check navigation status and wait for a relevant selector or page event before capture.

11. Frequently asked questions

Is test observability the same as test reporting?

No. A report summarizes test outcomes; observability adds signals that help investigate how the tested system behaved.

Does OpenTelemetry include a dashboard?

No. It provides instrumentation and collection components; teams select a separate backend for storage and visualization.

Can I use observability without modifying application code?

Zero-code instrumentation can help when changing the application is impractical, though the available application-specific context depends on the setup.

Should every test assert on a trace?

No. Add trace assertions where execution path is part of the behavior you need to verify. Keep simpler tests focused on their relevant outcome.

12. Sources