ScreenshotNeo

BlogGuides

How to Improve Developer Experience in Testing

Improve testing developer experience by shortening feedback loops, reducing flaky failures, and making results easier to diagnose across the delivery lifecycle.

By the ScreenshotNeo team4 October 202612 min read

Improve the developer experience of testing by making feedback fast, reliable, and easy to act on. Start by tracing the time from a code change to an actionable result, then fix the slowest or least trusted checks in small steps. DORA recommends automated test feedback in less than ten minutes both locally and in CI; treat that as guidance to adapt to your system, not a guarantee or a universal law.

A test suite is useful when developers can run it often, trust its failures, and locate the cause without guesswork. The goal is not simply to add more tests. It is to build a feedback loop that helps the team answer: “How do I know if my product is working?” Google’s Testing Blog describes speed, reliability, and failure isolation as desirable properties of that loop.

1. Measure the current feedback loop

Before changing the suite, observe how long it takes to get a result that helps someone decide what to do. Record the stages separately so one slow step does not get mistaken for another:

  1. A developer makes a change and starts the relevant local checks.
  2. The checks finish and report a result.
  3. If a check fails, the developer identifies the cause and makes a fix.
  4. CI runs its checks and reports whether the change is safe to integrate.
  5. A broken build is diagnosed and restored.

Useful signals include local test runtime, CI build and test execution time, feedback availability, time to diagnose a failure, and time to fix a broken build. DORA calls out feedback availability, build and test execution, and time to fix broken builds as useful CI factors. A single average can hide a painful tail, so note both typical runs and the slow or stuck cases that interrupt work.

What to observe Question it answers Possible next step
Time to first useful local result Can developers run a meaningful check during normal work? Split or optimize the most common checks; run independent checks in parallel where practical.
CI time to actionable feedback How long does a change wait before the author knows what failed? Keep fast checks early and move longer checks to a later stage when that preserves useful coverage.
Repeat or rerun rate Do failures reproduce, or are people rerunning jobs to get a green result? Investigate flaky tests and unstable dependencies instead of normalizing retries.
Time to locate the failing behavior Does the report point to a useful assertion, component, or reproduction? Improve failure messages, isolation, logs, and test boundaries.
Maintenance effort Do tests keep breaking for unrelated implementation changes? Decouple tests from internals, simplify setup, or prune checks with little value.

Do not turn a target into a performance promise. DORA’s test automation guidance recommends feedback in less than ten minutes, and its CI guidance describes a few minutes as the goal and about ten minutes as an approximate upper limit. Architecture, risk, and the kind of check matter; use the guidance to prompt investigation when the loop is slow.

2. Make the common checks fast enough to run often

A local check that developers avoid because it takes too long cannot provide frequent feedback. Start with the checks people need during everyday edits: focused unit tests, static checks, and the smallest relevant integration checks. Keep those commands discoverable and make it possible to run a narrow test or package without rebuilding unrelated parts of the system.

  • Remove repeated setup work where safe, such as rebuilding immutable dependencies for every run.
  • Run independent checks concurrently if shared resources and test isolation permit it.
  • Use test selection or package-level commands for focused local work, while retaining broader checks in CI.
  • Cache work only when invalidation is reliable; stale artifacts can make fast feedback misleading.
  • If the build remains too long, DORA suggests improving test efficiency, adding resources to parallelize checks, or moving long-running tests to a separate pipeline stage.

Parallel execution helps only when tests are independent and the environment has enough capacity. If tests share mutable data, ports, accounts, or global state, parallelism may increase flakiness. First isolate those dependencies or partition the resources.

3. Make failures trustworthy and easy to diagnose

A failure should point to a product defect or a clearly identified test or environment problem. When the same change alternates between red and green, developers lose confidence and can begin ignoring failures, including real regressions.

Investigate flaky failures

  1. Capture the failing test, environment, relevant logs, and any shared service or fixture state.
  2. Rerun the smallest reproduction to determine whether the failure is deterministic.
  3. Check for timing assumptions, shared mutable data, order dependence, external service instability, and cleanup that does not complete.
  4. Fix the cause, isolate the test, or quarantine it with a named owner and a plan to resolve it. Do not make repeated retries the normal definition of success.

Retries can help distinguish transient infrastructure problems, but a green retry does not establish that a test is reliable. Keep the original failure visible and track whether it is a code defect, a test defect, or an environment issue.

Improve failure isolation

  • Use assertions that state the expected behavior and include relevant actual values.
  • Report the smallest failing test or case, not only a large suite-level failure.
  • Preserve useful logs, request identifiers, and setup details without dumping secrets.
  • Make test data and prerequisites repeatable, with cleanup that runs after failure.
  • Prefer tests coupled to observable behavior over tests that assert incidental internal details.

Google’s Testing Blog summarizes the purpose of the loop: “Tests create a feedback loop that informs the developer whether the product is working or not.” The practical test of a failure report is whether the person receiving it can understand what behavior failed and find a useful next step.

4. Use fast checks early and broader checks later

Testing should run throughout the delivery lifecycle instead of appearing as a final phase after development. Put checks with short runtimes and clear feedback early in the workflow. Use broader acceptance and nonfunctional checks at stages where their extra runtime and setup are justified.

Stage Typical purpose Experience considerations
Local development Check a focused behavior and catch obvious errors while editing. Offer quick commands, clear output, and reliable fixtures.
Early CI Validate the change with fast automated checks before expensive work. Fail quickly on actionable issues and identify the failing check directly.
Later CI or pre-release Run broader acceptance, compatibility, and appropriate nonfunctional checks. Keep ownership and diagnostics clear; separate long duration from the common edit loop when useful.
Exploratory and usability work Find behavior and user experience issues that automated checks do not cover well. Include testers in planning and feedback; capture reproducible findings for the team.

There is no universal ideal ratio of test types in the cited guidance. Choose a mix that reflects system risks, architecture, and the failures the team needs to detect. More end-to-end checks are not automatically a better experience if they are slow, fragile, or hard to diagnose.

5. Share test ownership across developers and testers

Developers should participate in creating and maintaining automated tests. DORA warns that separating developers from test automation can leave suites broken and can encourage designs that are difficult to test. Testers still bring important exploratory, usability, and acceptance perspectives, and they can pair with developers to improve automated coverage and investigate failures.

Make ownership practical rather than merely assigning a team name:

  • Have the author of a behavior change update or add relevant automated checks.
  • Route failures to people who can act on them, with a documented escalation path for shared infrastructure problems.
  • Pair developers and testers when acceptance criteria are unclear or behavior spans several components.
  • Review recurring failure patterns and suite maintenance as part of normal engineering work.
  • Keep exploratory testing in the lifecycle; automation does not replace observing how a feature behaves for users.

6. Review and curate the suite continuously

A test suite accumulates cost as the product changes. Review checks that are flaky, excessively expensive, hard to maintain, redundant, or tightly coupled to implementation details. Keep tests that detect meaningful defects and provide useful feedback; redesign or remove checks that create ongoing noise without useful signal.

If a UI change breaks many acceptance tests, consider decoupling the tests from the system under test. DORA gives the page object pattern as one example. If tests repeatedly need edits for unrelated code changes, examine whether they depend too heavily on mocks or whether some checks should be pruned. The right response depends on what behavior a test protects; do not delete a check solely because it failed during a refactor.

7. Improve a brownfield system incrementally

A legacy system does not need a comprehensive test suite before the team can improve feedback. DORA recommends starting with a small, working pipeline and extending it as the product evolves.

  1. Choose one representative, high-risk user or system behavior.
  2. Add a small automated check that runs reliably in the current environment.
  3. Put it in a pipeline stage where its result arrives early enough to help.
  4. Record how often it fails, whether failures are actionable, and what it costs to maintain.
  5. Use the results to choose the next behavior or risk to cover.

This avoids holding improvement hostage to a large retrofit. It also creates an opportunity to make a feature more testable as the team changes it, instead of trying to redesign the entire system before getting any feedback.

8. Troubleshoot common testing experience problems

Symptom Likely cause Fix
Developers skip local tests The common command is slow, hard to find, or requires fragile setup. Document a focused command, reduce repeated setup, and keep broad checks in CI.
A test passes only after rerunning Timing, shared state, order dependence, or unstable infrastructure. Reproduce the failure, isolate dependencies, and fix the underlying cause; track retries explicitly.
A failure report says only that a large suite failed Results are aggregated without test-level context or useful logs. Expose the specific case, assertion, expected and actual values, and relevant safe diagnostics.
Small edits trigger many acceptance failures Tests may rely on UI structure or implementation details that change often. Test stable behavior and consider a page object pattern for UI interaction.
CI is slow despite parallel jobs Shared resources serialize work, tests contend, or setup dominates runtime. Measure the bottleneck, isolate shared state, and parallelize only independent work.
Mocks make the suite pass but defects reach users Tests may verify assumptions about collaborators without exercising important integration behavior. Add targeted integration or acceptance checks for the risky boundaries and review what each mock hides.
The team delays testing until a release phase Testing is treated as a separate final activity. Run automated checks continuously and involve testers during development and acceptance planning.
A legacy area has little coverage The team is treating full retrofit as a prerequisite. Start with a small working pipeline and add representative checks as the product evolves.

9. Capture browser behavior without maintaining a browser harness

Browser checks can be part of a useful feedback loop when they protect a behavior that lower-level tests cannot cover. They also introduce browser setup, page timing, and presentation details that need maintenance. For visual review, bug reports, or a targeted browser artifact, a screenshot API can provide a repeatable capture step without requiring a team to operate its own capture browser.

ScreenshotNeo is a website screenshot API and MCP server for developers. A GET request can return a PNG, JPEG, WebP, or PDF. The API accepts the parameter names used by other screenshot APIs, which can make switching easier. Use the API for capture tasks; keep assertions about product behavior in tests designed to verify that behavior.

DIY: run a browser capture in code

With a browser automation framework, load the target page, wait for the behavior that matters, and capture an artifact. This Playwright example is runnable with Node.js after installing the package and browser:

npm install playwright
npx playwright install chromium
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  try {
    await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 30000 });
    await page.screenshot({ path: 'shot.png', fullPage: true });
  } finally {
    await browser.close();
  }
})();

For a test, prefer waiting for a specific user-visible condition over assuming a fixed delay. Full-page captures can trigger lazy-loaded content, and third-party requests can prevent a network-idle condition from arriving. Use a bounded timeout and make the wait match the page being captured.

Or skip the browser setup

Make one request to capture a page. See the ScreenshotNeo API documentation for options and configuration.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers say the page verdict and whether the shot was billed. Its MCP server lets AI agents using Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Sign up for 1,000 free screenshots a month with no card.

10. Performance, reliability, and cost tradeoffs

Performance

Keep frequently used checks short, run independent work in parallel where safe, and separate long checks when their results do not need to block every local edit. Measure the whole path to actionable feedback, including setup and diagnosis, not only test execution time. For browser captures, page complexity, network behavior, and wait conditions affect completion time; use bounded waits and capture only the artifact needed.

Reliability

Optimize for a result people trust. A smaller set of deterministic checks can provide a better daily experience than a broader suite whose failures are frequently ignored. Preserve clear evidence when a job fails, distinguish code failures from infrastructure failures, and review flaky checks until they are repaired or intentionally isolated.

Cost

Test infrastructure, parallel runners, and browser execution all consume resources. Compare the cost of additional capacity with the delay and contention it removes. Moving a long test to a later stage may improve the common loop, but only if the team still runs it at a suitable point and responds to its result. Keep an eye on maintenance cost as well as compute cost: a brittle check can consume engineering time long after its initial implementation.

For API-based screenshot capture, ScreenshotNeo bills only clean shots. Its listed monthly plans are Free: 1,000 shots with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free. These capture costs do not replace the need to budget for CI and the tests that verify application behavior.

Frequently asked questions

Should every change run the full test suite?

Run fast, relevant checks early and broader checks at appropriate later stages. The right selection depends on risk and system design; the cited guidance does not prescribe one universal mix.

Is a flaky test better than no test?

A flaky test can provide useful evidence while its cause is being investigated, but repeated noise erodes trust. Make the failure visible, assign ownership, and resolve or isolate the source rather than treating retries as a permanent fix.

Does a screenshot prove a feature works?

No. A screenshot captures visual output at a point in time. Use behavioral tests and acceptance checks to verify interactions and requirements; use screenshots as artifacts for review or visual comparison.

Where should a team with no test pipeline begin?

Start with a small, working pipeline and a representative behavior, then extend it as the product and its risks evolve.

Sources