ScreenshotNeo

BlogEngineering

Continuous Testing: How to Improve Software Delivery

Learn how to build fast, reliable checks into every stage of delivery, balance automation with human testing, and improve release confidence.

By the ScreenshotNeo team4 October 202610 min read

Continuous testing improves software delivery by giving teams useful feedback throughout the delivery lifecycle. Start with a repeatable build and a small, reliable set of fast checks on every change; add broader validation as the software moves through test and release environments; and keep exploratory, usability, and acceptance testing in the process. The goal is to find problems early while keeping the software releasable with evidence appropriate to its risks.

There is no universal test-suite layout. Choose checks based on the system’s architecture, failure risks, data, and dependencies. Measure whether tests provide timely, trustworthy feedback and whether delivery outcomes improve—not simply how many tests run.

1. What is continuous testing?

Continuous testing means testing throughout software delivery rather than waiting for development to be declared complete and then entering a separate testing phase. Automated checks can run on each change, while broader automated validation and human testing continue as changes move toward release. DORA describes automated and manual testing as complementary parts of this lifecycle.

Testing belongs to the whole team. Developers help create and maintain automated checks; testers work alongside developers to explore behavior, assess usability, and identify risks that scripts may miss. Findings from production and user feedback should inform future tests.

2. How is it different from testing at the end?

In an end-loaded process, a large body of changes can accumulate before the team gets meaningful testing feedback. That makes failures harder to trace and can put release work behind a late queue. Continuous testing moves suitable checks earlier and repeats validation as the software changes and is deployed.

It does not mean every test runs on every commit. A fast presubmit suite gives an early signal; slower or environment-dependent checks can run after deployment to a suitable test environment or at a later release stage. Human exploration can happen throughout development and before release.

Continuous integration (CI) is the practice of integrating changes frequently and building and testing them, commonly on each change. Continuous delivery aims to keep software in a state where it can be released on demand. Continuous deployment goes further by automatically putting every eligible change into production. These terms describe related practices, but they are not interchangeable.

3. Build a staged testing workflow

  1. Map the current path. Write down what happens from a code change through build, test, qualification, release, and post-deployment checks. Include manual handoffs, external dependencies, and where failures are discovered.
  2. Make builds repeatable. A change should produce a build and test result that teammates can reproduce. Keep configuration under version control and make dependencies and test data explicit where practical.
  3. Start with high-value, fast presubmit checks. Run the build, focused unit tests, static analysis, and other checks that can give a dependable result quickly. Google Cloud documents examples such as unit tests, fuzz tests, hermetic integration tests, and static and dynamic code analysis in its own presubmit approach; these are examples, not a required template for every team.
  4. Keep the shared mainline usable. Treat a broken build as priority work. Find the failing change, repair it or revert it, and restore a passing shared branch before stacking more uncertain changes on top.
  5. Deploy the same package to a suitable test environment. Add broader acceptance or integration checks there, along with relevant performance or vulnerability tests. Tests that require deployed infrastructure or realistic external connections may be more useful at this stage than in the earliest presubmit run.
  6. Make room for human validation. Give testers and product stakeholders a passing build to explore. Use exploratory, usability, and acceptance testing to investigate workflows, confusing interactions, and unexpected combinations.
  7. Check after deployment. Run smoke checks that verify essential system behavior and reachability of required external services. Feed incidents and production discoveries back into test planning.
  8. Review and improve the pipeline. Remove or repair checks that produce noise, take disproportionate effort, or no longer cover meaningful risk. Add checks when new features, incidents, or changed dependencies reveal gaps.

DORA recommends fast automated feedback, with guidance describing a target of less than ten minutes. Treat that as a feedback objective to work toward, not a promise that every suite or architecture can meet it. If the reliable first signal is too slow, identify which checks are necessary on every change and which can run in later stages.

4. What tests should run in a CI/CD pipeline?

Stage Possible checks What they help answer
Change or presubmit Build, focused unit tests, static analysis, and other fast checks Does this change compile, and does it break behavior that can be checked quickly?
Presubmit or early pipeline, where suitable Fuzz tests, hermetic integration tests, dynamic analysis Does the change fail under varied inputs or controlled component interactions?
After deployment to a test environment Broader acceptance and integration checks; relevant performance or vulnerability tests Does the deployed package work across more realistic paths and constraints?
Qualification or before release Exploratory, usability, and acceptance testing Does the experience work as intended, including cases scripts do not anticipate?
After deployment Smoke checks and operational observation Can users reach critical functions and required services?

Use this as a menu, not a fixed test pyramid or vendor blueprint. The right mix depends on the system and its risks. A large, slow end-to-end suite as the only signal on every change can delay feedback; moving all manual testing out of the lifecycle can leave important usability and exploratory questions unanswered.

5. Keep test feedback fast and trustworthy

  • Prioritize signal over test count. A test suite is useful when it catches meaningful failures and passes code that is ready for its stage. More tests can add maintenance and runtime without adding confidence.
  • Separate fast feedback from broader coverage. Keep the earliest suite focused on checks that need to gate a change. Run slower checks later or in parallel when that gives useful coverage without holding up the first result.
  • Investigate unstable tests. A test that fails intermittently trains people to ignore failures. Identify whether the cause is shared state, timing, unreliable dependencies, or inadequate isolation; fix or quarantine the check with an owner and a plan to restore it.
  • Make failures diagnosable. Keep logs, test names, relevant artifacts, and the failing change available. A red status without enough information to locate the issue is slow feedback in practice.
  • Control suite complexity. Review runtime and maintenance burden alongside coverage. Retire obsolete checks and update tests as interfaces and risk areas change.
  • Coordinate ownership. Developers, testers, and operations should agree how failures are triaged, who maintains checks, and what evidence is needed at each release decision.

6. Keep builds and releases reproducible

Use the same build artifact as it passes through environments rather than rebuilding separately for each one. Version configuration and automate deployment steps where practical. This makes it easier to connect the evidence from qualification to the package that is released and to investigate differences between environments.

Define release criteria in terms of product risk and the checks that matter for the change. A passing pipeline is evidence, not proof that every possible defect is absent. Smoke checks after deployment should cover essential behavior and important external-service connections. Automation should be developed collaboratively; tools alone do not create effective continuous delivery.

7. Measure delivery outcomes, not just test activity

Use a small set of measures to see whether the workflow is improving. DORA identifies delivery measures including lead time, change failure rate, time to restore service, and release frequency. Pair them with pipeline measures such as time from commit to automated build and test feedback and time to fix a broken build.

Read measures together. A faster pipeline is not an improvement if it stops finding meaningful failures; more frequent releases are not a success if changes become fragile or recovery suffers. Establish a baseline, make a focused change, and review whether feedback quality and delivery outcomes move in the intended direction. The research dossier does not provide a verified statistic or effect size to promise from adopting continuous testing.

8. Common problems and fixes

Symptom Likely cause Practical fix
Developers wait too long for the first result Every check, including slow environment-dependent tests, blocks the earliest feedback Keep a dependable fast suite on each change; stage broader checks after deployment or run them in parallel where useful.
The pipeline is often red, but failures are hard to reproduce Unstable tests, shared state, timing assumptions, or inconsistent dependencies Investigate the failing condition, improve isolation and repeatability, and assign ownership to unstable checks.
People rerun failures until they pass False alarms have made the signal untrustworthy Track recurring failures, repair or temporarily quarantine them with a clear owner, and avoid treating a rerun alone as a fix.
Large suites give little release confidence Test quantity has grown without regular review of coverage, reliability, and relevance Map checks to important behavior and risk; remove obsolete tests and add targeted checks for known gaps.
The main branch stays broken Failures are allowed to accumulate while more changes land Make restoring the shared build a team priority; fix or revert the responsible change before adding more uncertain work.
Tests pass, but a deployed release fails The tested artifact or configuration differs from the release, or deployment-specific behavior was not checked Promote the same package through environments, version configuration, and run deployment smoke checks.
Automation misses confusing or unexpected user journeys Manual exploration and usability work were removed or left until too late Include testers and product stakeholders during development and before release; turn useful discoveries into focused repeatable checks where appropriate.
Release frequency rises while failures or fatigue increase Frequency is being pushed without addressing fragile process or architecture Review change risk, stability, recovery, and team workload together; improve the underlying delivery system rather than optimizing frequency alone.

9. Where browser checks and screenshots fit

For web products, browser-based checks can help validate rendered pages, important user journeys, and visual changes. Treat screenshots as one kind of evidence: they can help reviewers inspect a page or compare output, but they do not replace functional tests, accessibility checks, security validation, or human usability testing. Decide which pages and states are high risk, make capture conditions repeatable, and avoid relying on a screenshot alone to certify a release.

Do it yourself: capture a page in a browser

This runnable Playwright example opens a URL, waits for the page to load, and saves a full-page PNG. Install Playwright and its browser first. Use a stable test page or a page your team is authorized to access.

npm install --save-dev playwright
npx playwright install chromium
// screenshot.mjs
import { chromium } from 'playwright';

const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}
node screenshot.mjs https://example.com

Make browser capture repeatable

  • Set the viewport and device scale deliberately; uncontrolled dimensions can change layout and image output.
  • Wait for a meaningful condition. Network idle can be unsuitable for pages with long-running connections or continuous requests; consider waiting for a selector or a specific app-ready signal instead.
  • Use deterministic test data and state. Handle authentication in a secure test context and avoid embedding credentials in committed scripts.
  • For lazy-loaded content, scroll or use an appropriate full-page capture strategy and confirm that the intended content rendered.
  • Keep browser and dependency versions consistent in CI, and close the browser in a cleanup path so failed captures do not leak processes.
  • When comparing images, account for dynamic timestamps, rotating content, fonts, animation, and environment differences. Hide or stabilize known dynamic regions where possible.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. Its API accepts a URL in one GET request and returns an image or PDF. Cookie banners are accepted like a visitor would accept them, and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

See the ScreenshotNeo API documentation for request options and setup. One call avoids managing browser installation and capture code. Screenshots are free up to 1,000 per month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

10. Frequently asked questions

Does continuous testing mean every test runs on every commit?

No. Run the fast checks needed for early feedback on each change, and stage broader or slower checks where they give the most useful evidence.

Does continuous testing remove manual testing?

No. Exploratory, usability, and acceptance testing remain useful throughout development and before release.

Does continuous delivery mean each change reaches production automatically?

No. Continuous delivery keeps software releasable on demand; continuous deployment adds automatic production deployment for eligible changes.

Which test framework should a team choose?

Choose based on the system, the risks to cover, how quickly and reliably the checks run, and the effort needed to maintain them. The cited guidance does not establish one universally best tool or suite layout.

Further reading