Smoke Testing vs. Sanity Testing: Key Differences
Smoke testing checks whether a build is ready for planned testing. Sanity testing has no single universal meaning; learn how to choose a clear team convention and build a useful readiness check.
Smoke testing is a quick readiness check of a few critical application paths. It helps a team decide whether a build is stable enough for planned, deeper testing. Sanity testing is used inconsistently: some references use it as another name for smoke testing, while some teams use it for a focused check of a recent change. That narrower meaning is a team convention, not a universal standard.
Before comparing the terms, agree on what they mean in your team. If your team distinguishes them, a practical convention is to call the build-wide critical-path gate a smoke test and a scoped check of a recent change a sanity test. Document the convention so a test name communicates what was checked and what a pass means.
What is smoke testing?
A smoke test is a preliminary check intended to give enough confidence that a test object is ready for planned testing. The ISTQB Glossary definition describes that readiness purpose. Microsoft’s Engineering Fundamentals Playbook likewise describes smoke tests as quick acceptance checks of a few behaviors, rather than full functionality coverage.
For a web application, a small smoke suite might check that the home page loads, a user can sign in, and a key workflow reaches its expected completion state. It should catch obvious problems that make further testing unproductive, such as a broken deployment, unavailable service, or unusable critical path. It does not prove the whole application works.
What is sanity testing?
There is no consistent distinction that applies across all teams and references. Some sources use “sanity test” as a synonym for “smoke test.” Other teams reserve “sanity test” for a narrow check after a small change or fix: does the changed behavior work, and are nearby areas still sound?
When using that narrower definition, state it as your team’s convention. Do not assume that another team, interviewer, test plan, or certification syllabus uses the same distinction. A test’s name alone is not enough to infer its scope.
Smoke testing vs. sanity testing: key differences
| Question | Smoke testing | Sanity testing |
|---|---|---|
| Is the term consistent? | Commonly used for a preliminary readiness check. | No. Some references use it as a synonym for smoke testing; some teams define a narrower check. |
| Typical purpose | Decide whether essential behavior works well enough to begin planned testing. | If distinguished locally, check a recent change and its nearby risks. |
| Typical scope | A few critical paths across the application or system. | Under the narrower convention, the changed feature and selected related behavior. |
| When to run | Early after a build or deployment, before deeper testing or the next delivery stage. | When the team’s convention calls for a focused post-change check. |
| What a pass means | The build is ready to enter the planned test stage, not that it is defect-free. | The selected scope passed its checks, not that unrelated application areas are sound. |
So “smoke = broad, sanity = narrow” can be a useful local shorthand, but it is not an uncontested standard. If a test plan uses the terms, define them there. For the terminology caveat, see the Microsoft guidance and the ISTQB glossary entry.
When should you run each check?
Run a smoke test when a build needs a readiness decision
Run the smoke suite soon after a build is deployed to a test environment or reaches a stage where planned testing can begin. Choose a few stable, important paths that should work on every usable build. If one fails, stop the downstream test or deployment work for that version and investigate. A failed gate is useful information: it prevents a team from investing in a build whose basic behavior is broken.
Run a focused sanity check if your team defines one
After a small change or fix, a team may use “sanity test” for a targeted check of the modified behavior and nearby risks. Choose the scope from the change and its dependencies. For example, after changing password reset, check that reset flow and the relevant sign-in behavior still work. This focused check is not a replacement for a smoke gate or regression coverage unless the team explicitly designed it to serve that purpose.
How to design a useful smoke suite
- Pick critical paths. Start with behaviors whose failure would make the build unusable for the next test stage: page or service availability, authentication, and one core transaction are common candidates.
- Keep checks fast and shallow. Verify that each path can complete and reaches a meaningful expected state. Avoid turning the readiness gate into exhaustive field validation or a full regression suite.
- Make checks deterministic. Use controlled test data, explicit setup and cleanup, and stable assertions. Avoid depending on shared accounts or data another test can change.
- Define pass and fail conditions. A status code alone may not establish that a workflow is usable. Check the important result, such as an authenticated page or a created test record.
- Run at the right boundary. Trigger the suite against the deployed build and environment that the next stage will use. Record the build identifier and environment with the result.
- Make failure actionable. Report which path failed, the observed result, and enough context to reproduce it. Stop or hold downstream work when a critical gate fails.
- Review the suite as the product changes. Keep the paths representative and maintainable. Remove checks that no longer answer a readiness question; add a path when its failure would block planned testing.
A runnable browser smoke check in JavaScript
This example uses Playwright to open a site and verify that its main heading is visible. It is a minimal browser smoke check, not a complete test suite. Replace the sample URL and heading with a stable test environment and an assertion that reflects your application.
import { chromium } from 'playwright';
const baseUrl = process.env.BASE_URL ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
const response = await page.goto(baseUrl, {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
if (!response || !response.ok()) {
throw new Error(`Home page returned ${response?.status() ?? 'no response'}`);
}
const heading = page.getByRole('heading', { name: 'Example Domain' });
await heading.waitFor({ state: 'visible', timeout: 10_000 });
console.log(`Smoke check passed: ${baseUrl}`);
} catch (error) {
console.error(`Smoke check failed: ${error.message}`);
process.exitCode = 1;
} finally {
await browser.close();
}
Save as smoke.mjs, then run npm install playwright, npx playwright install chromium, and BASE_URL=https://example.com node smoke.mjs. In CI, set BASE_URL to the deployed test environment. Keep secrets out of source control and inject them through the CI secret store if the flow requires authentication.
For a real application, consider adding checks for a critical route, login with a dedicated test identity, and one essential workflow. Use role-based locators or stable test IDs instead of brittle selectors tied to layout. Keep test data isolated and clean it up where appropriate. A screenshot can help explain a visual failure, but it does not establish that an application’s data or business rules are correct.
Use screenshots as visual evidence, not as the readiness test
A browser screenshot can make a failed page load, unexpected overlay, or obvious rendering problem easier to review. It is supplementary evidence: a screenshot alone cannot verify authentication state, database writes, calculations, or other functional outcomes. Keep assertions in your test code and attach an image when visual context helps someone diagnose a failure.
For a manual capture, open the target page in a browser, wait for the relevant content to render, and capture the viewport or full page. Record the build and environment alongside the image. Avoid capturing pages containing real user data, credentials, or private tokens.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. A single request captures a URL as an image or PDF; the same capture can supply visual evidence for a human review, while your smoke test still checks behavior. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie and consent banners, newsletter popups, and chat widgets are removed before capture, and each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status. An MCP server lets AI agents use screenshot, page-info, and PDF tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
Options for adding a screenshot to a check
If your goal is to preserve a page image as diagnostic evidence, ScreenshotNeo supports options including full-page capture, CSS-selector element capture, viewport and device presets, dark mode, custom wait conditions, and custom JavaScript or CSS. Its API also supports caching, bulk capture, and asynchronous jobs with signed webhooks. These options configure capture; they do not determine whether the application passes your smoke-test assertions. Consult the docs for parameter names and request details.
For a CI artifact, use the captured image to help investigate a failure. Protect the API key as a secret, use a test URL that does not expose private data, and consider whether caching is appropriate when each build must capture the newest deployed state.
Performance, reliability, and cost considerations
- Keep the gate small. Every extra workflow adds execution time and maintenance. Select checks that answer whether the next stage can proceed.
- Separate application failures from test-environment failures. A DNS issue, unavailable dependency, expired test credential, or bad test-data setup can fail a check even when the application change is not the cause. Log the environment and useful response details.
- Use bounded waits and clear timeouts. An unbounded wait can stall a pipeline; overly short timeouts can make a healthy but slower environment appear broken. Set limits that fit the environment and report which operation timed out.
- Retry with care. A retry may help distinguish a transient infrastructure issue from a repeatable failure, but retries can also hide flaky tests or delay detection. Preserve the first failure and make retry policy visible.
- Keep test data isolated. Parallel runs that share mutable accounts or records can cause intermittent failures. Use per-run data or a cleanup strategy.
- Budget for evidence separately. Browser tests consume CI resources. Screenshot API costs depend on the selected service and plan; ScreenshotNeo’s listed plans range from 1,000 monthly free shots to paid tiers starting at $5 for 3,000. Only clean shots are billed under its stated billing rules. Verify current details in its docs and pricing before implementing a production workflow.
Troubleshooting common smoke-check failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Navigation returns an error or no response | Wrong base URL, deployment not ready, DNS/network issue, or service unavailable. | Log the target URL and response status; confirm the deployment and environment are reachable from the runner. |
| Element assertion times out | Wrong locator, page content changed, authentication redirected elsewhere, or the page is still loading. | Inspect the final URL and page state; use a stable role or test ID; wait for the specific expected state rather than an arbitrary long delay. |
| Passes locally but fails in CI | Different environment variables, network access, browser setup, data, or timing. | Compare configuration and build identifiers; install the required browser; use isolated test data and capture useful failure context. |
| Intermittent failures | Race conditions, shared mutable data, unstable dependencies, or excessive reliance on timing. | Wait for a meaningful condition, isolate records and accounts, and record the first failure even if a retry is allowed. |
| A smoke suite takes too long | The gate has accumulated deep regression tests or unnecessary setup. | Keep readiness checks to a few critical paths; move broader behavior coverage into the planned test stage. |
| Screenshot shows a challenge or blank page | The target may present a bot check or fail to load in the capture environment. | Check the target and response metadata. With ScreenshotNeo, inspect X-Page-Verdict and X-Billed; bot checks, blank pages, and failed loads are not billed. |
| Screenshot request is rejected | Missing or invalid access key, malformed request, or invalid URL. | Check the key, encode the URL, inspect the HTTP response, and compare request parameters with the API docs. |
| Screenshot is stale | A cached response may be returned when caching is enabled. | Review the configured cache behavior and TTL, or disable caching for captures that must reflect the current build. |
Frequently asked questions
Does a smoke test replace regression testing?
No. It is a quick readiness gate for a few essential paths. Planned functional and regression testing still covers the required breadth and depth.
Can a smoke test be automated?
Yes. Teams commonly automate stable critical-path checks and run them after a build or deployment. The important part is that the check answers a defined readiness question.
Is a failed smoke test always an application defect?
No. A failed check can come from the application, deployment, test data, credentials, network, or test code. Diagnose the failure before assigning its cause.
Should our team use the word “sanity”?
Use whichever label your team understands, and define its scope in the test plan. If communicating outside the team, describe what the check covers instead of relying on the label.
Key takeaways
- A smoke test is a quick gate asking whether essential behavior works well enough to start planned testing.
- “Sanity testing” is used both as a synonym and, by some teams, as a name for a narrow post-change check.
- Define local terminology and pass criteria; do not treat “smoke = broad, sanity = narrow” as universal.
- Keep readiness checks fast, deterministic, and actionable. Use screenshots as supporting visual evidence, not as functional assertions.


