7 Pitfalls to Avoid When Testing in Production
Production testing reveals real-world behavior, but unmanaged experiments can harm users and systems. Avoid these seven pitfalls with controlled exposure, useful signals, and a safe recovery plan.
Testing in production can reveal behavior that staging misses because real traffic, inputs, and mutable state are difficult to reproduce. It is safe only when exposure is controlled, success and failure are defined in advance, signals are useful, side effects are contained, and a responsible person can stop or reverse the change.
A canary is one way to do this: send a limited portion of traffic or service capacity to a new version, evaluate it, then decide whether to expand or stop. No single traffic percentage or observation period fits every service. Choose them based on traffic volume, risk, and whether the signals can support a useful decision. Google SRE’s canarying guidance and AWS ECS guidance both emphasize evaluation before wider rollout.
1. Sending the change to everyone at once
A full rollout gives a faulty change the widest possible blast radius. If the change behaves differently under real traffic, every user or dependent system can be affected before the team sees the problem.
Use gradual exposure where the architecture supports it. Options include a canary, traffic splitting, a one-box rollout, or blue/green deployment. The right choice depends on how traffic is routed, how quickly it can be switched back, the cost of running parallel capacity, and whether the change affects shared state.
- Start with a bounded group of users, requests, hosts, or regions.
- Keep the stable version available during evaluation when practical.
- Define who can pause expansion and how to do it.
- Expand in deliberate stages only after reviewing the agreed signals.
A canary limits initial exposure; it does not guarantee safety. Some failures affect all versions through a shared database, queue, cache, or external dependency. Include those shared components in the risk review. See AWS guidance on safe deployment practices.
2. Starting without a hypothesis or decision rule
“Watch it and see” leaves the team to interpret ambiguous dashboards after the change is live. People may disagree about what counts as a problem, or rationalize an early warning as noise.
Before deployment, write down:
- Change: What code, configuration, or feature is being evaluated?
- Hypothesis: What should improve or remain stable, and why?
- Success criteria: Which service and product outcomes must stay within acceptable bounds?
- Failure conditions: Which threshold, symptom, or user impact pauses or reverses the rollout?
- Decision owner: Who has authority to stop expansion, and who is on call?
Use thresholds that reflect the service’s normal variation and user impact. Avoid copying a threshold from another system without checking its baseline. AWS recommends predefined success criteria and failure conditions for production testing and automated reversal in its Well-Architected guidance.
3. Assuming a tiny sample proves safety
Small exposure limits the blast radius, but it can also produce too few observations to reveal a regression. This is especially likely for low-volume services, rare errors, infrequent workflows, or traffic segmented by region or customer type.
Estimate whether the canary will receive enough representative requests to evaluate the behavior in question. For a rare event, a short low-volume canary may observe no examples even when the underlying risk is meaningful. Consider a longer evaluation, a better targeted cohort, or a non-customer test method if the risk warrants it.
There is a real tradeoff: raising the traffic share improves the opportunity to observe behavior but exposes more users. Extending the evaluation can improve the sample while keeping the share bounded, but it lengthens deployment time. AWS ECS explicitly advises selecting a canary percentage that yields sufficient traffic for meaningful validation; it does not establish one universal minimum.
4. Watching dashboards informally or only after complaints
Waiting for support tickets means the test has already reached users without a timely detection path. Informal graph inspection can also miss subtle changes, especially when the team has no baseline or must decide on the fly whether a fluctuation matters.
Choose signals before rollout and compare the candidate with a baseline over the same period where possible. Useful signals commonly include:
- Request error rate and error type.
- Latency, including tail latency relevant to the service.
- Throughput and saturation or resource use.
- Service-specific outcomes such as successful checkout or completed jobs.
- Dependency failures, queue age, and other leading indicators tied to the change.
Set a review cadence or automated rule, and make sure someone can see and act on alerts during the rollout. A metric can be technically healthy while the user-facing outcome is broken, so pair infrastructure signals with the behavior the change is meant to preserve. AWS ECS discusses comparing canary and baseline task sets, while Google Cloud SRE’s release-canary account describes moving from manual graph review toward automated analysis.
5. Treating synthetic load as a perfect stand-in for production
Synthetic tests are useful, but generated traffic may not match organic traffic shifts, real inputs, or state-dependent conditions. Production behavior can depend on the mix of requests, customer history, caches, data shape, and interactions with other services.
Traffic teeing or request mirroring can make inputs more representative, but copied requests can still touch shared caches or state and distort results. A mirrored request must not accidentally charge a customer, send a message, change an account, or trigger another external action. Use a safe sink or suppress side effects before replaying traffic.
When direct customer exposure is too risky, use synthetic traffic or copied traffic in an isolated environment and be explicit about what the method cannot validate. For failure injection, apply guardrails and choose scope carefully; AWS’s resilience testing guidance calls for controlled experiments that limit harm.
6. Testing multiple moving parts without attribution
If several changes roll out together, an error increase may not reveal which change caused it. The same problem arises when logs and telemetry do not identify whether a request reached the baseline or candidate version.
Record the release, version, feature-flag state, and rollout cohort with relevant telemetry. Keep changes small enough to diagnose, or isolate independent changes into separate rollout groups. Run smoke checks and use logs, traces, and performance data to connect a symptom to the affected version and request path.
Microsoft’s incident management guidance recommends telemetry that links users to rollout phases alongside smoke checks, logs, tracing, and performance metrics. The practical goal is simple: when a signal moves, the team should be able to identify what changed and who or what was exposed.
7. Discovering rollback is unsafe or nobody is ready to act
A rollback plan is not real until the team knows the trigger, owner, steps, and communications path—and has checked that the old version can run against the state created by the new one.
Before rollout, verify:
- The rollback trigger is clear and tied to an observable signal or user impact.
- An authorized owner is available throughout the evaluation.
- The previous version can operate with current database schemas and persisted data.
- Reversal steps are documented and the needed deployment controls are accessible.
- For irreversible operations, there is a mitigation or forward-fix plan because reverting code may not undo the side effect.
- Stakeholders know how the team will communicate a pause, rollback, or incident.
Prefer backward-compatible, staged data changes when old and new versions may run at the same time. Automate reversal for well-defined signals when it is safe, but automation does not replace a recovery path that has been considered and operationally prepared. Google Cloud SRE’s account offers the concise operational advice “Rollback early, rollback often,” while AWS’s testing and rollback guidance emphasizes predefined conditions and recovery readiness.
A practical production-test checklist
- Scope the risk: List affected users, dependencies, data, and external side effects.
- Pick the rollout method: Choose a canary, traffic split, one-box, blue/green, synthetic test, or mirrored traffic based on exposure, fidelity, state, and reversibility.
- Write the decision rule: State the hypothesis, success criteria, failure conditions, evaluation window, and owner.
- Check signal quality: Confirm the sample can be informative and candidate metrics can be compared with a baseline.
- Tag the rollout: Make version and cohort visible in logs, traces, and metrics.
- Protect state: Prevent tests from causing unintended writes, customer charges, messages, or external actions.
- Prepare recovery: Verify rollback compatibility, access, ownership, and communications.
- Expand deliberately: Review the evidence at each stage; pause or reverse when a predefined condition is met.
How to choose a production testing approach
| Approach | Exposure | Fidelity | State and side effects | Operational tradeoff |
|---|---|---|---|---|
| Canary or traffic split | Limited at first, then increased | Real traffic for the exposed cohort | Can affect shared state and real users | Needs routing, comparison signals, and enough traffic; old and new capacity may overlap |
| Blue/green | Can switch traffic between environments | Can use production traffic after switch | Shared data and dependencies can still couple environments | Requires capacity for parallel environments and a reliable switch-back path |
| Synthetic traffic | Can be isolated from customers | Depends on how well scenarios match real use | Can be controlled if test data and destinations are isolated | Useful for specific paths, but may miss organic inputs and state |
| Mirrored or tee’d traffic | Original request serves users; copy is evaluated separately | Often more representative inputs | Copies may mutate shared state or invoke external actions unless prevented | Needs careful suppression, isolation, and result interpretation |
Compare approaches on exposure, fidelity, state effects, signal quality, attribution, operational cost, and reversibility. For example, ECS canary deployments keep old and new task sets during evaluation; this adds capacity and monitoring work, and a longer evaluation extends deployment time. Its example settings are product guidance, not universal thresholds.
Observing a live page during a rollout
Some release checks include capturing a page rendered by the candidate experience, such as a public landing page or a visual smoke check. A screenshot can help inspect the rendered result, but it does not establish that the backend is healthy, that an experiment has enough samples, or that the page represents every user segment. Treat it as one diagnostic artifact alongside telemetry and functional checks.
For a manual check, open the candidate URL in a browser, confirm the intended version and cohort, wait for the relevant content to settle, and save a screenshot with the release identifier and timestamp. If the page contains stateful actions, use a safe test account and avoid submitting irreversible actions. This method is useful for a quick visual comparison; automate it only when capture timing and test data are controlled.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and the response identifies the page verdict and billing status. Its MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These captures can document a rendered page during a rollout, but they do not replace release monitoring or a rollback decision rule.
Sign up for ScreenshotNeo and get 1,000 screenshots a month free, with no card required.
Troubleshooting production tests
| Symptom | Likely cause | What to do |
|---|---|---|
| No errors appear, but the sample is tiny | The canary received too little representative traffic or the event is rare | Check request counts and cohort mix; extend evaluation or change the test method without assuming zero observations means zero risk |
| Candidate and baseline dashboards disagree for unclear reasons | Different traffic mix, time window, or missing version labels | Align windows and cohorts, tag version and rollout group, and compare service-specific outcomes |
| Mirrored traffic changes production data | The copy is reaching shared state or external integrations | Stop mirroring; isolate storage and dependencies, suppress writes, and review any already-triggered effects |
| Alerts fire but the rollout owner is unavailable | Ownership and coverage were not arranged before deployment | Pause expansion; establish an on-call owner and escalation path before resuming |
| Rollback fails after a schema change | The new data shape is incompatible with the old application version | Use backward-compatible schema evolution; restore service with a forward fix or explicit data recovery plan if reversal is unsafe |
| A screenshot shows the wrong state or incomplete page | The capture happened before content settled, hit a different cohort, or saw a consent overlay | Verify the URL and rollout cohort, wait for a selector or settled state, and capture again; use screenshots as visual evidence, not health proof |
| A screenshot request is not billed but has no usable page image | The target may have returned a bot check, blank page, or failed load | Inspect the response verdict and billing headers, then check target availability and access requirements before treating it as a valid visual check |
Performance, reliability, and cost considerations
A production test adds work: parallel versions consume capacity, telemetry and comparisons need maintenance, and longer observation windows delay full rollout. Mirroring can increase load on downstream systems unless copies are isolated. Estimate these costs before choosing a method and include shared dependencies in capacity and side-effect reviews.
Reliability comes from limiting exposure, detecting meaningful changes, preserving attribution, and having a safe response path. A small canary without sufficient traffic is weak evidence; a large canary without a stop rule is broad exposure. Keep the evaluation long enough to observe the behavior of interest, but avoid treating any fixed duration as universally sufficient.
For page captures used as supplementary release evidence, ScreenshotNeo offers a free tier of 1,000 shots a month and paid tiers of 3,000 for $5, 15,000 for $15, 60,000 for $39, 250,000 for $99, and 1,000,000 for $249; yearly billing gives two months free. Every feature is available on every plan. Only clean shots are billed, including no charge for bot checks, blank pages, timeouts, failed loads, and cache hits. Confirm the response’s page verdict and billing headers when accounting for captures.
FAQ
How do you test in production safely?
Limit initial exposure, define success and stop conditions, compare candidate and baseline signals, protect shared state, and have an available owner with a recovery path. Choose synthetic or isolated traffic when real-user exposure is too risky.
What is canary testing?
It is a partial, time-limited release of a change that is evaluated before wider deployment. Its value depends on bounded exposure and enough representative observations to make a decision.
How much traffic should a canary receive?
There is no universal percentage. Choose a share that limits risk while producing enough relevant traffic for the behavior and error rate being evaluated.
When is rollback unsafe?
Rollback may be unsafe when the new version has changed persistent data or triggered irreversible external actions that the old version cannot handle or undo. Check compatibility before release and prepare a forward-recovery path where needed.


