ScreenshotNeo

BlogGuides

Website Monitoring Alert Error Handling: A Practical Guide

Design uptime alerts that catch real customer impact, filter flaky checks, and help you diagnose both failed checks and missing notifications.

By the ScreenshotNeo team29 September 20269 min read

Website Monitoring Alert Error Handling: A Practical Guide

A useful website alert answers three questions: what failed, whether customers are affected, and what the responder should do next. Start by defining the endpoint or user journey, expected response, timeout, probe locations, and confirmation policy. Then route urgent, actionable failures to a page and send lower-priority signals somewhere less disruptive. Keep the monitoring path under observation too: a healthy website does not help if checks stop running or notifications stop arriving.

1. Decide what “down” means

An HTTP uptime check requests a configured URL and evaluates it against criteria such as status code, response content, and latency. A synthetic monitor can run a sequence, such as opening a login page or exercising checkout. A passing request to a homepage does not prove every user workflow works. Google Cloud distinguishes uptime checks from scripted synthetic monitors and retains results such as failure details and latency for investigation (Google Cloud synthetic monitoring overview).

Write the check contract before enabling alerts. Record:

  • Target: the public endpoint or workflow and its environment.
  • Expected result: acceptable status codes and, when useful, a stable response marker. Avoid matching transient text or personalized content.
  • Timing: the timeout and check schedule. A timeout that is shorter than normal response variation creates noise; one that is too long delays detection.
  • Perspective: probe location or locations, authentication needs, and any network restrictions.
  • Maintenance behavior: planned windows, expected redirects, and how deployments should affect checks.

Use a basic reachability check for “can a client reach this endpoint?” Use a workflow check when the customer outcome depends on multiple steps, scripts, or a third-party service. A workflow check should use a safe test account and avoid creating real orders, sending messages, or changing customer data.

2. Reduce false alarms without hiding outages

A single failed probe can reflect a real outage, but it can also reflect a transient network problem between that probe and your service. Google Cloud’s documented default alert policy requires failures from at least two regions at the same time; it can be edited, and Google recommends the default to reduce notifications from transient failures (Google Cloud troubleshooting guidance). This is one vendor’s behavior, not a universal setting.

Comparing probe locations helps distinguish a broad outage from a single flaky check.
Comparing probe locations helps distinguish a broad outage from a single flaky check.

Confirmation rules trade noise for speed. Requiring multiple locations, consecutive failures, or a retest window can reduce alert flapping. Each extra confirmation can delay detection, and more probes or higher frequency can affect cost. Grafana notes that multiple probes help reduce flapping and recommends balancing probe count and frequency against cost (Grafana uptime and reachability).

Choice Helps with Tradeoff
One location, immediate condition Fast signal from that vantage point More exposure to isolated probe or route issues
Several locations must fail Confidence that the problem is broadly reachable Can miss a regional outage and may take longer to confirm
Consecutive failures or retest window Short blips and alert flapping Detection waits for confirmation
Endpoint plus synthetic workflow Different layers of customer experience More setup and separate failure modes to maintain

Choose based on the impact and recovery time your service can tolerate. A checkout failure may deserve a short confirmation and immediate page; a low-impact informational endpoint can tolerate a longer window and a ticket. Do not copy another provider’s defaults without understanding what they confirm and how long notification takes.

3. Make the alert actionable

Prometheus summarizes its guidance this way: “keep alerting simple, alert on symptoms, have good consoles to allow pinpointing causes, and avoid having pages where there is nothing to do” (Prometheus alerting practices). Build alert conditions around customer impact rather than every possible technical cause.

A page should include or link to the monitored target, the first failure time, current state, affected region or regions, recent result history, response or error details, and the relevant dashboard or runbook. Name an owner and a next step. If a responder cannot take action, use a ticket, dashboard, or lower-priority channel instead of waking them.

Separate the condition from its routing. For example, one policy can indicate a sustained customer-facing outage while a distinct lower-severity signal records a single probe failure. Add a maintenance window for planned work, but keep a way to surface an unexpected extension or a separate critical path that remains broken.

4. Monitor the monitor and notification path

The website is only one part of the system. The monitor can fail to execute, produce no data, or return a query error. The notification integration can be disconnected, misrouted, or delivered to an inactive contact. Decide deliberately how execution errors, timeouts, and missing data should appear to responders; there is no single correct state for every deployment. Grafana documents alert states for query execution errors and timeouts, and discusses synthetic checks for external availability (Grafana connectivity errors in alerts).

Use an independent signal to check that the primary monitoring path is alive. That can be a heartbeat or a separate check that expects a periodic result, paired with a distinct notification route if practical. A check that silently stops reporting must not look like a healthy target. Verify both that the monitor evaluates and that a notification can reach the intended person.

Google Cloud describes alert policies as conditions plus notification channels and documentation, and notes that data visibility, evaluation, retest windows, and channel behavior affect when an alert arrives (Alerting overview, metric alerting behavior). Report the outage start and notification arrival as separate timestamps. Do not promise instant delivery.

5. Troubleshoot a monitor that says the site is down

  1. Mark the first failure time. Compare it with deploys, DNS or certificate changes, incidents, and maintenance windows.
  2. Open the raw execution result. Inspect status, response snippet, error type, latency, logs, and the exact failure stage. Synthetic executions may retain error messages, line information, execution time, logs, and metrics; see Google Cloud’s result details.
  3. Compare probes. Check whether locations fail together, whether only one region is affected, and whether another independent client can reach the endpoint.
  4. Recheck the contract. Confirm URL, redirects, TLS, expected status/content, authentication, timeout, and the deployed response. A changed success page can fail a content assertion even while users can browse.
  5. Check your service path. Review application and proxy logs, dependency health, DNS resolution, certificate validity, firewall or allowlist changes, and recent releases.
  6. Decide whether action is needed. If evidence points to an isolated probe error, record it and watch the next result. If customer impact is plausible, follow the incident runbook while gathering more evidence.

A successful probe from one place is not proof that all users can reach the service. Conversely, one failed probe is not proof of a broad outage. Use multiple signals and customer-impact evidence to decide severity.

6. Troubleshoot a missing notification

First establish whether the condition fired. Inspect the monitor’s result history and alert or incident history around the expected time. If the check never met its failure condition, investigate the threshold, retries, location agreement, and retest window. If an incident exists but no person received it, follow the delivery path.

  1. Confirm the alert condition. Check whether it requires multiple regions, repeated failures, a latency threshold, or a sustained duration. Verify the monitor is assigned to the policy.
  2. Confirm recipient assignment and activation. The contact must be active and attached to the monitor or policy. UptimeRobot’s support guidance specifically calls out checking that alert contacts are active and attached; its monitor troubleshooting also describes retry behavior before marking a monitor down (UptimeRobot troubleshooting).
  3. Inspect channel health. Check integration credentials, destination address or number, routing rules, escalation policy, and provider delivery status.
  4. Check suppression and timing. Look for maintenance silences, deduplication, quiet hours, retry behavior, and evaluation or delivery delays.
  5. Test end to end. Trigger a test notification through the same policy and channel, then confirm receipt with the intended responder. A configuration screen showing “connected” does not prove the entire route works.

Keep the test route documented and repeat it after changing integrations, contacts, or escalation rules. A channel test should not depend on the same system whose failure it is meant to reveal.

7. Common errors and fixes

Symptom Likely cause What to check
Monitor says down, browser seems fine Probe-specific route, timeout, or response criterion mismatch Raw result, location comparison, redirect chain, TLS, status/content rule
Monitor says up, users report failure Check covers only a shallow endpoint or a different region/path Add a safe workflow check; compare user geography and dependencies
Alerts arrive late Confirmation window, evaluation delay, data visibility, or channel delivery time Review each stage’s timestamps and shorten confirmation only if the impact warrants it
Many pages for short blips Single-probe sensitivity, unstable assertion, or too-frequent evaluation Use measured confirmation, stabilize the response check, route brief events at lower severity
No incident appears Condition did not cross threshold, monitor is detached, or execution produced no data Inspect raw results, policy membership, no-data behavior, and evaluation history
Incident exists, no notification Inactive or unassigned contact, broken integration, suppression, or delivery failure Check routing, channel status, provider logs, and end-to-end test

8. Pick a monitoring setup that fits the service

Compare tools by the operational questions they answer: endpoint versus scripted workflow checks; probe geography; retry and confirmation behavior; latency and response validation; retained logs or failure detail; handling of no-data and execution errors; notification and incident integrations; maintenance controls; ownership; and cost at your intended frequency. Google Cloud Monitoring, Grafana Cloud Synthetic Monitoring, and UptimeRobot are examples in the research sources; their defaults and features differ. UptimeRobot describes endpoint checks, multiple locations, alert channels, recurrence settings, and maintenance windows on its website monitoring page. Verify current vendor configuration and pricing before adopting it.

Choose the smallest combination that gives the team an early, reliable signal and enough evidence to respond. A shallow HTTP check plus a carefully chosen customer journey often tells more than dozens of checks without clear ownership.

9. Capture the failure page for the incident record

A screenshot can preserve what a human-facing page displayed during an incident, such as a maintenance page, broken layout, or consent overlay. It complements status codes, logs, and synthetic results; it does not replace them. For repeatable captures, use a documented capture path and avoid placing secrets or customer data in the URL or screenshot.

A screenshot records page appearance alongside the monitor’s logs and timestamps.
A screenshot records page appearance alongside the monitor’s logs and timestamps.

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request can return a screenshot or PDF; its page cleanup accepts cookie banners and removes known consent platforms, newsletter popups, and chat widgets before capture. See ScreenshotNeo and its API documentation.

Or skip the browser setup

For a one-off incident capture, call the ScreenshotNeo API with the affected page URL. Keep your API key private and store the returned image with the incident record.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

The API accepts browser options including viewport or device, full page, selector capture, wait conditions, custom headers, cookies, and output format; consult the docs for parameter details. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.

FAQ

How do I get an alert when a website is down?

Configure an endpoint or workflow check with explicit success criteria, attach an active contact and notification channel, and test delivery end to end. Set confirmation to match the impact and acceptable detection delay.

Why does my monitor show as down?

It may have observed a genuine failure, or a probe timeout, regional route issue, or response assertion mismatch. Start with the raw execution result and compare locations before changing the alert threshold.

Why didn’t I receive a notification?

Check whether the alert condition fired first. If it did, inspect contact assignment, activation, integration health, suppression rules, and delivery logs, then send a test through the same route.

Does a successful uptime check prove the site works?

No. It proves only that the configured request met its criteria from the check’s perspective. Important user journeys need their own safe synthetic checks.