ScreenshotNeo

BlogHow-to

How to Monitor Websites and Automation Jobs

Learn how to monitor website availability, detect failed cron jobs, configure heartbeats, reduce false alerts, and capture visual evidence.

By the ScreenshotNeo team1 October 202610 min read

Use two complementary checks: an external HTTP or HTTPS monitor for every public website and business-critical endpoint, plus a heartbeat monitor that a scheduled job calls only after it finishes successfully. HTTP monitoring tells you whether a service can be reached and returns the expected result. Heartbeats tell you whether a backup, report, synchronization task, or other scheduled process actually ran.

A status code alone is not enough for many sites. A broken application can return a friendly error page with HTTP 200. Add a keyword or content assertion when the page must contain a known phrase, and use a browser or screenshot check when layout or visual output matters.

1. Decide what success means

Write the success condition before creating a monitor. Typical checks include:

  • Availability: DNS resolves, TLS succeeds, and the URL responds.
  • HTTP result: the response has an expected status such as 200, 204, or 3xx.
  • Content: the body contains a phrase such as “Order complete” or does not contain an error marker.
  • Performance: response time remains below a limit appropriate for the endpoint.
  • Business transaction: a login, checkout, search, or API request completes with the expected result.
  • Scheduled completion: a job sends a success-only heartbeat before its deadline.

Monitor public sites, APIs, login pages, checkout paths, and other URLs whose failure creates user or business impact. Track certificate and domain expiry where your monitoring service supports it. Use more than one monitoring location when regional reachability matters.

2. Add an external HTTP(S) monitor

  1. Choose the exact URL to check. Prefer a lightweight health endpoint for APIs and a representative page for websites.
  2. Set the accepted status codes and redirect policy.
  3. Configure a timeout shorter than the maximum delay your users can tolerate.
  4. Add a keyword or content assertion if a status code could hide an application error.
  5. Choose a check interval and at least one independent location.
  6. Route alerts to an actively watched address or incident integration.
  7. Configure repeated notifications for unresolved incidents and a maintenance window for planned work.

Minimal HTTP check with cURL

#!/usr/bin/env bash
set -euo pipefail

url="https://example.com/health"
status=$(curl --silent --show-error --output /tmp/health-body \
  --write-out "%{http_code}" \
  --max-time 15 "$url")

if [ "$status" != "200" ]; then
  echo "health check failed: HTTP $status" >&2
  exit 1
fi

if ! grep -q "ok" /tmp/health-body; then
  echo "health check failed: expected content missing" >&2
  exit 1
fi

echo "healthy"

HTTP and keyword check in Python

import sys
import requests

URL = "https://example.com/health"
try:
    response = requests.get(URL, timeout=15)
    response.raise_for_status()
except requests.RequestException as exc:
    print(f"request failed: {exc}", file=sys.stderr)
    raise SystemExit(1)

if "ok" not in response.text:
    print("expected keyword was not found", file=sys.stderr)
    raise SystemExit(1)

print(f"healthy: {response.status_code} in {response.elapsed.total_seconds():.3f}s")

HTTP and keyword check in Node.js

const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 15_000);

try {
  const response = await fetch('https://example.com/health', {
    signal: controller.signal,
  });
  const body = await response.text();

  if (!response.ok || !body.includes('ok')) {
    throw new Error(`health check failed: HTTP ${response.status}`);
  }

  console.log(`healthy: ${response.status}`);
} finally {
  clearTimeout(timeout);
}

3. Monitor cron jobs and automation with heartbeats

Cron job monitoring, also called heartbeat monitoring, checks whether scheduled tasks and background jobs are running on time. The job calls a unique URL after successful completion. If the ping is late or absent, the monitor alerts you. This catches a job that failed, exited early, was disabled, or never started.

Apply heartbeats to backups, report generation, data synchronization, SSL renewal, email delivery, cache rebuilds, and cleanup scripts. Create a separate monitor for each important job, with the job name, environment, owner, and expected schedule in the monitor name or tags.

Heartbeat pattern

  1. Create a unique heartbeat URL with an expected interval and grace period.
  2. Run the job’s normal work.
  3. Send the heartbeat only after every required step succeeds.
  4. Exit nonzero and skip the heartbeat when a step fails.
  5. Review the last successful heartbeat and expected next check-in during an incident.

Unix shell job with a success-only heartbeat

#!/usr/bin/env bash
set -euo pipefail

heartbeat_url="https://heartbeat.example.net/your-unique-token"

run_backup() {
  pg_dump "$DATABASE_URL" > /var/backups/app.sql
  aws s3 cp /var/backups/app.sql s3://example-backups/app.sql
}

run_backup
curl --fail --silent --show-error --max-time 15 "$heartbeat_url" >/dev/null
echo "backup completed and heartbeat sent"

Place the heartbeat command in the final success path. With set -e, a failed backup prevents the ping from being sent.

Crontab example

# Run every 15 minutes
*/15 * * * * /usr/local/bin/run-backup.sh >> /var/log/run-backup.log 2>&1

Windows PowerShell example

$ErrorActionPreference = "Stop"
$heartbeat = "https://heartbeat.example.net/your-unique-token"

& "C:\Tools\run-backup.exe"
if ($LASTEXITCODE -ne 0) {
  throw "backup failed with exit code $LASTEXITCODE"
}

Invoke-WebRequest -Uri $heartbeat -Method Get -TimeoutSec 15 | Out-Null

Python job with a success-only heartbeat

import os
import sys
import requests

HEARTBEAT_URL = os.environ["HEARTBEAT_URL"]

try:
    # Replace this with the real job.
    generate_report()
    upload_report()
    requests.get(HEARTBEAT_URL, timeout=15).raise_for_status()
except Exception as exc:
    print(f"job failed: {exc}", file=sys.stderr)
    raise SystemExit(1)

print("job completed and heartbeat sent")

Node.js job with a success-only heartbeat

const heartbeatUrl = process.env.HEARTBEAT_URL;

try {
  await generateReport();
  await uploadReport();

  const response = await fetch(heartbeatUrl);
  if (!response.ok) {
    throw new Error(`heartbeat failed: HTTP ${response.status}`);
  }
} catch (error) {
  console.error(error);
  process.exitCode = 1;
}

4. Choose intervals and grace periods

Set the interval from the job’s normal schedule and the maximum acceptable detection delay. A job that runs every five minutes can use a five-minute interval plus a short grace period. A nightly backup needs a much longer interval and a grace period that covers normal runtime variation.

For website checks, shorter intervals detect incidents sooner but create more requests and alert noise. UptimeRobot documents intervals of 5 minutes on its Free plan, 1 minute on Solo and Team, and 30 seconds on Scale. Recovery is recorded on the next scheduled check, so a service can be healthy again before the monitor records recovery.

Workload Starting interval Grace-period guidance
Health endpoint 1–5 minutes Allow for transient network failures and one retry.
Five-minute worker 5 minutes Add only enough time for normal runtime variation.
Hourly report 60 minutes Include the longest normal report duration.
Nightly backup 24 hours Allow for the normal backup window and upload time.

5. Reduce false positives and improve alert quality

  • Use retries for transient connection failures, but do not retry indefinitely.
  • Require more than one failed observation before paging for a low-risk endpoint.
  • Include monitor name, URL or job identity, expected check-in, last success, location, and failure reason in alerts.
  • Send recurring reminders while an incident remains unresolved, with a reasonable cap.
  • Pause monitors or schedule maintenance before deployments, migrations, and planned job pauses.
  • Prevent overlapping job runs. A late first run must not create a misleading heartbeat from a second run.
  • Keep credentials, heartbeat tokens, and API keys in a secret manager.
  • Tag monitors by service, environment, owner, and severity.

6. Manage monitors as infrastructure

When monitors are part of a deployment or platform, create and update them through an API or CLI instead of editing a dashboard by hand. UptimeRobot’s official CLI supports monitor, incident, status-page, maintenance-window, alert-contact, integration, and tag operations, with API-key authentication and JSON output for shell and CI workflows.

Store the API key in the CI secret manager, review monitor changes like infrastructure code, and keep names and tags consistent with service ownership. Export the monitor definition or configuration so a new environment can be rebuilt.

7. Add visual evidence when content or layout matters

An HTTP status and keyword assertion cannot detect every browser-facing failure. A page can return the expected text while a cookie banner covers the content, a chart is blank, or a client-side layout is broken. Add a browser or screenshot step when visual evidence is useful for release checks, reports, or incident review.

Or skip the browser setup

ScreenshotNeo is a website screenshot API that can be called from a monitoring or automation job. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Before capture, it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Failed loads, bot checks or CAPTCHAs, blank pages, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options, including full-page capture, CSS element capture, device presets, dark mode, retina scale, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous jobs, webhooks, bulk capture, and usage data.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Use the returned headers to record whether the page was cleanly captured and billed. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

8. Troubleshooting

Symptom Likely cause Fix
Monitor reports down but the site works in a browser Regional routing, bot protection, or an overly short timeout. Check from another location, review server logs, allow the monitoring user agent where appropriate, and increase the timeout only to a useful limit.
HTTP 200 with a broken page Application error rendered inside a successful response. Add a keyword, JSON field, or transaction assertion.
Heartbeat never alerts The ping runs at job start or from a finally block. Send it only after all required work and validation succeed.
Heartbeat alerts during normal long runs Interval or grace period is shorter than the job’s runtime variation. Measure normal duration and increase the grace period to cover it.
Duplicate success notifications Overlapping cron runs or retries invoke the heartbeat more than once. Use a lock, idempotency key, or single-run guard and document retry behavior.
Many alerts during deployment No maintenance window was configured. Schedule maintenance before planned changes and close it afterward.
Certificate errors from the monitor Expired certificate, incomplete chain, or incorrect hostname. Inspect the certificate chain from outside the server network and renew or correct the configuration.
Screenshot capture is blank or blocked Bot check, page timeout, client-side rendering delay, or a page that requires authentication. Use waits, custom headers or cookies where authorized, and inspect X-Page-Verdict; failed captures are not billed by ScreenshotNeo.

9. Performance, reliability, and cost considerations

  • Keep health endpoints cheap: avoid database-heavy checks unless the database is part of the failure you need to detect.
  • Separate liveness from readiness: liveness proves the process responds; readiness can include dependencies such as a database or queue.
  • Use independent locations: one network path cannot distinguish a regional outage from a site outage.
  • Control alert volume: retries, maintenance windows, escalation, and recurrence settings protect the signal-to-noise ratio.
  • Budget by checks: monitor count, interval, retention, locations, seats, and integrations affect monitoring cost; verify current plan limits before purchasing.
  • Budget job monitoring separately: a heartbeat usually adds one small request per successful run, while the operational value comes from detecting missed runs quickly.
  • Use screenshots selectively: capture only pages where visual evidence is needed, choose an appropriate cache TTL, and use asynchronous jobs or bulk capture for large batches.

10. Operational checklist

  • Every public site and critical endpoint has an external HTTP(S) monitor.
  • Important pages have content or transaction assertions.
  • Every material scheduled job has a unique success-only heartbeat.
  • Intervals and grace periods match real schedules and runtimes.
  • Alerts include ownership, severity, location, and the last successful check.
  • Maintenance windows are part of the deployment runbook.
  • Monitor definitions and credentials are managed through code and secret storage.
  • At least one alert path has been exercised deliberately.

FAQ

Should I use a heartbeat URL, HTTP check, or keyword monitor?

Use an HTTP check for an externally reachable service, a keyword or structured-content assertion when status alone is insufficient, and a heartbeat for a scheduled process that must report successful completion. Important systems often need more than one.

What happens if a job is still running when its heartbeat deadline passes?

The monitor can alert before completion. Set the grace period longer than the normal maximum runtime, or use a monitor design that matches the job’s completion window.

Can a heartbeat prove that every step succeeded?

Only if the heartbeat is sent after every required step and the script exits on failure. A ping placed at startup proves only that the process began.

How do I monitor a private endpoint?

Use an authorized monitoring path, VPN or agent-based check, or expose a narrowly scoped health endpoint. Never put long-lived secrets directly in a public URL unless the monitoring service is designed for that use.

When is a screenshot check worth adding?

Add one when visual layout, consent overlays, charts, rendered reports, or browser-only content affects the outcome. Keep HTTP and heartbeat checks as the fast baseline and use screenshots for evidence or visual assertions.