ScreenshotNeo

BlogComparisons

The Best DNS Monitoring Tools for Website Performance

Compare DNSPerf, ThousandEyes, Catchpoint, Cloudflare and Google Cloud DNS, then build monitoring that catches latency, outages, bad records and DNSSEC failures.

By the ScreenshotNeo team1 October 20268 min read

Short answer: there is no single best DNS monitoring tool. Use DNSPerf for a free public comparison of provider performance, ThousandEyes for distributed DNS infrastructure and path visibility, Catchpoint for synthetic checks with resolver and retry controls, Cloudflare DNS analytics for activity on Cloudflare zones, and the Google Cloud DNS monitoring dashboard for Cloud DNS private zones. Add origin or HTTP checks when you need to know whether the application itself is reachable.

DNS and origin monitoring answer different questions. A domain can resolve correctly while its web server is down, or the origin can be healthy while users receive bad or missing DNS answers. Cloudflare describes its Health Checks as origin monitoring with uptime, latency and failure-reason analytics. Select tools by the failure layer you need to detect.

What DNS monitoring should measure

A useful DNS monitoring design measures more than one latency number:

  • Availability: did the resolver receive a valid answer before the timeout?
  • Resolution time: how long did the lookup take from the selected vantage point?
  • Record correctness: does the answer contain the expected A, AAAA, CNAME, MX, TXT or other values?
  • DNSSEC: do signatures validate, and are failures reported separately from ordinary NXDOMAIN responses?
  • Path and server health: which resolver, authoritative server or network segment failed?
  • Cache behavior: was the result served from cache, and what TTL and retry policy shaped the observation?
  • Alerting and history: can the team correlate a DNS event with a deployment or an application outage?

Best DNS monitoring tools compared

Tool Best for Documented strengths Limits and questions
DNSPerf Public provider comparisons Free results, selectable geography and period, provider, resolver and root-server views. Tests every minute from 200+ locations over IPv4 with a 1-second timeout. A benchmark snapshot, not a monitor configured for your records, alerts or audience.
ThousandEyes Distributed DNS infrastructure visibility Server, trace and DNSSEC tests; availability, resolution speed, record mapping and DNS-path diagnostics. Confirm agent locations, test types, retention and plan details for your environment.
Catchpoint Synthetic DNS checks Direct and experience tests, DNS resolution time and server availability, with configurable caching and retries. Cache and retry settings change what each result represents.
Cloudflare DNS analytics Investigating Cloudflare-managed zones Query counts, dimensions, average processing time, dashboard and GraphQL access. Processing time excludes parts of the client-to-resolver and resolver-to-authoritative path. History and intervals vary by plan.
Google Cloud DNS monitoring dashboard Google Cloud DNS private zones Queries, error rate, queries per second and 99th-percentile latency charts. Scoped to Cloud DNS private zones; it is not a general external-user monitor.

Which tool should you choose?

  1. Need a quick public comparison? Start with DNSPerf. Select the geography and period that resemble your users and record the timestamp with the result.
  2. Need to locate failures across networks? Use ThousandEyes with agents in the regions, clouds and offices that matter. Add DNSSEC and trace tests when delegation or path failures are possible.
  3. Need control over resolver behavior? Use Catchpoint. Configure caching and retries to match the incident you want to detect, then document those settings with the alert.
  4. Already use Cloudflare DNS? Use its analytics for query activity and provider-side processing. Pair it with external synthetic tests because the dashboard does not represent the whole resolver path.
  5. Use Google Cloud private DNS? The Cloud DNS dashboard supplies service-side charts for those private zones. Add tests from the client networks that consume the zones.
  6. Need application availability? Add an HTTP or origin check. A healthy origin check does not prove DNS is healthy, and a healthy DNS result does not prove the origin responds.

Build a practical DNS monitor yourself

A small scheduled script can catch record changes and lookup failures. Run it from at least two networks or regions, store raw answers, and alert only after a short confirmation window to avoid transient noise.

One-off checks with dig

dig +time=1 +tries=1 A example.com
 dig +time=1 +tries=1 AAAA example.com
 dig +time=1 +tries=1 +dnssec example.com
 dig +time=1 +tries=1 @1.1.1.1 A example.com
 dig +time=1 +tries=1 @8.8.8.8 A example.com

Compare the status, answer section, query time and server address. Test both A and AAAA records; an incorrect IPv6 answer can make only some users fail.

Continuous checks with dnsperf

The open-source dnsperf/resperf tools are load-testing utilities, not a turnkey alerting service. Supply a realistic query input file and interpret packet loss beside latency. Requests with no response may be omitted from latency graphs, which can make a server that drops requests appear faster.

printf 'example.com A\nexample.com AAAA\n' > queries.txt
dnsperf -s 1.1.1.1 -d queries.txt -l 60 -c 1

Use a low query rate for a health check. Use load tests only with authorization and a rate that cannot harm the DNS service.

Python check with record validation

import dns.resolver
import time

name = "example.com"
expected = {"203.0.113.10"}
resolver = dns.resolver.Resolver(configure=True)
resolver.timeout = 1
resolver.lifetime = 2

started = time.perf_counter()
try:
    answer = resolver.resolve(name, "A")
    values = {r.address for r in answer}
    elapsed_ms = (time.perf_counter() - started) * 1000
    if not values & expected:
        raise RuntimeError(f"unexpected A record: {sorted(values)}")
    print({"ok": True, "ms": round(elapsed_ms, 1), "answers": sorted(values)})
except Exception as exc:
    print({"ok": False, "error": str(exc)})
    raise

Node.js check using the built-in resolver

import dns from "node:dns/promises";

const name = "example.com";
const started = performance.now();
try {
  const addresses = await dns.resolve4(name);
  console.log({ ok: true, ms: +(performance.now() - started).toFixed(1), addresses });
} catch (error) {
  console.error({ ok: false, code: error.code, message: error.message });
  process.exitCode = 1;
}

cURL for an HTTP follow-up check

curl --fail --silent --show-error --connect-timeout 3 --max-time 10 \
  -o /dev/null -w 'http=%{http_code} dns_connect=%{time_connect} total=%{time_total}\n' \
  https://example.com/

Keep this check separate from DNS results so an HTTP failure does not get misclassified as a DNS failure.

Configure tests so results mean what you think

  • Vantage points: include user regions, cloud regions, office networks and IPv6 where applicable.
  • Resolvers: test the recursive resolvers users actually use, plus direct authoritative queries when delegation is in scope.
  • Timeouts: record the timeout and retry policy. DNSPerf’s published method uses IPv4 and a one-second timeout.
  • Cache: decide whether you are measuring cached-user experience or a fresh resolution path.
  • Expected answers: validate address sets, CNAME targets, TTL ranges and DNSSEC status, not only a successful response code.
  • Alert windows: require repeated failures from one location or a failure from several locations before paging.
  • Retention: keep enough history to compare incidents with deployments and TTL changes.

How to interpret latency and rankings

DNSPerf reports public benchmark results that update hourly and test every minute from 200+ locations. Its worldwide authoritative-provider page has displayed volatile snapshots such as 12.1 ms for ClouDNS and 12.12 ms for Cloudflare over a stated 30-day period. Treat those values as a dated comparison, not a promise for your users. Geography, resolver choice, protocol, timeout and failed-query handling can change the ranking.

Provider dashboards have the same limitation from another angle. Cloudflare says its processing time is not end-to-end response time because it excludes all client-to-resolver and resolver-to-authoritative-provider time. Google Cloud’s 99th-percentile chart is useful for tail behavior, but it is a metric view rather than a latency guarantee.

Troubleshooting common DNS monitoring failures

Symptom Likely cause Fix
Intermittent NXDOMAIN Different authoritative servers have different zones, or delegation is mid-change. Query each authoritative server directly and compare serials and answers.
High latency only from one region Resolver path, peering or an overloaded local authoritative node. Repeat from another network, run a trace test, and compare recursive versus authoritative timings.
Monitor says healthy but users report outage The monitor used a warm cache or a resolver unlike the affected users. Add relevant recursive resolvers, cold-cache tests and HTTP checks.
False alarms during DNS changes TTL propagation and retry behavior were not included in the alert design. Use a change window, require repeated failures and alert on record correctness after the expected TTL.
DNSSEC failures Broken signatures, expired keys or incorrect DS records. Run DNSSEC-aware tests from multiple locations and inspect the delegation chain.
Latency graph looks unusually fast Timed-out requests are omitted by the load-test visualization. Report success rate and dropped requests beside latency; inspect raw output.
Cloudflare numbers disagree with external tests Cloudflare analytics measures provider-side processing, not the complete resolver path. Use both service-side analytics and external vantage-point tests.

Performance, reliability and cost considerations

Run frequent lightweight checks for critical records and slower, broader tests for every record type. Keep query volume low enough that monitoring cannot become a denial-of-service source. Sample from several locations instead of increasing frequency from one machine. Store raw status, resolver, location, answer, DNSSEC result, latency and retry count so an aggregate percentile remains explainable.

DNSPerf is free for public comparison. ThousandEyes, Catchpoint, Cloudflare and Google Cloud capabilities depend on the service, plan and configuration; verify current limits, retention and pricing before committing. Cloudflare documents plan-dependent history and maximum intervals: Free zone history is 8 days with a 7-day maximum interval; Pro and Business provide 31 days; Enterprise provides 62 days. Confirm current documentation before publishing an operational policy.

Or skip the browser setup

If your DNS workflow also needs visual checks of status pages, dashboards or rendered pages, ScreenshotNeo returns a screenshot or PDF with one GET request. Its API can remove cookie banners, newsletter popups and chat widgets before capture, and bot checks, blank pages and failed loads are not billed. An MCP server lets AI agents take screenshots. See the ScreenshotNeo API docs for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Free accounts include 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Is DNS monitoring the same as uptime monitoring?

No. DNS monitoring checks resolution, answers, DNSSEC and related paths. Uptime monitoring checks the origin or application after DNS resolution.

Should I monitor a recursive resolver or authoritative servers?

Use both when possible. Recursive checks represent user experience; authoritative checks isolate your DNS hosting and delegation.

How many locations are enough?

Start with the regions and networks that serve your users, then add a second independent vantage point for every critical alert.

Can a public DNS ranking identify the best provider for my site?

No. Public rankings are useful context, but your answer depends on geography, resolver paths, records, protocol and failure behavior.

What should an alert contain?

Include the location, resolver, record type, expected and received answers, status, DNSSEC result, latency, timeout and retry count.