ScreenshotNeo

BlogEngineering

How DNS Resolution Affects Website Scraping

DNS lookups can add hundreds of milliseconds, return stale servers, or fail before HTTP starts. Learn how caching, TTLs and resolver choice affect scrapers.

By the ScreenshotNeo team1 October 20267 min read

DNS resolution happens before a scraper opens a TCP connection or sends an HTTP request. The hostname is converted into one or more IP addresses by a recursive resolver, which may answer from cache or query authoritative DNS servers. A cache hit is usually fast; a cache miss adds network round trips and can vary with geography, packet loss and authoritative-server health.

For reliable scraping, measure DNS separately from TCP, TLS, server response and body transfer; reuse a normal resolver cache; honor TTLs instead of pinning CDN addresses forever; and classify DNS failures independently from HTTP errors.

1. What happens before the HTTP request

  1. Your scraper asks a configured recursive resolver for example.com.
  2. The resolver checks its cache. If the record is absent or expired, it queries the DNS hierarchy and authoritative servers, then stores the answer for its TTL. Google Cloud explains this recursive-resolver process.
  3. The resolver returns an A (IPv4) and/or AAAA (IPv6) address, plus TTL data.
  4. The scraper opens a TCP connection, performs TLS when using HTTPS, sends the HTTP request and downloads the response.

DNS time is therefore part of end-to-end page time, but it is not HTTP server time. A lookup timeout, SERVFAIL or NXDOMAIN can occur with no HTTP status at all.

2. Why DNS can make a scraper slow

Cache hits and misses

A shared recursive cache can answer immediately. A miss may require several network round trips. Google Public DNS notes that DNS lookups significantly affect loading speed on pages that reference many domains, and reports 300–400 ms average end-to-end resolution under conditions including packet loss, dead name servers and configuration failures. Treat that figure as an observed, failure-inclusive average, not a universal timeout or guarantee. Google Public DNS performance documentation

Many hostnames per page

A browser-like scraper may resolve the page host, API endpoints, font hosts, image CDNs, analytics domains and ad domains. Serial lookups accumulate; concurrent lookups reduce wall-clock time but increase resolver load.

Resolver geography

The resolver’s network location and the worker’s location affect round-trip time and the CDN address selected. A laptop result is not representative of workers running in another region.

Failures before TCP or TLS

Resolver timeouts and SERVFAIL prevent connection setup. NXDOMAIN means the chosen resolver believes the name does not exist; causes include a typo, an uncreated record, a delegation problem or propagation through caches.

3. TTL, caching and CDN address changes

TTL is the number of seconds a recursive resolver may cache an answer. Longer TTLs reduce repeated DNS traffic and misses; short or zero TTLs expose changes sooner but increase lookup work. RFC 9199 describes TTL as a direct control on cache duration, latency, resilience and CDN server selection: RFC 9199.

CDNs can change the address returned for the same hostname. Cloudflare documents a 300-second (five-minute) TTL for proxied anycast IP changes, while noting that local caches can delay what clients observe: Cloudflare TTL documentation.

Strategy Latency Freshness Operational effect
Reuse normal resolver cache Usually lowest TTL bounded Good default for workers
Force lookup every request Higher and variable Fresher More resolver traffic and failure opportunities
Pin an IP indefinitely No repeated lookup Can become stale Breaks CDN steering and failover
Serve stale answers Continues during some DNS outages May be old Availability versus freshness trade-off

RFC 8767 defines serve-stale behavior so recursive resolvers can keep using expired data when authoritative servers cannot be reached; its amended TTL definition recommends a 604800-second (seven-day) cap. RFC 8767. Stale data can preserve availability during an outage but send scraping traffic to an old migration target.

4. Instrument DNS separately

Record the resolver address, returned A/AAAA records, observed TTL, error code and timestamp. Also record DNS duration, TCP connect duration, TLS duration, time to first byte and body transfer time.

cURL timing probe

curl -sS -o /dev/null \
  -w 'dns=%{time_namelookup}s connect=%{time_connect}s tls=%{time_appconnect}s ttfb=%{time_starttransfer}s total=%{time_total}s ip=%{remote_ip}\n' \
  https://example.com/

Run it repeatedly from the same region as production workers. Compare a warm run with a new process or a different resolver to expose cache effects.

Python with explicit DNS timing

import socket
import time
import requests

host = "example.com"
start = time.perf_counter()
infos = socket.getaddrinfo(host, 443, type=socket.SOCK_STREAM)
dns_seconds = time.perf_counter() - start
addresses = sorted({item[4][0] for item in infos})

request_start = time.perf_counter()
response = requests.get(f"https://{host}/", timeout=(5, 30))
request_seconds = time.perf_counter() - request_start

print({
    "dns_seconds": round(dns_seconds, 4),
    "addresses": addresses,
    "status": response.status_code,
    "request_seconds": round(request_seconds, 4),
})

Node.js lookup timing

import dns from "node:dns/promises";

const host = "example.com";
const started = performance.now();
const addresses = await dns.lookup(host, { all: true });
const dnsMs = performance.now() - started;

const response = await fetch(`https://${host}/`, { signal: AbortSignal.timeout(30000) });
console.log({ dnsMs, addresses, status: response.status });

5. Resolver choices

Use the resolver supplied by your host or network unless you have a measured reason to change it. A different recursive resolver can change latency, cache state, filtering behavior and CDN answers. Compare from the production network path, not only from a developer workstation.

Conventional DNS versus DNS over HTTPS

DNS over HTTPS (DoH) encrypts DNS queries inside HTTPS. RFC 8484 specifies the transport; encryption does not guarantee lower latency because it adds an HTTPS path and may use a resolver farther away. RFC 8484. Choose DoH for a documented privacy or network-policy requirement, then measure its actual lookup time.

Shared versus per-worker caches

A shared recursive cache avoids duplicate misses across workers. Per-worker caches isolate failures and can be useful for separate tenants, but cold-start bursts create repeated lookups. Keep a bounded cache and let the resolver enforce TTLs rather than writing an unbounded application cache.

6. Scraper configuration checklist

  • Set separate, bounded DNS, connect, TLS, response and overall deadlines.
  • Reuse HTTP sessions and connection pools so DNS and TCP setup are not repeated unnecessarily.
  • Allow both IPv4 and IPv6 unless your environment has a known broken family; measure each path.
  • Do not permanently pin CDN or failover IPs.
  • Honor normal TTL behavior; refresh early only for a documented freshness requirement.
  • Log resolver, records, TTL, error class and region.
  • Retry transient resolver failures with capped exponential backoff and jitter; do not rapidly retry NXDOMAIN.
  • Compare multiple resolvers and authoritative answers during incidents.
  • After an address change, verify the TLS certificate and HTTP Host/SNI behavior.

7. Common errors and fixes

Symptom Likely cause Fix
Lookup timeout Unreachable resolver, packet loss or overloaded network Check resolver reachability, use a bounded retry, and compare from the worker region.
SERVFAIL Authoritative failure, DNSSEC or delegation problem Query another resolver and authoritative servers; fix the zone or delegation.
NXDOMAIN Typo, missing record or negative-cache entry Confirm the hostname and wait for the negative TTL before retrying.
Old CDN or failover server TTL cache or serve-stale answer Check TTL and resolver age; avoid permanent IP pinning and verify authoritative answers.
Large run-to-run variance Different cache state, geography, packet loss or name-server reachability Measure DNS as its own phase and group results by resolver and region.
HTTP appears down but no status exists DNS failed before TCP/TLS Classify DNS errors separately from HTTP status codes.
HTTPS certificate error after DNS change Address does not serve the expected certificate or SNI host Check the certificate and Host/SNI on the returned address; correct routing or wait for propagation.

8. Reliability, performance and cost

DNS caching usually improves throughput and reduces latency, but cache state is part of your failure model. A resolver outage can affect every worker sharing it; isolated caches reduce blast radius while increasing misses. Alert on DNS error rate and lookup latency separately from HTTP metrics.

DNS traffic itself is inexpensive compared with browser rendering and page transfer, but repeated cold lookups consume time and resolver capacity. Reusing a session, batching work by hostname and avoiding needless re-resolution generally lowers compute cost. Do not trade away required freshness: a stale address can produce incorrect pages or missed failover.

9. Or skip the browser setup

If your goal is a reliable image or PDF of a URL, ScreenshotNeo handles browser startup and capture through one request. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed. Its response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. An MCP server lets Claude, Cursor and other MCP clients use take_screenshot, get_page_info and capture_pdf.

See the ScreenshotNeo API documentation for all options. This call returns a WebP image:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page and element captures, device and viewport controls, custom headers and cookies, waits, request blocking, caching, asynchronous jobs, bulk capture and PDF options. 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

10. FAQ

Does DNS caching always make scraping faster?

No. Warm cache answers are faster, but an incorrect or stale answer can route requests poorly or miss failover. Measure both latency and correctness.

Should I resolve every URL before requesting it?

Usually no. Let a normal recursive resolver cache answers and honor TTLs. Resolve early only when you need explicit DNS timing or validation.

Can I use an IP address to avoid DNS?

Only with care. HTTPS needs the original hostname for certificate validation and SNI, and CDNs may return the wrong site or edge when addressed by IP.

Why does changing DNS not immediately affect my scraper?

Recursive and local caches retain the old answer until its TTL expires; serve-stale policies can extend that window during authoritative outages.

Is DoH faster than normal DNS?

Not inherently. DoH encrypts transport; resolver distance and connection overhead determine the measured latency.