ScreenshotNeo

BlogHow-to

How to Use a Broken Link Checker API to Find Dead Links

Use single checks, queued batches, polling and webhooks to find dead links reliably with a broken link checker API.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: send one URL to a link-checking API for an immediate report, or submit a batch and poll its batch ID (or receive a webhook) until processing finishes. Treat pending as unfinished, inspect redirects and diagnostics instead of labeling every non-200 response as dead, and store the check timestamp with each finding. The documented GOV.UK Link Checker API exposes GET /check for one URI, POST /batch for up to 5,000 URIs, and GET /batch/:id for a later report.

A checker evaluates whether a supplied URI is suitable to link to according to that service’s rules. The result is evidence from one checker at one point in time, not a permanent guarantee that the destination will remain reachable. Status-code checks also do not prove that a page is useful, that a JavaScript-only destination works for real users, or that every soft 404 and bot challenge has been detected.

The GOV.UK documentation describes the single endpoint as one that checks a single link. Its report can include a state, timestamp, errors, warnings, danger details, problem summaries and suggested fixes. Use those fields in your workflow instead of reducing the response to a Boolean.

GOV.UK Link Checker API documentation is the primary reference; its page says it was updated on 18 March 2025 and may become out of date, so verify endpoint details before deploying.

2. Choose single or batch mode

Need Endpoint When to use it
Check one URL now GET /check?uri=... Editor validation, CI checks, or an interactive tool.
Check many URLs POST /batch Site exports, content migrations, and scheduled audits.
Read a submitted batch GET /batch/:id Polling after an asynchronous submission or resuming a job.

The API’s checked_within value is in seconds and defaults to 14,400 (four hours). Send checked_within=0 when you need a fresh check. A synchronous single check can be requested, but the documentation warns that it may be slow or time out.

3. Check one URL

Set LINK_CHECKER_BASE to the base URL from the current API reference. Keeping it configurable avoids hard-coding a possibly changed service address.

cURL

export LINK_CHECKER_BASE="https://YOUR-LINK-CHECKER-HOST"
curl --fail-with-body --get "$LINK_CHECKER_BASE/check" \
  --data-urlencode "uri=https://example.com/docs" \
  --data-urlencode "checked_within=0"

Python

import os
import requests

base = os.environ["LINK_CHECKER_BASE"].rstrip("/")
params = {
    "uri": "https://example.com/docs",
    "checked_within": 0,  # force a fresh check
}
response = requests.get(f"{base}/check", params=params, timeout=60)
response.raise_for_status()
report = response.json()
print(report)

Node.js

const base = process.env.LINK_CHECKER_BASE.replace(/\/$/, "");
const params = new URLSearchParams({
  uri: "https://example.com/docs",
  checked_within: "0"
});

const response = await fetch(`${base}/check?${params}`);
if (!response.ok) {
  throw new Error(`HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.json());

Save the complete report, including its checked timestamp and diagnostic fields. If a response is asynchronous or contains a pending state, do not mark the link healthy; continue with the batch workflow below.

4. Submit and retrieve a batch

Send a JSON object with a uris array. The documented maximum is 5,000 URIs per batch. Optional fields include checked_within, priority, webhook_uri and webhook_secret_token. For large scheduled runs, the documentation recommends low priority so interactive checks are not blocked.

cURL submission and polling

export LINK_CHECKER_BASE="https://YOUR-LINK-CHECKER-HOST"
curl --fail-with-body -X POST "$LINK_CHECKER_BASE/batch" \
  -H "Content-Type: application/json" \
  --data '{
    "uris": [
      "https://example.com/",
      "https://example.com/docs",
      "https://example.com/missing"
    ],
    "checked_within": 0,
    "priority": "low"
  }'

# Copy the returned batch ID into BATCH_ID, then poll:
export BATCH_ID="YOUR_BATCH_ID"
curl --fail-with-body "$LINK_CHECKER_BASE/batch/$BATCH_ID"

Python batch worker

import os
import time
import requests

base = os.environ["LINK_CHECKER_BASE"].rstrip("/")
uris = [
    "https://example.com/",
    "https://example.com/docs",
    "https://example.com/missing",
]

created = requests.post(
    f"{base}/batch",
    json={
        "uris": uris,
        "checked_within": 0,
        "priority": "low",
    },
    timeout=60,
)
created.raise_for_status()
batch = created.json()
batch_id = batch["id"]

while True:
    report_response = requests.get(
        f"{base}/batch/{batch_id}", timeout=60
    )
    report_response.raise_for_status()
    report = report_response.json()
    states = [item.get("state") for item in report.get("links", [])]
    if states and all(state != "pending" for state in states):
        print(report)
        break
    time.sleep(5)

Node.js batch worker

const base = process.env.LINK_CHECKER_BASE.replace(/\/$/, "");
const body = {
  uris: [
    "https://example.com/",
    "https://example.com/docs",
    "https://example.com/missing"
  ],
  checked_within: 0,
  priority: "low"
};

const created = await fetch(`${base}/batch`, {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify(body)
});
if (!created.ok) throw new Error(`Submit failed: ${created.status}`);
const batch = await created.json();

for (;;) {
  const response = await fetch(`${base}/batch/${batch.id}`);
  if (!response.ok) throw new Error(`Poll failed: ${response.status}`);
  const report = await response.json();
  const links = report.links || [];
  if (links.length && links.every(link => link.state !== "pending")) {
    console.log(JSON.stringify(report, null, 2));
    break;
  }
  await new Promise(resolve => setTimeout(resolve, 5000));
}

The documentation describes an in-progress batch response as HTTP 202 and a completed response as HTTP 201. Treat those as documented examples and confirm the current reference before making status codes part of a strict client contract.

5. Use a webhook instead of polling

When submitting a batch, provide webhook_uri and, if you need authenticity verification, webhook_secret_token. The documented callback includes the batch report and an X-LinkCheckerApi-Signature HMAC-SHA1 header when a secret token is configured.

  1. Read the request body as raw bytes before parsing JSON.
  2. Compute the HMAC-SHA1 digest with the shared secret.
  3. Compare your digest with the signature using a constant-time comparison.
  4. Only then parse and persist the report.
  5. Make the handler idempotent because delivery retries can repeat a batch.
# Pseudocode for verification
expected = HMAC_SHA1(secret, raw_request_body)
if constant_time_compare(expected, request.headers["X-LinkCheckerApi-Signature"]):
    report = parse_json(raw_request_body)
    save_once(report.batch_id, report)
else:
    return HTTP 401

6. Interpret states, redirects and diagnostics

State Meaning and action
pending Processing is unfinished. Poll the batch or wait for its webhook.
ok No problem was reported by this check.
caution Review warnings, redirects or other non-fatal concerns.
broken Errors were detected. Inspect the error and suggested fix.
danger The destination was classified as dangerous. Apply your security and editorial policy before retaining the link.

A redirect is not automatically a dead link. The W3C Link Checker documentation describes warnings for HTTP redirects, including directory redirects, and checks fragments. Review the final destination, redirect chain and warning details.

Store at least:

  • the original URI and its referring page or content location;
  • the state and all error, warning, danger and suggestion fields;
  • the checker timestamp;
  • the batch ID or request correlation ID;
  • the action taken and the next review date.

7. Freshness, scale and performance

  • Freshness: use the four-hour default when repeated checks can reuse recent results; use checked_within=0 for releases, migrations or incident response.
  • Batch size: keep each request at or below 5,000 URIs. Split larger exports into numbered batches and persist every ID.
  • Concurrency: prefer a small number of low-priority batches over thousands of individual synchronous calls.
  • Polling: use exponential backoff with a ceiling, stop after a deadline, and resume from the saved batch ID after a worker restart.
  • Timeouts: set client timeouts longer than ordinary API calls for synchronous checks, and make retries idempotent.
  • Rate limits: the reviewed documentation does not establish a universal quota for the GOV.UK service. Read the live service guidance before selecting worker concurrency.

For a broader hosted site audit, Ahrefs documents page-explorer data such as status codes and internal/external link fields, plus HTTP 429 responses when request limits or dynamic throttling apply. That is a site-audit product rather than an identical drop-in replacement for a supplied-URL checker; compare discovery, queueing, freshness, report detail, limits and workflow integration. See the Ahrefs API reference.

8. Reliability and correctness limits

  • A successful response is time-specific. Recheck high-value links before publication and on a schedule.
  • Bot protection, authentication, robots policies, transient DNS failures and network timeouts can produce ambiguous results.
  • Content can be technically reachable but obsolete, misleading or a soft 404. The reviewed API sources do not establish comprehensive semantic detection.
  • Use a browser or a second appropriate method for uncertain, dangerous or business-critical destinations.
  • Never overwrite the previous report without retaining its timestamp; trend data helps distinguish a one-off outage from a persistent break.

9. Troubleshooting common errors

Symptom Likely cause Fix
Every item remains pending The batch is asynchronous or polling is too aggressive. Save the ID, back off between GET /batch/:id calls, and check the webhook configuration.
A redirect is reported as a problem The checker warns about redirects even when the destination works. Inspect the final URL, chain length and warning; update the link when the redirect is permanent or unnecessary.
Fresh results are not returned A non-zero checked_within allowed a cached recent check. Send checked_within=0 for a forced refresh.
Batch submission is rejected The URI list exceeds 5,000 items or the JSON is malformed. Validate URLs, split the list, and send Content-Type: application/json.
Webhook accepted an untrusted report The raw body was parsed before signature verification or compared insecurely. Verify HMAC-SHA1 over the raw body with constant-time comparison before parsing.
HTTP 429 from an audit provider Per-minute or dynamic throttling. Reduce concurrency, add backoff and follow that provider’s quota documentation.
Link is marked healthy but users report failure Transient outage, bot challenge, JavaScript behavior or semantic soft 404. Recheck, capture diagnostics, validate in a browser and route the case for manual review.

10. Cost and operational planning

The research sources do not provide a pricing comparison for the GOV.UK service or Ahrefs. Budget for your own worker, storage, retries and webhook endpoint, then confirm any provider quotas or commercial terms in the current documentation. Reusing recent checks reduces unnecessary work; forcing every URL fresh increases load and may slow a large audit.

11. Or skip the browser setup

Broken-link reports tell you whether a URL responds, but reviewing a suspicious destination often still requires a clean visual capture. ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF. Cookie and consent banners, newsletter popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are not billed; and each response identifies the verdict and billing result in headers. Its MCP server lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf.

See the ScreenshotNeo API documentation for options such as waits, custom headers, cookies, JavaScript, element capture, full-page capture, PDF settings, caching, signed links, async jobs and bulk capture.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Plans include 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

12. FAQ

No. It means the check has not finished. Poll the batch report or process the completion webhook.

Should every redirect be fixed?

No. Inspect the destination and warning. A redirect can be valid, although removing an unnecessary or unstable redirect is usually preferable.

When should I force a fresh check?

Use checked_within=0 for release gates, migrations and incident investigations. Reuse the default four-hour window for routine monitoring.

Can an API prove that a page is still useful?

No. Reachability is different from content quality. Review soft 404s, bot challenges, JavaScript-only pages and high-value destinations manually or with another suitable check.

How do I audit more than 5,000 URLs?

Split the input into batches of no more than 5,000, persist each batch ID, and process them with bounded concurrency and backoff.

Where should I keep results?

Store the original URL, referring location, state, diagnostics, timestamp, batch ID and remediation status so every finding can be reproduced and reviewed later.