ScreenshotNeo

BlogHow-to

How to Fetch a Web Page Programmatically

Learn how to fetch HTML with Python, browser JavaScript, cURL and Node.js, handle CORS and errors, and capture JavaScript-rendered pages.

By the ScreenshotNeo team29 September 20269 min read

How to Fetch a Web Page Programmatically

Fetching a web page programmatically means making an HTTP request, checking the response, and reading its body. For a static page, a GET request is enough. The smallest reliable implementation also sets a timeout, checks the status code, inspects the content type, and handles network and decoding failures.

Use a server-side HTTP client when your application needs to retrieve another origin’s HTML. Use the browser Fetch API when the request is part of a permitted web application and the target server allows your origin with CORS. If the page builds its content with JavaScript, an ordinary HTTP client returns the initial response, not the final rendered page; use a documented data endpoint or an allowed browser automation or rendering service.

1. The request-response pattern

Every implementation follows the same sequence:

A fetch retrieves the response body; rendering requires a browser or rendering service.
A fetch retrieves the response body; rendering requires a browser or rendering service.
  1. Validate and normalize the URL.
  2. Send a GET request with a finite timeout.
  3. Follow redirects according to your client’s policy.
  4. Check the HTTP status before parsing.
  5. Inspect Content-Type and character encoding.
  6. Read the body, while enforcing a response-size limit.
  7. Classify errors so callers can retry only the cases that make sense.

GET requests ask for a representation of a resource. They have no request body and are safe, idempotent and cacheable according to HTTP semantics. Query parameters belong in the URL and must be encoded. Use another method only when the site’s API contract requires it.

2. Fetch HTML with Python

Python’s standard library includes urllib.request, so no dependency is required. A Request object lets you send a truthful User-Agent and other headers. The Python documentation demonstrates the same urlopen pattern used below: Python urllib HOWTO.

from urllib.request import Request, urlopen
from urllib.error import HTTPError, URLError

url = "https://example.org/"
request = Request(url, headers={"User-Agent": "my-fetcher/1.0"})

try:
    with urlopen(request, timeout=10) as response:
        status = response.status
        content_type = response.headers.get("Content-Type", "")
        html_bytes = response.read()
        if status < 200 or status >= 300:
            raise RuntimeError(f"HTTP status {status}")
        if "text/html" not in content_type.lower():
            raise RuntimeError(f"Unexpected content type: {content_type}")
        charset = response.headers.get_content_charset() or "utf-8"
        html = html_bytes.decode(charset, errors="replace")
        print(html)
except HTTPError as exc:
    print(f"HTTP error: {exc.code} {exc.reason}")
except URLError as exc:
    print(f"Network or URL error: {exc.reason}")
except UnicodeDecodeError as exc:
    print(f"Character decoding error: {exc}")

HTTPError covers an HTTP response such as 404 or 500. URLError covers URL and network failures. A response can have a successful transport connection but still be the wrong resource, content type or encoding, so check those separately.

Python with requests

The third-party requests package has a convenient API and connection pooling controls. Set both a connect and read timeout in production, and call raise_for_status() before parsing.

import requests

response = requests.get(
    "https://example.org/",
    headers={"User-Agent": "my-fetcher/1.0"},
    timeout=(5, 15),
)
response.raise_for_status()
content_type = response.headers.get("content-type", "")
if "text/html" not in content_type.lower():
    raise ValueError(f"Unexpected content type: {content_type}")
html = response.text
print(html)

3. Fetch a page in browser JavaScript

The browser Fetch API is promise-based. It resolves to a Response even when the server returns 404 or 504, so you must check ok or status. The API and body readers are documented by MDN’s Fetch guide.

async function fetchPage(url) {
  const response = await fetch(url, { method: "GET" });
  if (!response.ok) {
    throw new Error(`HTTP ${response.status}`);
  }
  const contentType = response.headers.get("content-type") || "";
  if (!contentType.includes("text/html")) {
    throw new Error(`Unexpected content type: ${contentType}`);
  }
  return await response.text();
}

fetchPage("https://example.org/")
  .then(html => console.log(html))
  .catch(error => console.error(error));

Use response.json() for JSON, response.arrayBuffer() for binary data, and response.blob() when you need a browser Blob. A response body is a stream and can normally be consumed once.

4. CORS and the browser boundary

Browser scripts are restricted by the same-origin policy. A cross-origin Fetch request is readable only when the server returns an appropriate Access-Control-Allow-Origin header. A failed preflight, missing header or disallowed method produces a CORS error in the browser even if the server is reachable.

mode: "no-cors" is not a solution for reading another site’s HTML. It usually returns an opaque response whose headers and body are unavailable to JavaScript. When you control the application, use one of these designs:

  • Expose a documented API that permits your origin.
  • Fetch from your own same-origin backend and return only the data your frontend needs.
  • Run the retrieval in a server-side worker where browser CORS rules do not apply.

These choices still require permission to access the target, and you should follow authentication requirements, rate limits, robots guidance and site terms.

5. Fetch with cURL

cURL is useful for debugging headers, redirects and status codes from a shell.

curl --fail-with-body --location --max-time 30 \
  --user-agent "my-fetcher/1.0" \
  --header "Accept: text/html" \
  "https://example.org/" \
  --output page.html

Add --dump-header headers.txt to save response headers. Use --head for a HEAD request when the server supports it, but do not assume a HEAD response proves that a subsequent GET will succeed. --location follows redirects; inspect the final URL when redirects matter.

6. Fetch with Node.js

Modern Node.js provides a promise-based global fetch. This example enforces a timeout with AbortController and limits the response size while reading it.

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 15000);

try {
  const response = await fetch("https://example.org/", {
    headers: {
      "user-agent": "my-fetcher/1.0",
      "accept": "text/html"
    },
    signal: controller.signal
  });
  if (!response.ok) throw new Error(`HTTP ${response.status}`);
  const type = response.headers.get("content-type") || "";
  if (!type.toLowerCase().includes("text/html")) {
    throw new Error(`Unexpected content type: ${type}`);
  }
  const html = await response.text();
  console.log(html);
} finally {
  clearTimeout(timer);
}

For large or untrusted responses, consume the stream incrementally and stop after your configured byte limit instead of calling text() without a cap.

7. Headers, cookies, authentication and URL options

Only send headers the target documents or your application requires. A truthful, identifiable User-Agent helps site operators diagnose traffic; do not impersonate a browser to bypass controls.

Need Implementation Common mistake
Query parameters Use a URL builder or parameter object Concatenating unescaped values
Authentication Use the documented Authorization or cookie mechanism Logging secrets in URLs or errors
Cookies Use a cookie jar on the server Assuming browser cookies exist in a server process
Redirects Enable and limit redirects Following an untrusted redirect indefinitely
Encoding Read the charset from Content-Type when available Assuming every page is UTF-8

Validate schemes and hosts before fetching user-supplied URLs. Restrict your application to https when appropriate, and protect internal networks from server-side request forgery by applying an allowlist or network egress policy.

8. Static HTML versus JavaScript-rendered pages

An HTTP client downloads bytes; it does not execute scripts, create a DOM, run layout, preserve browser storage or click controls. Many applications return a minimal HTML shell and fill it after load. A 200 status therefore does not prove that the visible page has been reproduced.

First inspect the HTML for a documented JSON or GraphQL endpoint. Calling that endpoint is usually faster and more stable than rendering a page. If no permitted endpoint exists, use browser automation or a rendering service that can wait for the application to become ready. Define a readiness condition such as a selector, network idle, or a bounded delay. Treat CAPTCHA and bot checks as access controls, not as errors to bypass.

9. Or skip the browser setup

If your goal is a screenshot or PDF of the finished page, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP or PDF. Its capture pipeline accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off.

Consent banners and overlays can be handled before a rendered screenshot.
Consent banners and overlays can be handled before a rendered screenshot.

See the ScreenshotNeo API documentation for all options. A minimal call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, 12 device presets and custom viewports, retina scale, PDF paper size and margins, custom CSS and JavaScript, clicks, selector or network-idle waits, blocked ads and trackers, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, image resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

Free usage includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.

10. Reliability, performance and cost

Timeouts and retries

Set separate connect and read limits where your client supports them. Retry only transient DNS, connection-reset, 408, 429 and selected 5xx failures. Use exponential backoff with jitter and a maximum attempt count. Do not retry malformed URLs, authentication failures or most 4xx responses without changing the request.

Connection reuse

Reuse a session or HTTP agent for repeated requests. Pooling reduces handshake overhead and keeps throughput predictable. Respect the target’s rate limits and avoid unbounded concurrency.

Memory and response size

Cap bytes before buffering a response. Reject unexpected content types before attempting expensive parsing. Store large bodies on disk or stream them to downstream processing.

Caching

Honor cache headers when possible. An application cache can reduce latency and origin load, but include the URL, relevant headers and authentication context in the cache key. Never share a private response across users.

Rendered capture costs

Rendering requires more work than downloading HTML because a browser must load resources and execute scripts. Keep waits bounded, block resources you do not need, capture one element instead of a full page when that is sufficient, and use a chosen cache TTL for repeated captures. ScreenshotNeo reports whether a response was billed, so your accounting can distinguish clean captures from failed or cached attempts.

11. Troubleshooting checklist

Symptom Likely cause Fix
Browser says “blocked by CORS policy” Target did not authorize your origin Use a permitted API, same-origin backend or server-side fetch.
Fetch resolves but status is 404 or 500 Fetch does not reject on HTTP errors Check response.ok or status before reading.
HTML is a nearly empty shell Content is inserted by JavaScript Find the data endpoint or use permitted rendering.
Python raises HTTPError Server returned an HTTP error status Log the status, handle the class explicitly and avoid blind retries.
Timeout or connection reset Slow origin, network path or overloaded service Set bounded timeouts, retry transient failures with backoff and reduce concurrency.
Garbled accented characters Wrong assumed encoding Read the declared charset and decode with a controlled fallback.
Unexpected JSON, PDF or image URL redirected or content type changed Inspect final URL and Content-Type before parsing.
Screenshot shows a consent banner Capture happened before consent handling or the site uses an unsupported flow Enable the relevant ScreenshotNeo consent step and wait for a stable selector.

12. Production checklist

  • Accept only URL schemes and hosts your application is designed to fetch.
  • Use finite connect, read and total timeouts.
  • Check status, content type, final URL and encoding before parsing.
  • Limit response bytes and concurrency.
  • Use a truthful User-Agent and documented authentication.
  • Reuse connections and apply bounded backoff for transient errors.
  • Respect robots.txt guidance, rate limits and site terms.
  • Record request duration, status, retry count and failure category without logging secrets.
  • For rendered pages, define a readiness condition and a maximum wait.

13. FAQ

Can I fetch any URL from browser JavaScript?

No. The target must permit your origin through CORS, and the request must satisfy browser security rules.

Does a 200 response mean I downloaded what a user sees?

No. It means the server returned a successful response. JavaScript-rendered content may not be present in that body.

Should I use GET or POST?

Use GET for retrieving a representation. Follow the target API’s contract when it requires POST or another method.

How do I avoid hanging workers?

Set finite timeouts, cancel requests that exceed them, cap body size and keep retries bounded.

When is a screenshot API preferable to raw HTML?

Use one when you need the rendered visual result, browser waits, PDF output or consistent consent and popup cleanup rather than source HTML.