ScreenshotNeo

BlogHow-to

How to Download a Website’s HTML Code

Save a page's original HTML with curl, inspect rendered DOM in browser tools, or mirror its assets for offline use.

By the ScreenshotNeo team1 October 20267 min read

There are three different jobs people mean by “download a website’s HTML code”:

  • Save the original HTML response: use curl or another HTTP client.
  • Inspect the page in a browser: use View Source for the server response or developer tools for the rendered DOM.
  • Keep a page usable offline: mirror the HTML and its linked assets instead of saving one file.

For one page, the shortest reliable command is:

curl -o page.html https://example.com/

The -o option writes the response to the filename you choose. This saves the HTML response; it does not automatically download every stylesheet, image, font, script, or linked page.

1. Save the original HTML with curl

curl’s official tutorial documents saving a webpage under an explicit filename with -o.

Basic command

curl -o page.html https://example.com/

Replace the URL and filename:

curl -o product.html https://shop.example/products/widget
curl -o article.html 'https://example.com/article?id=42'

Quote URLs containing &, spaces, or shell metacharacters.

Use the remote filename

Use uppercase -O when you want curl to use the filename from the URL:

curl -O https://example.com/index.html

If the URL ends in a slash or does not provide a useful filename, choose one explicitly with -o.

Follow redirects

Many sites redirect HTTP to HTTPS or redirect old paths. Add -L:

curl -L -o page.html https://example.com/old-path

Without -L, curl may save a small redirect response instead of the destination page.

See headers and status

curl -L -D headers.txt -o page.html https://example.com/
cat headers.txt

For a compact diagnostic without writing the body:

curl -L -I https://example.com/

Check the final status, Content-Type, redirects, compression, and caching headers. A successful TCP connection does not guarantee that the response is HTML.

Preserve compressed responses correctly

Ask the server for compression and let curl decompress it before writing the file:

curl -L --compressed -o page.html https://example.com/

Set a timeout and fail on HTTP errors

curl --fail --show-error --location \
  --connect-timeout 10 --max-time 60 \
  -o page.html https://example.com/

--fail makes HTTP 4xx and 5xx responses errors, --show-error keeps the diagnostic, and the timeout flags prevent a stalled request from running forever.

Send a browser-like user agent when required

curl -L \
  -A 'Mozilla/5.0 (compatible; HTML downloader)' \
  -o page.html https://example.com/

A user agent can change what a server returns, but it does not bypass authentication, bot checks, or access controls. Respect the site’s terms and robots policy.

2. Download HTML with Python

Python’s requests library saves the response body as text. This example follows redirects, checks the status, and writes UTF-8 text.

import requests

url = "https://example.com/"
response = requests.get(
    url,
    headers={"User-Agent": "HTML downloader"},
    timeout=60,
)
response.raise_for_status()

# response.text decodes using the server's declared encoding.
with open("page.html", "w", encoding=response.encoding or "utf-8") as f:
    f.write(response.text)

print(response.status_code, response.url)

Install the dependency with python -m pip install requests. For a binary-safe copy of the exact body, write bytes instead:

with open("page.html", "wb") as f:
    f.write(response.content)

Use response.text when you need decoded text; use response.content when you want the raw response bytes.

3. Download HTML with Node.js

Modern Node.js versions include fetch. This complete script follows the server’s response but does not implement browser JavaScript execution.

import { writeFile } from "node:fs/promises";

const url = "https://example.com/";
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 60_000);

try {
  const response = await fetch(url, {
    redirect: "follow",
    headers: { "user-agent": "HTML downloader" },
    signal: controller.signal,
  });

  if (!response.ok) {
    throw new Error(`HTTP ${response.status} ${response.statusText}`);
  }

  await writeFile("page.html", await response.text(), "utf8");
  console.log(`Saved ${response.url}`);
} finally {
  clearTimeout(timer);
}

Save it as download.mjs and run node download.mjs. The fetch API returns the server response; it does not render a JavaScript application.

4. Inspect HTML in a browser

Downloading a response and inspecting a browser page answer different questions.

Need Use What you see
Original server response View Source / show-source HTML initially returned before scripts and other resources run
Final page structure Developer tools → Elements The DOM after scripts, framework rendering, and browser changes

Google documents this distinction in its guidance on viewing source and rendered HTML. A JavaScript-heavy page can have very little useful markup in View Source while containing a complete interface in the Elements panel.

View Source

  1. Open the page in your browser.
  2. Choose View Source, or enter view-source:https://example.com/ in the address bar.
  3. Save the source with the browser’s save command if you need a local copy.

Inspect the rendered DOM

  1. Open developer tools.
  2. Select the Elements or Inspector panel.
  3. Right-click the root element and choose the copy or save option available in your browser.

The rendered DOM is not necessarily a faithful downloadable source file: browser extensions, user state, JavaScript, consent choices, and runtime data can all affect it.

5. Download a page for offline viewing

One HTML file usually contains links to external resources. To create an offline copy, use a mirroring workflow that rewrites links and downloads dependencies. A typical curl request is insufficient for this job.

Before mirroring, decide whether you need:

  • Only the initial HTML response.
  • The page plus images, CSS, fonts, and scripts.
  • Several pages or an entire site.
  • Authenticated or personalized content.

Mirrors can become large, miss resources loaded by JavaScript, or copy private data. Follow the site’s terms, rate limits, and access rules. Test the result by disconnecting from the network and opening the local entry file.

6. What a downloaded HTML file contains

The file can include elements such as headings, links, forms, inline styles, inline scripts, and references to external resources:

<link rel="stylesheet" href="styles.css">
<img src="images/hero.jpg" alt="...">
<script src="app.js" defer></script>

Those references are not the referenced files themselves. Relative URLs may also stop working when the file is opened from file://. An HTML response can therefore be valid while its offline copy looks unstyled or incomplete.

7. JavaScript-rendered pages and dynamic content

HTTP clients such as curl, Requests, and Node fetch do not run page JavaScript. They receive the server response. If the site renders content in the browser, use developer tools to inspect the final DOM or use a browser automation tool that waits for the application to finish rendering.

Common signs that you downloaded only a shell include a nearly empty root element, script bundles with no visible content, or data that appears only after an API call. Save the initial HTML and inspect the browser’s Network panel to identify the requests that populate the page.

8. Troubleshooting

Symptom Likely cause Fix
Saved file is a redirect or error page Redirects were not followed or the server returned an error Use -L, inspect headers, and use --fail
File is empty or truncated Timeout, interrupted transfer, or server connection failure Set --max-time, retry carefully, and check the exit code
Garbled characters Incorrect character decoding Inspect the response charset; write raw bytes or use the declared encoding
HTML has no visible application Content is rendered by JavaScript Use View Source for the initial response and Elements or browser automation for the rendered DOM
Images and styles are missing offline Only the HTML document was downloaded Use a resource-aware mirroring workflow and verify relative paths
403, 429, or CAPTCHA Access policy, rate limiting, or bot protection Authenticate where permitted, slow requests, or use the site’s approved export/API
Different content than the browser Cookies, headers, location, device, or personalization differ Compare request headers and cookies; do not assume curl reproduces a browser session

9. Reliability, performance, and cost

  • Reliability: record the final URL, status code, timestamp, and response headers alongside the file when reproducibility matters.
  • Performance: reuse connections for batches, set finite timeouts, and avoid downloading assets you do not need.
  • Retries: retry transient network failures with backoff; avoid rapidly retrying 401, 403, 404, or 429 responses.
  • Integrity: calculate a checksum if the file is an input to another build or archival process.
  • Cost: direct HTTP downloads consume your network and server resources. Full mirrors can multiply transfer, storage, and request counts.
curl -L --fail --show-error --retry 3 --retry-delay 2 \
  --connect-timeout 10 --max-time 60 \
  -o page.html https://example.com/

Retries are most useful for temporary connection failures. Keep concurrency low enough to respect the site and your own bandwidth.

10. Or skip the browser setup

If your real goal is a clean visual capture rather than the HTML source, ScreenshotNeo returns a screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to try it without a card.

11. FAQ

Does downloading HTML download the whole website?

No. It downloads one HTTP response. Linked resources and other pages require additional requests or a mirroring tool.

Should I use View Source or Elements?

Use View Source for the original response and Elements for the DOM after JavaScript and browser processing.

Why does curl show less content than my browser?

The browser runs JavaScript and sends cookies and other context. curl normally does none of those things.

Can I download a page that requires login?

Only with authorized credentials or an approved export/API. A public URL alone does not grant access to private content.

What filename should I use?

Use a descriptive name such as page.html with -o; use -O only when the URL’s remote filename is suitable.