ScreenshotNeo

BlogHow-to

How to Download a Website’s HTML and CSS Online

Save a page’s raw HTML with curl, or download its CSS and assets for offline viewing with Wget. Learn what survives and what does not.

By the ScreenshotNeo team1 October 20268 min read

Use curl for one HTML or CSS response. Use GNU Wget with page-requisites and link conversion when you need a page’s linked stylesheets, images, and other assets for offline viewing. Browser DevTools is better when you need to inspect the rendered DOM after JavaScript runs.

These methods retrieve files you are authorized to download, archive, test, or migrate. A static copy does not reproduce server-side code, databases, logins, checkout, APIs, or every JavaScript-generated value.

1. Choose the right method

Method Use it for What you get Main limitation
Browser DevTools Inspecting what the browser received and rendered Rendered DOM, stylesheet requests, runtime resources Manual; the DOM can differ from the original response
curl One known HTML or CSS URL The URL’s HTTP response Does not discover linked assets
Wget page-requisites One page with its CSS and referenced assets HTML, stylesheets, images and other exposed requisites Dynamic or API-created content may be missing
Wget recursive An authorized static mirror or multi-page archive Files within a deliberately bounded crawl Can download far more data than intended

2. Download one HTML document with curl

The simplest raw download is:

curl -L -o page.html https://example.com/

-L follows redirects and -o writes the response to page.html. The curl tutorial documents this output-file pattern: curl tutorial.

Inspect the response before saving

# Show response headers
curl -I -L https://example.com/

# Print the HTML to the terminal
curl -L https://example.com/

# Keep the server's filename when possible
curl -L -O https://example.com/page.html

Check the Content-Type, redirects, authentication requirements, and response status. Curl saves what the server returns; it does not execute JavaScript or build the browser’s final DOM.

3. Download a known CSS file with curl

If the page source or Network panel gives you a stylesheet URL, download it directly:

curl -L -o styles.css https://example.com/css/styles.css

For a stylesheet on another host, use its complete URL. If it is relative, resolve it against the page URL first. Curl will not automatically find every <link rel="stylesheet">, image, font, or CSS url() reference.

4. Download HTML, CSS, and page assets with GNU Wget

For one page intended for local viewing, use Wget’s page-requisites and link-conversion modes:

wget --page-requisites --convert-links --adjust-extension --no-parent https://example.com/page.html
  • --page-requisites retrieves files needed to display the page, including references exposed in HTML and CSS.
  • --convert-links rewrites links so the downloaded page can find local files.
  • --adjust-extension gives downloaded HTML and CSS sensible extensions.
  • --no-parent prevents traversal above the target path.

Wget describes itself as “a free utility for non-interactive download of files from the Web” in its GNU manual. Option behavior can vary by installed version, so run wget --help or read the local manual when a command behaves differently.

Keep the output in a named directory

mkdir -p site-copy
cd site-copy
wget --page-requisites --convert-links --adjust-extension --no-parent https://example.com/page.html

Open the resulting HTML file in a browser. If assets are missing, inspect the page source and Network panel for URLs Wget could not discover or access.

5. Create a broader, authorized static mirror

Recursive downloading changes the task from “save one page” to “crawl a site.” Define the host, path, depth, and file types before starting:

wget --recursive --level=2 --page-requisites --convert-links \
  --adjust-extension --no-parent \
  --domains example.com https://example.com/docs/

--level=2 limits link depth, while --domains keeps traversal on the named host. Use a narrow starting path and stop if the output grows unexpectedly. Do not mirror areas you are not authorized to archive.

6. Inspect the rendered page in Chrome DevTools

  1. Open the page in Chrome.
  2. Right-click and choose Inspect.
  3. Use Elements to inspect the DOM after browser parsing and script changes.
  4. Use Network and filter by CSS to find stylesheet responses.
  5. Open a stylesheet request in a new tab and save it if you are authorized to copy it.

The HTML shown in Elements is the rendered DOM, not necessarily the original server response. Search Console’s explanation of how Google receives HTML and Inspect exposes the rendered page is useful background: Google Search documentation.

7. Download with Python

This script saves one HTML response, follows redirects, and raises an error for HTTP failures:

import requests

url = "https://example.com/"
response = requests.get(url, allow_redirects=True, timeout=30)
response.raise_for_status()

with open("page.html", "wb") as file:
    file.write(response.content)

print(response.url, response.headers.get("content-type"))

Download a known stylesheet the same way:

import requests

css_url = "https://example.com/css/styles.css"
response = requests.get(css_url, timeout=30)
response.raise_for_status()
open("styles.css", "wb").write(response.content)

Python requests does not execute page JavaScript or crawl linked resources. For a full asset set, call Wget from your deployment process or implement URL discovery and downloading with the same scope controls described above.

8. Download with Node.js

Node.js 18 or newer includes fetch:

import { writeFile } from "node:fs/promises";

const response = await fetch("https://example.com/", { redirect: "follow" });
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
await writeFile("page.html", Buffer.from(await response.arrayBuffer()));
console.log(response.url, response.headers.get("content-type"));

To save a known CSS URL, replace the URL and output filename. Fetch follows redirects only when configured, and it does not discover or download page requisites.

9. Why the offline page looks different

  • Raw versus rendered HTML: curl and Wget save the server response. DevTools Elements shows a DOM modified by parsing and JavaScript.
  • Missing CSS: a stylesheet may be injected by JavaScript, blocked by authentication, or referenced from a URL your downloader did not reach.
  • Relative paths: moving files without preserving directory structure can break href, src, and CSS url() references.
  • Fonts and images: cross-origin restrictions, signed URLs, lazy loading, or CSS-generated assets can prevent retrieval.
  • Dynamic content: API responses, timestamps, personalization, ads, and client-side rendering may not exist in a static copy.
  • Server behavior: logins, checkout, forms, search, uploads, and backend functions require the original service.

Browsers load HTML, construct a DOM, apply CSS, and execute JavaScript in sequence. MDN’s overview of the browser loading process explains why a downloaded response and a live page can diverge.

10. Troubleshooting

Symptom Likely cause Fix
Only a login page downloads The URL requires authentication or a session cookie Use an authorized authenticated session and supply the required cookies or headers; do not bypass access controls.
HTTP 301/302 or an empty file Redirects were not followed or the response was an error page Use -L with curl, inspect headers with -I, and check the final URL.
Styles are missing CSS is injected, cross-origin, blocked, or outside the crawl scope Find the stylesheet request in DevTools, download it directly, or widen Wget scope only when authorized.
Images are broken offline Lazy loading, signed URLs, absolute paths, or un-downloaded CSS assets Check Network requests, preserve directory paths, and verify CSS url() references.
Page is blank offline The app needs JavaScript and API responses Save the raw source for reference, or use a browser capture for a visual snapshot. A static download cannot recreate the backend.
Wget downloads too much Recursive mode has a broad start path or depth Use page-requisites for one page, add --level, set --domains, and use --no-parent.
403 or 429 responses Access policy, rate limiting, or bot protection Respect the site’s policy, slow requests, authenticate where permitted, and avoid attempts to defeat protections.
Character encoding looks wrong Missing or conflicting charset metadata Inspect the HTTP Content-Type and HTML charset declaration; preserve bytes and decode explicitly when processing text.

11. Performance, reliability, and cost

  • Start narrow: one URL with curl is fastest and easiest to audit.
  • Reuse connections: a scripted downloader can keep a session and rate requests instead of opening an unbounded number of connections.
  • Set timeouts: Python, Node.js, and automated jobs should fail predictably rather than wait forever.
  • Retry carefully: retry transient network failures with backoff, but do not repeatedly retry authorization failures or rate limits.
  • Record provenance: keep the source URL, retrieval time, redirect chain, status, and response headers beside an archive.
  • Budget storage: recursive mirrors can grow quickly because pages share fonts, images, scripts, and query-string variants.
  • Expect change: a later download may differ because the server, JavaScript bundle, personalization, or assets changed.

12. Or skip the browser setup

If your goal is a clean visual capture rather than a source archive, ScreenshotNeo returns a screenshot or PDF from one request. It accepts the cookie or consent banner before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports its verdict in X-Page-Verdict and X-Billed headers. It also provides an MCP server for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There are 1,000 free screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

13. Rights and safe use checklist

  • Confirm that you own the site or have permission to download and archive it.
  • Check terms, robots guidance, rate limits, and access requirements.
  • Keep credentials and private cookies out of shared archives and logs.
  • Do not republish copied HTML, CSS, images, fonts, or text without the required rights.
  • Document the URL, date, scope, and purpose of the download.

14. FAQ

Can I download a whole website with curl?

Curl downloads the URL you give it. Use Wget with deliberate recursive limits when you need multiple pages and assets.

Does saving HTML include CSS?

Only if CSS is inline. Linked stylesheets require separate downloads or Wget page-requisites.

Can an offline copy run a website’s backend?

No. Static files do not reproduce databases, server functions, authentication, checkout, or API services.

What is the difference between View Source and Inspect?

View Source exposes the original response. Inspect shows the DOM after parsing and JavaScript changes.

How do I save only one element?

Use DevTools to inspect and copy the element and its relevant styles, or use a capture tool that supports CSS-selector element capture when you need a visual result.