ScreenshotNeo

BlogHow-to

How to Download a Website’s HTML, CSS, and JavaScript

Learn how to save one page or mirror an entire website, including HTML, CSS, JavaScript, images, crawl scope, and common download failures.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: For one page, use your browser’s Save page command or download the HTML and linked resources with developer tools. For multiple pages, use a recursive mirror such as HTTrack or GNU Wget. A mirror can rewrite links for offline browsing, but it cannot guarantee a complete copy of a JavaScript application or resources hosted outside your allowed scope.

1. Decide what you need to download

“Download the HTML, CSS, and JavaScript” can mean two different jobs:

Goal Best starting point What you receive
Read one rendered page offline Browser Save Page One HTML file plus a resource folder, depending on browser
Inspect the source and network files Developer tools Individual HTML, CSS, JavaScript, fonts, and images
Mirror linked pages HTTrack or Wget recursive mode A directory tree with downloaded files and, where supported, rewritten links
Capture the visual result Screenshot API or browser automation PNG, JPEG, WebP, or PDF rather than editable source

Saving a page is not the same as copying a site. A recursive tool must discover links, obey its scope, fetch referenced CSS resources, and decide whether external hosts are allowed.

2. Save one page with a browser

  1. Open the final URL in your browser and wait for the page to finish loading.
  2. Choose Save page as (or the equivalent command).
  3. Select a complete-page option if offered. The browser normally creates an HTML file and a companion directory.
  4. Open the saved HTML file locally and check styles, scripts, fonts, images, and links.

This method is convenient for a page that is mostly server-rendered. It may miss resources loaded after the save operation, resources blocked by authentication, or content assembled only after JavaScript runs.

3. Download source files with developer tools

Developer tools help when you need the exact files requested by the browser:

  1. Open DevTools and select the Network panel.
  2. Reload the page with the network log preserved.
  3. Filter by Doc, CSS, JS, Img, or Font.
  4. Open a request and use Save response or Copy as cURL.
  5. Repeat for resources that are actually needed by the page.

The Sources panel is useful for reading loaded files, but viewing a file there does not package a browsable offline copy. “View source” shows the original response HTML; the Elements panel shows the live DOM after scripts have modified it.

4. Download a single page from the command line

Using cURL

curl -L --fail --remote-name https://example.com/

This saves the response body, usually as a file named after the URL. It does not fetch CSS, JavaScript, images, or linked pages.

Using GNU Wget

wget --page-requisites --convert-links --adjust-extension --no-parent https://example.com/

--page-requisites requests resources needed by the page, while --convert-links changes links for local viewing. Review the output directory before sharing it.

Python: save HTML and inspect linked assets

from pathlib import Path
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup

url = "https://example.com/"
out = Path("site-one")
out.mkdir(exist_ok=True)

response = requests.get(url, timeout=30)
response.raise_for_status()
html = response.text
(out / "index.html").write_text(html, encoding=response.encoding or "utf-8")

soup = BeautifulSoup(html, "html.parser")
for tag, attribute in (("link", "href"), ("script", "src")):
    for node in soup.find_all(tag):
        value = node.get(attribute)
        if not value or value.startswith(("data:", "#", "javascript:")):
            continue
        asset_url = urljoin(response.url, value)
        try:
            asset = requests.get(asset_url, timeout=30)
            asset.raise_for_status()
        except requests.RequestException as exc:
            print(f"Skipping {asset_url}: {exc}")
            continue
        name = asset_url.split("?")[0].rstrip("/").split("/")[-1] or "asset"
        (out / name).write_bytes(asset.content)
        print(asset_url)

This small script is intentionally conservative: it downloads only link and script references and does not rewrite URLs, parse CSS url() references, or crawl additional pages. Install the dependencies with python -m pip install requests beautifulsoup4.

Node.js: save the HTML response

import { mkdir, writeFile } from 'node:fs/promises';

const url = 'https://example.com/';
const response = await fetch(url, { redirect: 'follow' });
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
await mkdir('site-one', { recursive: true });
await writeFile('site-one/index.html', await response.text());
console.log(`Saved ${response.url}`);

Node’s built-in fetch gets the document only. Use a crawler or browser automation package when you need to execute scripts or collect every network request.

5. Mirror multiple pages with HTTrack

HTTrack’s documentation describes it as a copier that saves a site to disk and rewrites links so the local copy browses like the original. It can resume interrupted downloads and update an existing project.

Basic same-host mirror

httrack https://example.com/ --path mirror

Limit the crawl depth

httrack https://example.com/ --depth=2 --path mirror

In the documented example, the start page counts as depth one. A depth limit prevents an accidental crawl of the whole site, but it also excludes pages deeper than the limit.

Scope, redirects, and external resources

HTTrack’s default scope is commonly tied to the starting host. If the URL redirects from an apex domain to www, or from HTTP to HTTPS, start with the final URL or explicitly allow the destination host. CSS, fonts, analytics, images, and APIs may live on different hosts; allowing them can improve fidelity while increasing the crawl size.

Use filters and rate controls before widening scope. HTTrack documents robots.txt handling, connection limits, filters, sitemap support, and update behavior in its command-line guide. A broader filter can fetch files you did not intend to copy.

6. Mirror with GNU Wget

GNU Wget supports recursive retrieval and link conversion. It parses HTML and CSS references such as href, src, and CSS url() values.

wget \
  --recursive \
  --page-requisites \
  --convert-links \
  --adjust-extension \
  --no-parent \
  --domains example.com \
  --directory-prefix mirror \
  https://example.com/

--recursive follows links, --no-parent keeps the crawl below the starting path, and --domains constrains hostnames. Consult the official manual for depth, include/exclude, wait, and robots settings. Do not disable robots restrictions or evade an HTTP refusal.

7. Understand what crawlers can and cannot discover

Traditional crawlers discover URLs in HTML links and CSS references. They do not automatically know every route in a single-page application. HTTrack explicitly documents that it does not execute JavaScript, so a URL assembled only at runtime may never be found. Lazy-loaded images and API responses can also be absent.

  • Add known entry points or a sitemap when the tool supports them.
  • Inspect the network log for runtime requests and download those resources separately.
  • Expect authentication, geo restrictions, signed URLs, and session-only content to require a configured session.
  • Keep an untouched copy of an existing mirror before running an update; update behavior can remove files no longer included.

8. Or skip the browser setup

ScreenshotNeo returns a clean screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

9. Troubleshooting missing files and incomplete mirrors

Symptom Likely cause Fix
Only the start page appears The URL redirected to another host or crawl depth is one Start at the final URL, allow the destination host, and raise depth deliberately
Styles or images are missing Assets are on a CDN or another domain Check Network requests and add the required host within your scope
JavaScript application is incomplete The crawler does not execute JavaScript Collect runtime URLs from DevTools, provide a sitemap, or use browser automation
Lazy-loaded media is absent The resource appears only after scrolling or interaction Trigger the interaction in a browser session and save the resulting requests
Links work online but not offline URLs were not converted or point to absolute hosts Use link-conversion options and inspect generated paths
HTTP 403 or 429 The server refused or rate-limited the request Stop, reduce request rate, check permissions, and follow the site’s policies; do not bypass controls
Mirror update removed files The updated crawl no longer includes those files Back up the existing tree before updating and compare manifests
Private page redirects to login Authentication cookies or headers were not supplied Use an authorized session and configure credentials only where permitted

10. Performance, reliability, and cost considerations

  • Scope controls: Restrict hosts, paths, depth, and file types before starting. This reduces bandwidth and prevents accidental crawling.
  • Rate controls: Use conservative delays and connection limits. A fast crawl can trigger throttling or create load for the origin.
  • Resuming: HTTrack projects can resume interrupted downloads. Keep project files and logs so you can diagnose a partial run.
  • Verification: Open representative pages offline, check browser console errors, compare network requests, and record missing status codes.
  • Storage: JavaScript bundles, source maps, fonts, video, and image variants can dominate disk usage. Exclude unnecessary media when the goal is source inspection.
  • Dynamic sites: A static mirror may never reproduce server-side sessions, API data, personalization, payments, or runtime-generated routes.
  • Screenshot cost: ScreenshotNeo bills only clean shots; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Choose caching TTL, viewport, device, full-page, PDF, and blocking options to control output and request volume.

11. Responsible use checklist

  • Confirm you are allowed to copy and reuse the site’s content and code.
  • Respect robots.txt, terms, access controls, copyright, and applicable policies.
  • Identify the owner before crawling a third-party domain.
  • Use a narrow scope and polite rate limits.
  • Do not treat a 403 response or robots rule as an invitation to evade restrictions.
  • Protect downloaded credentials, private data, and session cookies.

HTTrack’s documentation places responsibility for copying a website on the user. The legal result of copying an unspecified site depends on the content, permission, purpose, and jurisdiction.

12. FAQ

Can I download a website’s source without downloading its assets?

Yes. Save the HTML response with cURL, Wget, Python, or Node.js. The page will usually reference online CSS and scripts, so it may not work offline.

Does “view source” include JavaScript-generated HTML?

No. View source shows the original document response. The live DOM can contain nodes inserted after scripts execute.

Can HTTrack copy a React or Vue site?

It can copy discoverable HTML, CSS, and linked files, but its documentation says it does not execute JavaScript. Runtime routes and API data can therefore be missing.

How do I copy pages that are not linked?

Provide their URLs directly, use a supported sitemap source, or collect routes from an authorized crawl. A link-following crawler cannot discover an unlinked page by itself.

Is a screenshot an editable copy of HTML, CSS, and JavaScript?

No. A screenshot or PDF records the rendered appearance. Use browser tools or a crawler when you need source files.