How to Download Website Content
Choose the right way to save one page, mirror a site, or download an Internet Archive item—with commands, limits, permissions, and troubleshooting.
The right method depends on your scope: use a browser save or print workflow for one page, HTTrack for a bounded offline mirror, and an archive’s own download controls for Internet Archive items. No method guarantees a complete or authorized copy of every dynamic, protected, or copyrighted resource.
Choose a method by scope
| Goal | Best fit | Main limitation |
|---|---|---|
| Save one page for later | Browser save or print-to-file | Dynamic content and linked assets may not be preserved |
| Browse a bounded site offline | HTTrack | JavaScript-generated links and content can be missed |
| Download an archived item | Internet Archive Download Options | Some items or files are restricted |
| Capture a visual snapshot | ScreenshotNeo | A screenshot is an image or PDF, not a navigable site copy |
1. Save a single web page
For a small number of pages, use your browser’s built-in save or print-to-file workflow. Browser menus and available formats vary by browser and operating system, so confirm the resulting files before relying on them.
- Open the page and wait for its important content to finish loading.
- Use the browser’s save-page or print-to-file command.
- Open the saved result without an internet connection.
- Check images, styles, links, forms, embedded media, and downloadable files separately.
A saved page can contain references to remote resources, session-dependent content, or scripts that no longer work offline. If you need a visual record rather than an interactive copy, export a PDF or image instead.
2. Mirror a website with HTTrack
HTTrack is a free, GPL-licensed offline browser utility. It recursively retrieves HTML, images, and other files into a local directory, rewrites links for offline browsing, can resume interrupted downloads, and can update an existing mirror.
Basic command
httrack https://example.com/ --path mydir
The default behavior stays on the same host, follows links to any depth, and stores the project, logs, and cache in the output directory. HTTrack documents a default throttle of about 100 KB/s; actual speed varies with configuration and later versions. It identifies itself as HTTrack and obeys robots.txt.
Bound the crawl before starting
- Start with one host and a clearly defined path, such as
https://example.com/docs/. - Exclude account areas, search results, calendars, infinite-scroll endpoints, and query-string combinations unless they are required.
- Set an output directory with enough free space for HTML, images, stylesheets, scripts, fonts, and media.
- Run a small test scope first, then inspect the result before expanding it.
HTTrack discovers links by parsing HTML and CSS. It does not execute JavaScript, so URLs or resources assembled only at runtime can be missed. Broader parsing options may find awkward links present in source, but cannot discover links that do not exist until a script runs.
Inspect the result
Review hts-log.txt and hts-err.txt for refused, redirected, filtered, or failed URLs. Test representative pages while offline:
- Home page and navigation
- Deep pages and directory indexes
- Images, CSS, fonts, and downloads
- Pages that require JavaScript
- Links that leave the original host
A crawl completing without an error does not prove that the site is complete. Compare the discovered URL list with your intended scope and record exclusions.
Update a mirror
HTTrack can update an existing project. Its update reporting classifies files as new, changed, unchanged, or gone, which helps you document what changed between captures.
Use WARC for archival capture
When archival fidelity matters, enable HTTrack’s documented WARC output and retain the WARC alongside the browsable mirror. A rewritten mirror is convenient for local navigation; WARC is an archival capture format. Neither guarantees every dynamic or access-controlled feature.
3. Download an Internet Archive item
The Internet Archive does not make every item downloadable. Restricted books and some collections may have no downloadable files. For an available item:
- Open the item’s page.
- Find the Download Options area.
- Choose one file, download multiple files in a format, or use an available bulk method.
- For large collections, follow the Archive’s documented guidance for
wgetor its command-line tool.
Check the item’s access status and the exact file format before writing automation. An archive download is governed by that item’s availability and rights, not merely by the fact that its page is publicly visible.
4. Or skip the browser setup: capture a clean visual copy
If your goal is a reliable image or PDF of a page rather than an offline, navigable mirror, ScreenshotNeo makes one GET request and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report X-Page-Verdict and X-Billed.
You can also capture full pages with lazy images loaded, select one element by CSS selector, set a viewport or device preset, use dark mode and retina scale, provide custom CSS or JavaScript, click or hide elements, wait for a selector, delay, or network idle, block ads, trackers, requests, or resource types, set headers, cookies, user agents, authorization, timezone, and geolocation, produce PDFs, resize images, cache with a chosen TTL, create signed links, run async jobs with signed webhooks, capture up to 100 URLs per call, and query usage. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Plans include 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.
Permissions, robots.txt, and copyright
Confirm that you are allowed to copy, store, and share the material. HTTrack’s documentation places responsibility for copying on the user. A public URL is not automatically permission to reproduce or redistribute its contents.
Google Search Central explains that “A robots.txt file tells search engine crawlers which URLs the crawler can access on your site.” Treat robots.txt as crawler guidance and traffic management, not a license to copy, a security barrier, or a replacement for authentication. Rules may be interpreted differently by crawlers.
Website writing, artwork, and photographs may be protected by copyright. The U.S. Copyright Office FAQ discusses protection and explains that the section 117 archival-copy provision concerns computer programs under specific conditions; it does not create a general archival exception for other website works. Local law and the site’s terms may differ.
Troubleshooting
The mirror finishes but pages are missing
Cause: links were generated by JavaScript, filtered by scope rules, blocked, or hosted on another domain.
Fix: inspect hts-log.txt and hts-err.txt, verify filters and host boundaries, and identify runtime URLs separately. HTTrack cannot discover links that do not exist in fetched HTML or CSS.
Images or styles work online but not offline
Cause: the page references remote, authenticated, lazy-loaded, or script-injected resources.
Fix: test offline, check the logs, and decide whether a visual capture is a better fit than a navigable mirror.
The crawl is too slow or large
Cause: broad scope, duplicate query URLs, media files, or the documented default throttle of about 100 KB/s.
Fix: narrow paths, exclude unneeded resource types and query patterns, schedule a bounded crawl, and monitor disk usage.
Access is denied
Cause: authentication, rate limits, robots rules, server policy, or an unavailable Archive item.
Fix: obtain permission and credentials where appropriate, reduce request load, respect restrictions, or use the item’s official download controls. Do not attempt to bypass access controls.
An Internet Archive item has no download button
Cause: the item or collection is restricted.
Fix: read the item’s access information and use only an available option; contact the rights holder or Archive support if you need authorized access.
The screenshot response is not the expected file
Cause: an invalid API key, inaccessible URL, timeout, bot check, or non-success response.
Fix: check the HTTP status and X-Page-Verdict/X-Billed headers, increase waits for slow pages, and verify the URL and credentials in the API docs.
Performance, reliability, and cost checklist
- Define the exact scope before downloading.
- Run a small pilot and inspect files offline.
- Keep logs with the output and record the capture date, scope, and configuration.
- Reserve storage for media and duplicate versions.
- Use a browsable mirror for navigation and WARC when archival capture matters.
- Expect dynamic, authenticated, and JavaScript-generated content to need separate handling.
- Respect robots guidance, terms, access controls, copyright, and request load.
- For visual snapshots, use ScreenshotNeo’s caching, waits, blocking, signed links, and verdict headers to make automation predictable.
FAQ
Can I download an entire website?
You can mirror a bounded, accessible portion with a crawler such as HTTrack, but no crawl guarantees every dynamic, private, or externally hosted resource.
Does robots.txt give permission to copy?
No. It communicates crawler access preferences and traffic rules; permission and rights come from the site owner, terms, and applicable law.
What is the difference between a mirror and a screenshot?
A mirror aims to preserve files and offline navigation. A screenshot preserves a rendered visual state as an image or PDF.
Is a completed HTTrack crawl proof that nothing was missed?
No. Review logs, scope rules, and offline behavior, especially for JavaScript-generated links.
When should I use WARC?
Use WARC when you need an archival capture format; use a rewritten mirror when convenient offline browsing is the priority.


