How to Download a Website as an HTML File
Save one webpage as HTML, or mirror an entire site with assets, rewritten links, and practical limits explained.
Short answer: For one open page, use your browser’s Save page as… command. In Firefox, choose Web page, HTML only for one HTML file without images, or Web page, complete for an HTML file plus a folder containing its images and other resources. To download many linked pages, use a site mirror such as HTTrack; it crawls URLs, saves assets, and rewrites retained links for offline browsing.
A saved page is a snapshot, not automatically a complete copy of a site’s database, server code, or JavaScript-generated content. Decide first whether you need one document, a page with assets, or a multi-page offline mirror.
1. Save one webpage as an HTML file
Firefox
- Open the page you are allowed to save.
- Open the menu and choose Save Page As.
- Choose a filename ending in
.htmland select a location. - Choose one of the page types:
| Firefox option | What you get | Use it when |
|---|---|---|
| Web page, HTML only | One HTML file without pictures | You need the markup or want the smallest portable file |
| Web page, complete | An HTML file and a companion directory containing pictures and other resources | You want the page to look closer to the original offline |
Mozilla documents these two modes and notes that complete-page saving creates a directory beside the HTML file. It also cautions that the original HTML link structure may not be preserved. See Mozilla’s save-page documentation.
Chrome
In Chrome, open More → Cast, save and share → Save page as…, then select a filename and folder. Google’s help describes this workflow for offline reading; the exact file packaging can vary by browser version. See Chrome’s official help.
Command line: save the response itself
If you only need the server’s returned HTML, use a plain HTTP client. This does not render the page or download its referenced assets:
curl -L "https://example.com/" -o page.html
-L follows redirects. The resulting file may contain relative links to resources that are not on your computer.
2. Choose the right package
The phrase “HTML file” can describe several different results:
- HTML only: one document, with images, CSS, fonts, and scripts still remote or absent.
- HTML plus resources: a document and a companion folder. Keep both together and preserve the generated folder name.
- Embedded single file: a crawler can rewrite supported assets into
data:URLs. This is convenient for sharing one file but can make it large and cannot embed every resource. - MIME archive: an
.mhtor MIME package stores related resources in an archive. It is not an ordinary.htmlfile.
Saving one page does not create a navigable copy of the entire site. Links to pages you did not save will still point to the live web or fail offline.
3. Mirror an entire site with HTTrack
HTTrack copies a website to disk, retrieves linked HTML, images, and other files, and rewrites retained links so the local copy can be browsed offline. It provides WinHTTrack for Windows, WebHTTrack for macOS and Linux/Unix/BSD, and a command-line program. Read the official HTTrack documentation.
Basic same-host mirror
httrack https://example.com/ --path mydir
By default, HTTrack stays on the starting host. Start with the final canonical URL when possible: a redirect from www to the apex domain, or from HTTP to HTTPS, can move the crawl to another host. Configure the allowed host if that redirect is expected.
Limit crawl depth
httrack https://example.com/ --depth=2 --path mydir
HTTrack counts the starting page as level one. A depth of two includes links found on the start page, subject to the crawler’s other filters.
Request one HTML file per page with embedded assets
httrack https://example.com/ --single-file --path mydir
--single-file embeds supported stylesheets, scripts, images, and fonts as data: URIs. Assets above the default 10 MB per-asset limit remain linked. Audio, video, and links from one page to another also remain links. This option creates one self-contained file per saved page; it does not turn a multi-page site into one giant HTML document.
Use a MIME archive when that is the required format
httrack https://example.com/ --mime-html --path mydir
--mime-html produces a MIME-encapsulated archive such as .mht. Treat it as a different archive format, not as ordinary HTML. It can avoid duplicating shared assets when a compatible browser is available.
4. What a mirror can and cannot capture
JavaScript-generated URLs
HTTrack discovers URLs by parsing HTML and CSS; it does not run JavaScript. A route created only after a script executes, an infinite-scroll request, or an API response may be absent. Extended parsing can find links already present in unusual markup, but it cannot discover links that do not exist until runtime.
Logins and private pages
Only mirror content you are authorized to access. HTTrack documents importing a Netscape-format cookies.txt file and capturing browser requests when a login form requires a POST request. Treat exported cookies as credentials: store them securely, do not commit them to a repository, and delete them when finished.
External hosts and filters
Decide whether third-party CDNs, analytics hosts, downloads, and linked subdomains belong in the mirror. Broader host scope increases size and can pull in unrelated content. Narrow filters are safer for reproducible archives.
5. Resume, update, and protect an archive
HTTrack can resume an interrupted download and can update an existing mirror. These modes have different trust assumptions:
- Resume: continues from cached data and does not recheck every page already stored.
- Update: revalidates pages and downloads changes.
- Old-file protection: an update normally removes local files no longer in the current mirror list. If an update might finish incompletely, use
--purge-old=0to protect older files.
Keep the original command, start URL, date, filters, and cookie policy beside the mirror so another person can reproduce or audit it.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The HTML opens but has no images | You selected HTML-only or used a plain HTTP download | Use Firefox’s complete mode, or mirror the required assets with HTTrack |
| Styles are broken offline | CSS, fonts, or relative paths were not saved | Keep the companion resource folder beside the HTML file; check host and asset filters |
| Only the home page was copied | Crawl depth or scope is too narrow, or links are generated by JavaScript | Increase depth carefully, allow the final host, and remember that runtime-only links are invisible to HTTrack |
| Links lead back online | The target page was excluded or never crawled | Include that host/path, or accept that the mirror is partial |
| Login pages are missing | The crawler has no valid session | Use an authorized cookie file or captured request, and keep credentials private |
| An update deleted local files | Those files were absent from the latest mirror list | Restore from backup or use --purge-old=0 before retrying |
| The page looks empty or stale | Content is loaded by scripts, blocked by bot protection, or dependent on an API | Save a rendered capture, export required data separately, or use a screenshot service |
7. Performance, reliability, and cost considerations
- Scope controls speed and storage: depth, host filters, file-type filters, and asset limits determine how much is retrieved.
- Rendering is different from downloading: a browser save records what the browser can serialize; a crawler parses files without executing application JavaScript.
- Repeatability: record the URL, timestamp, browser or HTTrack version, options, and authorization used.
- Backups: keep a copy before running an update, especially when preserving historical files matters.
- Server impact: crawl politely, obey the site’s rules, and avoid sending unnecessary requests.
- Permission: technical ability does not grant permission to copy a site. HTTrack explicitly advises reading its responsible-use guidance before crawling a server you do not own. Read the HTTrack FAQ and responsible-use notes.
8. Or skip the browser setup
If your actual goal is a reliable rendered image or PDF of a URL rather than an editable HTML mirror, ScreenshotNeo provides a single website screenshot API request. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
9. FAQ
Does saving a page give me the site’s source code?
It gives you the HTML delivered to the browser, plus whatever resources you save. It does not include server-side source code, databases, or private APIs.
Can I combine every page into one HTML file?
Not with normal browser saving. HTTrack’s --single-file option creates one file per page with supported assets embedded.
Why does the offline copy behave differently?
Interactive features may depend on JavaScript, network APIs, authentication, service workers, or server-side state that a static copy does not contain.
Which format should I archive?
Use HTML-only for markup, complete-page folders for a visually useful single page, a crawler mirror for linked navigation, and MHT only when an archive format is specifically required.


