How to Download an Entire Website Including Links for Offline Use
Mirror a website for offline browsing with HTTrack or Wget, keep internal links working, seed sitemaps, verify assets, and fix incomplete crawls.
Direct answer: Use HTTrack Website Copier for a guided mirror. Give it the site’s final HTTPS URL, choose Download web site(s), select a local folder, and start. HTTrack crawls reachable resources recursively and rewrites links so the saved copy can be browsed offline. For a scriptable alternative, GNU Wget can recursively download pages and convert links for offline viewing. Neither tool exports server databases, accounts, or live application behavior.
Only mirror sites you are authorized to copy, and check copyright, license, terms, and robots.txt requirements. A mirror contains files exposed to the crawler; it is not automatically a complete backup.
1. What an offline website mirror contains
A mirror is a directory of retrieved HTML, images, stylesheets, scripts, and other files. The crawler discovers URLs from links and downloads them within its scope. HTTrack builds directories recursively and arranges relative links for local browsing; it can also resume an interrupted mirror or update an existing project. Read the HTTrack product documentation.
Offline copies usually cannot reproduce server-side search, logins, shopping carts, databases, personalization, API calls, payment flows, or JavaScript that requires a live origin. Treat those parts as application behavior rather than downloaded files.
2. Before you start: choose the right URL and scope
- Open the site in a browser and note its final URL after redirects. Start with that URL, such as
https://www.example.com/, instead of a bare HTTP address that redirects to another host. - Decide whether you need one section or the whole host. A conservative same-host crawl is safer and smaller.
- Confirm that you have permission to retrieve the pages and assets. Keep HTTrack’s robots.txt behavior enabled unless the site owner has given you different instructions.
- Estimate disk space. Large images, video, downloadable documents, and duplicated query URLs can make a mirror much larger than the visible pages.
3. HTTrack: graphical workflow
- Install HTTrack from its official project page.
- Open it, create a project name, and choose an output directory.
- Enter the final site URL.
- Choose Download web site(s), the normal action for copying a selected site with the current options.
- Review scope, filters, and robots settings. Keep defaults for a first pass.
- Start the mirror and wait for the job to finish.
- Read the log, open the saved index page, and click through representative links while offline.
The HTTrack interface guide recommends checking logs because a project can look complete while still missing resources. Use Continue after a cancelled or crashed job, and Update to recheck a previous project and download changed content.
4. HTTrack: command line
The documented quick-start form is:
httrack https://example.com/ --path mydir
This keeps the default same-host scope and follows links to any depth from the starting location. Put the final redirected host in the command when possible. Consult the HTTrack command-line guide for filters, limits, and scope syntax before widening a crawl.
Seed pages from a sitemap
Link following cannot discover pages that are listed only in a sitemap. HTTrack’s sitemap seeding option is off by default. Enable it when important URLs are not linked from the pages you crawl, then review the resulting scope and filters because a sitemap may contain many URLs.
Keep a crawl inside the intended site
Redirects, language hosts, asset CDNs, and external documentation can cross host boundaries. If those resources are required, allow the specific hosts deliberately; otherwise keep the initial project restricted. A redirect to another host is a common reason that “only the home page came down.”
5. GNU Wget: a scriptable alternative
GNU Wget’s official manual documents recursive downloads, robots.txt behavior, and link conversion for offline viewing. A typical starting command is:
wget \
--recursive \
--page-requisites \
--convert-links \
--adjust-extension \
--no-parent \
--domains example.com \
https://example.com/
--recursive follows links, --page-requisites fetches resources needed by pages, --convert-links changes downloaded links for local viewing, --adjust-extension gives saved documents suitable extensions, and --no-parent prevents climbing above the starting path. Review the current Wget manual for exact flag behavior and add limits appropriate to the site.
6. Verify that the mirror really works offline
- Read the crawler log for failed requests, blocked resources, redirects, and skipped files.
- Disconnect from the network if practical.
- Open the local
index.htmlor project start page. - Follow navigation links several levels deep.
- Check representative images, fonts, stylesheets, scripts, PDFs, and other downloads.
- Try URLs that are known to be important but may not be linked from the home page.
- Record pages that require a live server so users do not mistake them for broken files.
Verification is necessary because a successful completion message does not prove that every asset was retrieved.
7. Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Only the start page appears | The URL redirected to another host, or scope is too narrow. | Start with the final URL and inspect the log. Allow only the required alias or host. |
| Linked pages are missing | Filters exclude them, or they appear only in a sitemap. | Review filters and scope; seed the sitemap explicitly when appropriate. |
| Images or CSS are absent | Resource requests failed, were filtered, or came from another host. | Read the log, check scan rules, and allow required asset hosts deliberately. |
| Links still point online | URLs were not converted or the page builds links in JavaScript. | Use HTTrack’s normal mirror mode or Wget’s link-conversion option; dynamic links may require a live site. |
| Login content is missing | The crawler did not have an authenticated session. | Use documented login or browser-assisted capture options only with authorization. Expect interactive applications to remain partly live. |
| HTTP 403 or refusal | The server rejected the request. This is not the same as robots.txt. | Do not try to evade access controls. Contact the site owner or use an approved export. |
| Very large or never-ending crawl | Calendars, query parameters, faceted navigation, media, or duplicate URLs create an unbounded graph. | Limit paths and hosts, exclude duplicate URL patterns, cap file sizes, and crawl a section first. |
| JavaScript app is blank offline | The page needs APIs, service workers, or server rendering at runtime. | Mirror the required static resources where possible, but plan for a live environment or an application-specific export. |
8. Scope, filters, and edge cases
Redirects and aliases
HTTP-to-HTTPS and bare-domain-to-www redirects can change the host before discovery starts. Use the final destination and include only the aliases you actually need.
Sitemaps and unlinked pages
A crawler sees discovery sources, not a site’s complete URL database. Sitemap seeding helps find URLs that navigation never exposes, but it can also add obsolete or out-of-scope addresses.
Query strings and duplicate content
Tracking parameters, sort orders, filters, calendars, and session IDs can produce many URLs with nearly identical content. Exclude or normalize those patterns where your tool permits it.
Authentication and forms
Credentials and submitted forms can expose private data. Use a dedicated authorized account, avoid saving secrets in project files, and assume that POST actions, CSRF tokens, and stateful workflows will not become reliable offline features.
Media and downloads
Video, archives, and large PDFs dominate storage and bandwidth. Decide whether the offline copy needs them, and apply size or path limits before starting.
Robots.txt and server rules
HTTrack obeys robots.txt by default. Respect the site’s stated rules and access controls; changing a crawler’s robots setting will not make a server-issued 403 legitimate or successful.
9. Performance, reliability, and storage planning
- Start small: mirror one section first to validate scope, filters, and disk usage.
- Prefer resumable projects: HTTrack can continue interrupted work and update an existing cache, avoiding a full restart.
- Throttle responsibly: conservative request rates reduce load and the chance of blocks. Follow the site’s policies.
- Watch the log: repeated timeouts, retries, and redirects usually indicate a site rule or network problem rather than missing link conversion.
- Budget storage: keep space for temporary files and multiple versions if you plan updates. A mirror with media can be many times larger than its HTML.
- Preserve provenance: save the source URL, crawl date, tool version, filters, and log beside the mirror so another person can reproduce or audit it.
HTTrack and Wget are free software. Your practical costs are local disk, bandwidth, time, and any permitted infrastructure used to run the crawl. No source reviewed for this guide publishes a universal completion rate or speed; results depend on site architecture, network conditions, limits, and access rules.
10. Or skip the browser setup
If you need clean visual captures of pages rather than a navigable file mirror, ScreenshotNeo makes one GET request and returns a PNG, JPEG, WebP, or PDF. It is not a replacement for exporting a site’s database or application state, but it avoids maintaining a browser crawler for screenshots.
Cookie and consent banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Create a free ScreenshotNeo account with 1,000 screenshots per month and no card.
11. Short FAQ
Does a mirror include every page?
Only pages reachable through allowed discovery sources, such as links and any sitemap you seed. Private, unlinked, blocked, or dynamically generated URLs may be absent.
Can I mirror a site I do not own?
Get permission and check applicable copyright, license, contract, and robots.txt rules before copying or redistributing it.
Why does the offline copy need internet?
That usually means a script, font, image, API, redirect, or form still points to a live origin, or the application depends on server state.
Should I use HTTrack or Wget?
HTTrack is the clearer guided mirror workflow with documented scope and resume/update features. Wget fits command-line jobs and automation. Test either against the target site’s behavior.
Can ScreenshotNeo create a browsable offline website?
No. It returns individual screenshots or PDFs. Use it when visual snapshots are the deliverable; use a crawler or an approved site export for linked offline browsing.


