How to Download All Documents from a Website
Learn how to find and download discoverable website documents with HTTrack or Wget, control scope, handle failures, and verify your archive.
Short answer: use an official export or bulk-download feature when the site provides one. Otherwise, use HTTrack to create a local mirror, or GNU Wget when you want a command-line workflow. Both tools retrieve documents they can discover through links and access within the scope you configure. Neither can guarantee every document on an arbitrary site.
“All” normally means all discoverable, accessible documents in a defined scope. Files behind a login, loaded only after a search or JavaScript request, hosted on another domain, or never linked from a page may be missed.
1. Check for an official download first
- Look for Export, Download all, Archive, Sitemap, API, or Data export in the site interface.
- Check the documentation and account settings for a bulk-download option.
- Use the official export when available; it usually defines the collection more reliably than crawling links.
- Confirm that you are allowed to copy and store the documents. HTTrack states that copying a website is the user’s responsibility, and Wget documents robots.txt behavior.
2. Define what “all documents” means
Write down the boundary before downloading. Useful boundaries include:
| Boundary | Example decision |
|---|---|
| Starting URL | https://example.com/resources/ |
| Hosts | Only example.com, or approved subdomains too |
| File types | PDF only, or PDF, DOCX, XLSX, and ZIP |
| Depth | One directory level, or the complete section |
| Authentication | Public pages only, or an authorized logged-in session |
| Storage | A local directory with enough free space |
Start with the smallest relevant section, inspect the result, and expand only when the URLs and file types are correct. Both HTTrack and Wget provide scope controls; choosing them deliberately prevents an unexpectedly large mirror.
3. Option A: mirror the site with HTTrack
HTTrack is designed to download a website to a local directory, recursively building directories and retrieving HTML, images, and other files. Its documentation also describes resuming interrupted projects and updating an existing mirror.
Graphical workflow
- Install HTTrack for your operating system from the official site.
- Create a new project and choose a local destination.
- Enter the narrowest starting URL that contains the documents.
- Choose the default download action first. Add URL or file filters when you need a tighter scope.
- Start the mirror, then open the local project and inspect its file tree and logs.
- Resume the project if the connection stops. Use the project update action when you need to fetch changes later.
Command-line example
httrack 'https://example.com/resources/' -O './site-archive'
The exact filter syntax depends on your scope. Use the HTTrack manual’s URL filters to include only an approved path or selected extensions. A restrictive starting URL is safer than beginning at a site’s home page.
PDF-only filtering
Use an allow rule for the document path or extension after checking the syntax for your installed version. For example, a project can be constrained to URLs ending in .pdf and to the resources section. Test the rule on a small project before a large run.
4. Option B: recursively retrieve links with GNU Wget
Wget is a non-interactive command-line downloader. For HTTP, its manual explains that it parses retrieved HTML and CSS, downloads linked resources, and follows further HTML, XHTML, or CSS documents recursively.
Download a document section
wget --recursive --level=1 --no-parent --accept='pdf,doc,docx,xls,xlsx,zip' --directory-prefix='./site-archive' 'https://example.com/resources/'
--recursivefollows links.--level=1limits recursion to one level. Increase it only when the section requires deeper pages.--no-parentprevents climbing above the starting directory.--acceptlimits saved files to selected extensions.--directory-prefixchooses the local destination.
For a deeper mirror, remove or increase --level after checking the first result. Wget’s manual documents additional host, domain, timestamp, and recursion controls.
Record URLs before downloading
Use a dry run to review the URLs Wget would request:
wget --recursive --level=1 --no-parent --accept='pdf,doc,docx,xls,xlsx,zip' --spider --server-response 'https://example.com/resources/' 2> wget-review.log
Read wget-review.log, remove out-of-scope hosts or paths from your command, then run the real download.
Resume an interrupted run
wget --continue --recursive --level=1 --no-parent --accept='pdf,doc,docx,xls,xlsx,zip' --directory-prefix='./site-archive' 'https://example.com/resources/'
--continue lets Wget continue partially downloaded files where the server supports range requests. It does not make an inaccessible or dynamically generated document discoverable.
5. A small Python inventory script
If you need a repeatable list of document links before downloading, this script extracts links from one HTML page. It does not crawl JavaScript applications or bypass access controls.
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
start = 'https://example.com/resources/'
allowed_ext = ('.pdf', '.doc', '.docx', '.xls', '.xlsx', '.zip')
response = requests.get(start, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
seen = set()
for tag in soup.find_all('a', href=True):
url = urljoin(start, tag['href'])
parsed = urlparse(url)
if parsed.netloc == urlparse(start).netloc and parsed.path.lower().endswith(allowed_ext):
seen.add(url)
for url in sorted(seen):
print(url)
Install dependencies with python -m pip install requests beautifulsoup4. Treat this as an inventory aid: follow linked index pages separately, and verify the final archive against the site’s own document index.
6. Handle JavaScript, search pages, and external hosts
Recursive tools follow URLs they can discover in retrieved documents. They may miss documents that appear only after a search form submission, client-side rendering, an API call, or an interaction. A document hosted on a separate approved domain also requires an explicit host or domain rule.
- Inspect the page source and browser network panel for document URLs.
- Check for a sitemap or documented API.
- Export search results to a URL list when the site allows it.
- Download that reviewed list with Wget or another approved tool.
- Do not guess hidden URLs or bypass authentication, bot checks, paywalls, or other technical controls.
7. Verify the archive
A successful command only means requests completed. Verify the collection itself:
- Compare the number of local files with the site’s index or export, if one exists.
- Search logs for HTTP errors, redirects, skipped URLs, and denied requests.
- Check extensions and file sizes; an HTML error page can be saved with a document-looking URL.
- Open a sample from each directory and confirm that PDFs and office files are readable.
- Record the source URL, download date, command, and scope for future updates.
- Check free disk space before widening the crawl. An external drive is optional when local space is insufficient.
8. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Only the home page was saved | Recursion was disabled or the starting page has no crawlable links | Enable recursion, choose the document section, and inspect the page source. |
| Many HTML files but no PDFs | The PDFs are loaded by JavaScript, blocked by a login, or hosted elsewhere | Use the site’s export/API, inspect network requests, or obtain authorized direct URLs. |
| The crawl leaves the intended section | The starting URL or filters are too broad | Use --no-parent, a lower depth, host restrictions, and explicit accept rules; test with --spider. |
| 403, 401, or CAPTCHA responses | The server requires authentication or rejects automated access | Use an official authenticated export or ask the site owner. Do not bypass the control. |
| Broken relative links locally | The mirror lacks a required asset or the application constructs URLs dynamically | Check logs and capture the complete approved asset scope, or use the site’s export. |
| Disk fills during the run | The scope includes large assets or too many hosts | Stop, narrow the path/extensions, and move the destination only after checking storage capacity. |
| Downloads stop partway through | Network interruption, server timeout, or rate limiting | Resume with the tool’s continuation/update feature and reduce scope or request rate. |
9. Performance, reliability, and cost
- Performance: narrower paths, lower recursion depth, and selected extensions reduce requests and storage.
- Reliability: keep logs, use resumable projects, and rerun an update rather than deleting a partial archive.
- Server impact: schedule large collections responsibly and follow robots.txt, terms, and site instructions.
- Cost: HTTrack and Wget save files locally, but bandwidth, storage, and any site-specific service charges still apply. The reviewed documentation provides no universal time or file-count estimate.
- Completeness: a recursive mirror is evidence of what was discoverable and accessible under your rules, not proof that a site has no other documents.
10. Or skip the browser setup
If your goal is a visual record or PDF of each page rather than the original downloadable files, ScreenshotNeo can capture pages through one API request. It removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its response identifies the page verdict and billing status in headers. The service also provides an MCP server for AI agents with take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options, including full-page capture, PDF settings, custom headers and cookies, waiting rules, blocking, caching, bulk capture, and signed webhooks.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo is for rendered screenshots and PDFs, not a replacement for downloading original DOCX or XLSX files. It can be useful when you need an auditable visual archive of linked pages. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
11. FAQ
Can I guarantee that I downloaded every document?
No. You can guarantee only the URLs discovered and accessible within your configured scope. Use an authoritative export when completeness matters.
Is Wget better than HTTrack?
HTTrack is purpose-built for an offline site copy and offers graphical and command-line interfaces. Wget is a general command-line downloader with detailed recursion controls. Choose the interface and scope controls that fit your workflow.
Can these tools download documents behind a login?
Only when you have authorized access and configure the session correctly; neither tool should be used to bypass authentication or other controls. An official account export is preferable.
Do I need an external hard drive?
No. Save to any directory with enough free space. An external drive is an optional destination for a large archive.
Why are some files missing even though links appear in the browser?
The browser may create those links with JavaScript, retrieve them from an API, or load them from another host. Inspect the approved request flow and use an export or reviewed URL list.


