ScreenshotNeo

BlogHow-to

How to Download an Entire Website With cURL

cURL cannot crawl a whole site by itself. Learn the correct Wget mirror command, custom Python and Node crawlers, limits, and a screenshot shortcut.

By the ScreenshotNeo team1 October 20267 min read

Short answer: cURL cannot download an entire website by itself. The curl project FAQ says it has no built-in recursive operation. cURL can fetch each URL you provide, but discovering links, preventing loops, defining scope, and rewriting links require another program or script. For a conventional static site, GNU Wget is the simplest documented mirror tool.

If your goal is a visual screenshot rather than an offline copy, skip recursive downloading and use ScreenshotNeo to capture a page directly.

What “download an entire website” means

A website archive may include HTML pages, stylesheets, images, fonts, rewritten local links, redirects, and a defined host or path boundary. Plain cURL handles the HTTP transfer for one URL at a time. It does not parse HTML or CSS and enqueue discovered links. The official answer is explicit: “No. curl itself has no code that performs recursive operations, such as those performed by Wget and similar tools.” See the curl FAQ.

Use GNU Wget for a linked static-site mirror

For a permitted site whose pages are linked under the starting URL, begin with:

wget --mirror --convert-links --adjust-extension --page-requisites --no-parent --wait=1 https://example.com/docs/

This is a Wget command, not a recursive cURL command. The GNU Wget manual documents these options:

Option Purpose
--mirror Enables recursive retrieval with infinite depth and timestamping.
--page-requisites Fetches resources needed to render pages, such as stylesheets and inline images.
--convert-links Changes downloaded links to point at local files.
--adjust-extension Helps map downloaded HTML to locally viewable extensions; verify behavior in your Wget version.
--no-parent Prevents traversal above the starting hierarchy.
--wait=1 Waits one second between requests to reduce load.

Run a bounded first pass

mkdir -p site-mirror
cd site-mirror
wget --recursive --level=2 --convert-links --adjust-extension \
  --page-requisites --no-parent --wait=1 \
  https://example.com/docs/

Wget’s normal recursive depth is five; an explicit --level makes the limit visible. Increase it only after checking disk use and URL scope. Replace the level with --mirror only when unlimited recursion is genuinely required.

Keep the crawl in scope

  • Start at the deepest directory you are authorized to archive, such as /docs/.
  • Use host restrictions when links leave the site: --domains=example.com.
  • Exclude unwanted paths with --reject-regex, or include known suffixes with --accept.
  • Use --delete-after only for link checking; it discards the downloaded files.
  • Review robots.txt, terms, and authorization. Wget respects the Robot Exclusion Standard, but that alone does not determine whether your use is permitted.

Where cURL still fits

cURL is useful for individual resources or an explicit URL list:

# One response, saved with the remote filename
curl -L -O 'https://example.com/index.html'

# Several known URLs
curl -L -O 'https://example.com/index.html' \
  -O 'https://example.com/docs/start.html' \
  -O 'https://example.com/assets/site.css'

# URLs listed one per line in urls.txt
curl -L --remote-name-all --config urls.txt

A cURL config file uses one option per line:

url = 'https://example.com/index.html'
url = 'https://example.com/docs/start.html'
remote-name
location

This fetches an explicit list. It does not discover additional links or make the copy internally consistent.

Build a small crawler when you need cURL’s transfer behavior

A custom crawler must define scope, normalization, deduplication, depth, retries, timeouts, and file naming. The following examples are conservative starting points for sites you are allowed to crawl.

Python standard-library crawler

#!/usr/bin/env python3
from collections import deque
from html.parser import HTMLParser
from pathlib import Path
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.request import Request, urlopen

START = 'https://example.com/docs/'
ROOT = urlparse(START)
OUT = Path('mirror')
MAX_DEPTH = 2

class Links(HTMLParser):
    def __init__(self):
        super().__init__(); self.links = []
    def handle_starttag(self, tag, attrs):
        attrs = dict(attrs)
        for key in ('href', 'src'):
            if attrs.get(key): self.links.append(attrs[key])

def in_scope(url):
    p = urlparse(url)
    return p.scheme in ('http', 'https') and p.netloc == ROOT.netloc and p.path.startswith(ROOT.path)

def filename(url):
    path = urlparse(url).path
    if not path or path.endswith('/'): path += 'index.html'
    return OUT / path.lstrip('/')

queue, seen = deque([(START, 0)]), set()
while queue:
    url, depth = queue.popleft()
    url = urldefrag(url)[0]
    if url in seen or not in_scope(url): continue
    seen.add(url)
    try:
        data = urlopen(Request(url, headers={'User-Agent': 'site-archiver/1.0'}), timeout=30).read()
    except Exception as exc:
        print(f'skip {url}: {exc}'); continue
    dest = filename(url); dest.parent.mkdir(parents=True, exist_ok=True); dest.write_bytes(data)
    if depth >= MAX_DEPTH: continue
    parser = Links(); parser.feed(data.decode('utf-8', 'ignore'))
    for link in parser.links:
        child = urljoin(url, link)
        if in_scope(child): queue.append((child, depth + 1))
print(f'saved {len(seen)} URLs')

This parses HTML attributes only. It does not execute JavaScript, rewrite links, parse CSS, authenticate, or guarantee that binary assets are discovered. Use Wget for the standard mirror workflow or add each behavior deliberately.

Node.js crawler with built-in fetch

import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const start = new URL('https://example.com/docs/');
const maxDepth = 2, queue = [[start, 0]], seen = new Set();
const inScope = u => u.protocol.startsWith('http') && u.host === start.host && u.pathname.startsWith(start.pathname);
const target = u => path.join('mirror', u.pathname.endsWith('/') ? u.pathname + 'index.html' : u.pathname);
while (queue.length) {
  const [u, depth] = queue.shift(); u.hash = '';
  if (seen.has(u.href) || !inScope(u)) continue; seen.add(u.href);
  const res = await fetch(u); if (!res.ok) { console.error(res.status, u.href); continue; }
  const bytes = Buffer.from(await res.arrayBuffer()); const file = target(u);
  await mkdir(path.dirname(file), { recursive: true }); await writeFile(file, bytes);
  if (depth >= maxDepth || !res.headers.get('content-type')?.includes('text/html')) continue;
  for (const raw of bytes.toString('utf8').matchAll(/(?:href|src)=["']([^"']+)["']/gi)) {
    try { const child = new URL(raw[1], u); if (inScope(child)) queue.push([child, depth + 1]); } catch {}
  }
}
console.log(`saved ${seen.size} URLs`);

Run this with a recent Node.js release that provides global fetch. Like the Python sample, it is a link collector, not a complete browser or offline-link rewriter.

What recursive downloads miss

  • JavaScript-rendered routes: links or content created after page load are invisible to an HTML/CSS crawler.
  • Authenticated areas: sessions, CSRF tokens, and POST workflows need application-specific code.
  • APIs and infinite scroll: data requested after load is not necessarily represented by a link.
  • Cross-origin assets: fonts, images, or scripts on another host need an explicit allowlist.
  • Query-driven duplicates: calendars, tracking parameters, and search pages can create effectively infinite URL spaces.
  • Robots and rate limits: a crawl can be denied or slowed even when pages are public.

Reliability, performance, and cost controls

  1. Measure the scope first with shallow depth and a narrow path.
  2. Throttle requests with --wait, plus delays and timeouts in scripts.
  3. Keep logs and failed URLs so you can retry only errors.
  4. Watch disk, bandwidth, memory, and CPU. GNU warns recursive retrieval can burden both the remote server and your machine.
  5. Use timestamps or a saved URL set for incremental updates.
  6. Prefer a deterministic host and path allowlist over “the whole domain.”

There is no universal size or speed. The result depends on link count, response sizes, redirects, server limits, and your network.

Troubleshooting

Symptom Cause Fix
Only one page downloads with cURL cURL has no recursive crawler. Use Wget or provide an explicit URL list or script.
Pages look unstyled Assets were not fetched. Add --page-requisites and inspect cross-origin assets.
Links open the live site URLs were not rewritten. Use --convert-links or implement rewriting.
Parent directories are copied Start URL is too high. Start at the allowed subdirectory and keep --no-parent.
Crawl never ends Unlimited depth or query variants. Set --level, reject query paths, and deduplicate normalized URLs.
403, 429, or timeouts Access controls or request rate. Stop, confirm authorization, slow down, and follow site rules.
Important content is missing JavaScript, login, or API-only content. Use an authorized browser workflow or API export.
Disk fills unexpectedly Scope or asset size is larger than expected. Stop, inspect logs, constrain hosts and paths, and set a depth limit.

Or skip the browser setup

If you need a clean image or PDF of a page, recursive mirroring is unnecessary. ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and whether it was billed. An MCP server lets Claude, Cursor, and other MCP clients take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can a single cURL flag mirror a site?

No. Recursive discovery is not built into the cURL CLI.

Is Wget guaranteed to copy every page?

No. It follows discoverable HTML and CSS references. JavaScript-only, authenticated, and API-driven content can remain outside the mirror.

Should I use --mirror immediately?

Usually start with a depth limit and narrow path. Use unlimited recursion only after checking scope and resource impact.

How do I archive a private site?

Only with authorization and an approved authentication method. Login flows may require cookies, tokens, and POST requests.

Does downloading a mirror create screenshots?

No. It saves responses and referenced assets. Use a browser-based capture service when you need rendered pixels.