ScreenshotNeo

BlogHow-to

How to Copy a Website for Local Development

Learn how to mirror an authorized website with HTTrack, inspect the result, handle dynamic behavior, and use screenshots when a full local copy is unnecessary.

By the ScreenshotNeo team1 October 20267 min read

Short answer: For an authorized static snapshot, use HTTrack to download the pages and assets into a bounded directory. It recursively follows links, rewrites retained links for local browsing, and can resume or update a mirror. A downloaded HTML file is not automatically a working copy of the original application: server routes, authentication, APIs and JavaScript-driven behavior may need to be rebuilt separately.

1. Decide what “copy” means

Choose the smallest result that meets your development goal:

Goal Best starting point What you still need to verify
Inspect one page’s markup and styles Capture one URL Images, fonts, scripts and relative links
Prototype a section of a site Mirror one directory with crawl limits Cross-page links and shared assets
Browse static pages offline Bounded HTTrack mirror Forms, search, API calls and client-side routes
Reproduce a production application Use the mirror only as an asset reference Backend routes, data, authentication, jobs and third-party services

Only copy a site and content you own or are authorized to reproduce. HTTrack’s command-line guide says the crawler identifies itself and obeys robots.txt by default; that does not replace permission or a review of the site’s terms.

2. Install HTTrack and create a small first mirror

Download the version for your operating system from the official HTTrack documentation or the official project site. The project currently lists HTTrack 3.50-4 dated 2026-09-25; check the download page before installing because release details can change.

Run a first capture against a page or directory you are allowed to copy:

httrack "https://example.com/docs/" -O "./mirror"

The first argument is the starting URL. -O selects the output directory. Use a new, empty directory for the first run so you can tell which files the crawl produced.

If you prefer a graphical workflow, start HTTrack, create a new project, enter the starting URL, choose a local destination, and review the action and limits before starting. The official walkthrough and command-line guide document the same core model: recursive download, retained-link rewriting, scope controls and resumable updates.

3. Bound the crawl before downloading

A narrow scope reduces accidental downloads and makes the result easier to understand.

  1. Start at the exact page or directory needed for your task.
  2. Set a maximum depth or other crawl limits in the wizard or command-line options documented in the HTTrack command-line guide.
  3. Use filters to include the host or path you need and exclude areas such as large downloads, account pages, calendars or search URLs.
  4. Run a small capture, inspect it, then widen the scope only when a required asset or linked page is missing.

Do not assume that a command copied from wget or curl will work in HTTrack. Its options, filters and limits are specific to HTTrack.

4. Inspect the local copy

Open the generated entry page in a browser from the mirror directory. Check all of the following:

  • HTML pages load without network access.
  • Stylesheets, fonts and images appear at the expected sizes.
  • Navigation stays inside the mirror instead of returning to the live site.
  • JavaScript files load and do not immediately fail because an API or module is missing.
  • Forms, search, login, checkout and other interactive paths behave as expected for your development task.

HTTrack rewrites retained links for offline browsing and can preserve relative structure, but the presence of an HTML file does not prove that its scripts, images, styles or interactions were captured correctly.

5. Refresh or resume a mirror

HTTrack documents resuming interrupted downloads and updating an existing mirror. Re-run the project with the same destination when you need to continue or refresh it. Keep the original scope and filters so an update does not unexpectedly expand into a site-wide crawl. Review the log after an update and spot-check pages whose assets or behavior changed.

6. Choose an output format

Ordinary output keeps downloaded files in a directory and is usually easiest to edit in a local project. The manual also documents MIME-HTML and single-file modes. MIME-HTML can keep resource URLs while storing a shared asset once, which may help with a large mirror when a Chromium-family browser is available. Check browser compatibility and your build workflow before selecting a packaged format.

7. Handle dynamic and protected sites realistically

HTTrack’s documented model is a crawler that downloads responses and rewrites links. It does not establish that every modern application can be reproduced offline. Expect extra implementation work when the site depends on:

  • Server-side routes that are generated only after a request.
  • Authenticated state, sessions or signed URLs.
  • Data loaded from APIs after the initial HTML response.
  • Client-side routing that needs a running server fallback.
  • WebSockets, payments, search indexes, queues or other backend services.
  • Consent, bot checks or other interstitials that change what a crawler receives.

For these cases, use the mirror as a reference for markup and assets, then recreate the backend behavior with fixtures or a development API. Test the exact flows your local project must support.

8. Capture one page as a visual reference

If your goal is a design reference rather than editable source files, a screenshot can be faster and more reliable than mirroring every asset. A screenshot records the rendered result at a chosen viewport; it does not reproduce the site’s code or interactions.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP or PDF. Before capture it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for authentication and all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, custom headers and cookies, user-agent and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed links, async jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API and an OpenAPI specification.

An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Parameter names used by other screenshot APIs also work, which can simplify a migration.

There is a free tier of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.

10. Troubleshooting

Symptom Likely cause Fix
Only the entry page downloaded Links were outside the selected scope, blocked by filters or generated by JavaScript Check scope and filters, capture the required directory, then inspect whether links exist in the original HTML.
Images or CSS are missing Asset hosts or paths were excluded, or requests require runtime state Review the log and allowed hosts; fetch a small asset scope separately and verify the original response is publicly reachable.
Navigation jumps to the live site The link was not retained or was outside the mirror Keep the linked path inside the crawl scope or replace that route in your local application.
JavaScript shows API errors The mirror contains scripts but not the backend responses Provide local fixtures or implement the API routes required by the page.
Login or personalized pages are blank Authentication and session state were not reproduced Do not copy private data without authorization; build a local authenticated test state and mock services.
The crawl grows unexpectedly Unbounded links such as search, calendars or generated parameters Stop the run, narrow the starting path, add documented filters and limits, and restart in a clean directory.
The command fails immediately Invalid HTTrack syntax or an installation problem Use the official command-line guide for option names and run the graphical workflow to validate the project settings.

11. Performance, reliability and cost

  • Performance: Narrow scopes finish sooner and produce smaller repositories. A first small capture exposes missing dependencies before you spend time downloading the rest.
  • Reliability: Keep the project directory and logs. HTTrack’s resume and update behavior helps recover from interruptions, but always inspect the resulting pages.
  • Storage: Images, video, archives and duplicated query URLs can dominate disk usage. Add exclusions and limits before broadening a crawl.
  • Reproducibility: Record the starting URL, date, scope, filters and tool version alongside the mirror.
  • Cost: HTTrack is software you run locally, so your practical costs are storage, bandwidth and development time. ScreenshotNeo charges only for clean shots; bot checks, blank pages, timeouts, failed loads and cache hits are not billed.

12. FAQ

Can I copy any public website?

No. Public accessibility does not automatically grant permission to reproduce content. Obtain authorization and review applicable terms.

Will HTTrack copy a WordPress or React site?

It may capture publicly served HTML and assets, but the sources do not guarantee complete support for server routes, APIs or JavaScript-rendered interactions. Test the exact pages and rebuild missing backend behavior.

Is a mirror the same as source code?

No. A mirror is a downloaded set of responses with links adjusted for local browsing. It does not include the original server code, database or deployment configuration.

When should I use a screenshot instead?

Use a screenshot when you need a visual reference, review artifact or PDF and do not need editable site behavior. Use a mirror when you need to inspect and modify downloaded files.

Can I update a mirror later?

Yes. HTTrack documents resuming interrupted work and updating an existing mirror. Reuse the same project settings and inspect changes after each update.

Sources