ScreenshotNeo

BlogComparisons

Save Page WE vs WebScrapBook for Archiving Complete Websites

Save Page WE and WebScrapBook capture pages, but that does not make either a whole-site crawler. Compare their formats and workflows, then choose the right tool for your archive.

By the ScreenshotNeo team4 October 20268 min read

Save Page WE and WebScrapBook are browser extensions for capturing rendered pages. Neither is established by the reviewed documentation as a crawler that discovers and archives every page on a website. If you need a complete website mirror, use a crawler with an explicit URL scope and crawl policy; use either extension for selected pages that need browser rendering or manual selection.

The key distinction is complete page versus complete website. Save Page WE focuses on saving a page as displayed into one HTML file. WebScrapBook offers several capture formats and collection features. Choose between them based on how you want to store, organize, and retrieve selected page captures—not on an assumption that either automatically mirrors a site. See the Save Page WE project documentation, WebScrapBook introduction, and the maintainer discussion of whole-site capture.

What “complete website” means here

A page archive records a page that you selected or opened. A website archive must also discover which pages belong in scope, follow links according to rules, and decide how to handle changing URLs, authentication, and dynamic content. Saving one page—even if its resources are embedded—does not show that every other page was found or saved.

The reviewed evidence supports page-level capture workflows for both extensions. WebScrapBook’s maintainer distinguishes capturing the currently rendered page from headless whole-site crawling and points to spider tools for the latter. Therefore, treat these extensions as complements to a crawler when a page needs a browser-rendered capture, not as substitutes for a site crawl.

Save Page WE vs WebScrapBook at a glance

Need Save Page WE WebScrapBook
Capture model A complete page as currently displayed, saved as one HTML file. A rendered page with configurable resource handling and multiple output forms.
Output choices Single HTML file. Folder, ZIP-based HTZ or MAFF archive, or single HTML.
Multiple pages Can save selected tabs or a list of URLs. That does not establish automatic site-wide crawling. Can capture and organize pages into collections. The reviewed evidence does not establish automatic whole-site crawling.
Organization and retrieval The reviewed source describes single-file saving; it does not document WebScrapBook-style collection search. Hierarchical collections, annotations and editing, full-text search, and optional remote access or backend workflow.
Whole-site mirroring Not established as a crawler by the reviewed product description. The maintainer describes current-page capture and points to spider tools for whole-site crawling.

When Save Page WE is the better fit

Choose Save Page WE when your goal is a portable, single HTML artifact for a page you have rendered, or when you want to save a chosen group of open tabs or a list of URLs. A single file is easy to move and open in a browser. It is still a capture of selected pages, not evidence that every linked page on a site was included.

  1. Open or select the page you want to preserve.
  2. Use the extension’s page, selected-tab, or listed-URL workflow that fits your task.
  3. Open the saved HTML file in a browser and verify the content you need, especially content that loads dynamically or from external resources.
  4. Keep the saved artifact with enough context to identify its source and capture date if that matters to your archive.

For important records, verify the output rather than assuming that every interactive feature, login state, or remote resource will work later. The available documentation supports page capture, but does not promise preservation of every resource or future behavior.

When WebScrapBook is the better fit

Choose WebScrapBook when you want to build a collection of page captures and value output choices, annotations, search, or optional backend access. Its formats let you choose between a folder structure and more consolidated files. The WebScrapBook FAQ describes folder capture as the most reliable and broadly usable option.

  1. Capture the rendered pages you want to keep.
  2. Organize them into collections that reflect your project or research.
  3. Choose folder output when broad compatibility and reliability are priorities.
  4. Use HTZ, MAFF, or single-HTML output when a consolidated artifact better suits your workflow, keeping their documented limitations in mind.
  5. Use the collection’s annotation, editing, search, or optional backend workflow where those features help with retrieval and access.

WebScrapBook’s single-file formats have size and merge-capture tradeoffs. Its FAQ also describes limitations of its built-in archive page viewer: it cannot read a very large ZIP archive (around 2 GiB), large files inside an archive (around 400–500 MiB) can exhaust memory, and scripts do not run in that viewer for security reasons. These are viewer-specific documented limitations, not universal capture-size limits or guarantees about other viewers.

How to archive an entire website reliably

If “complete” means every in-scope page, plan a crawler workflow before collecting anything. The following decisions are practical guidance inferred from the whole-site requirement; the product sources establish the need for a crawler but do not prescribe a particular crawler or settings.

  1. Define scope. Decide which hostnames, URL paths, and page types belong in the archive. Explicitly exclude areas you do not want crawled.
  2. Set crawl depth and link rules. Determine how far from starting URLs the crawler may follow links and how it handles query parameters, duplicate URLs, and redirects.
  3. Plan access. Identify whether the target requires authentication, and whether your chosen tool can capture the needed pages in that state.
  4. Check dynamic rendering needs. Some pages may need browser rendering or interaction before their content appears. Select a crawler that supports the rendering your pages require; use a manual browser capture for exceptions if needed.
  5. Choose an archival format. Decide whether you need a browsable local mirror, individual page files, or another format supported by your chosen crawler.
  6. Validate coverage. Compare the crawl results against known important URLs, inspect representative pages, and check links and assets in the resulting archive.
  7. Record the policy. Keep the start URLs, scope, crawl depth, access assumptions, and capture date with the archive so its limits are clear.

A crawler’s results depend on the defined scope and its handling of dynamic pages, access controls, and resources. Do not describe a crawl as complete without defining what “in scope” means and checking the result.

Common mistakes and troubleshooting

Symptom or assumption Likely cause What to do
Only one page appears in the saved result. A page capture saves the selected or rendered page; it does not necessarily discover linked pages. Use a crawler for site-wide discovery. Use the extension for specific pages that need manual or browser-rendered capture.
A saved page is missing content that appeared later. The content may load dynamically or from an external resource after the page was captured. Wait for the content to render before capture, then inspect the saved output. For a site archive, choose a crawler with suitable rendering support or capture the exceptional page manually.
A consolidated WebScrapBook archive is difficult to use. Single-file formats have size or merge-capture tradeoffs, and the built-in viewer has documented limits. Try folder output for reliability and broad compatibility, or use a suitable viewer while respecting its own limitations.
Scripts do not run in WebScrapBook’s built-in archive viewer. The FAQ says scripts are disabled there for security reasons. Do not treat the viewer as a live reproduction of the original site. Use the archive to inspect preserved content, and avoid relying on archived scripts.
A crawl contains too many pages or misses expected pages. The URL scope, depth, or link rules may not match the intended site boundary. Review starting URLs, allowed hosts and paths, depth, and duplicate or parameter handling; then validate against known URLs.
A page behind a login is absent or incomplete. The capture workflow may not have access to the same authenticated session or content state. Confirm the chosen tool’s access workflow and capture the page in the intended state. Treat login-dependent coverage as a separate requirement to validate.

Performance, reliability, and storage considerations

There are no comparable performance benchmarks in the reviewed sources, so choose based on workflow and validate with representative pages. A single-file capture is convenient to move, while a folder archive is the WebScrapBook FAQ’s recommended choice for reliable, broad use. Large archives and large files can run into the built-in viewer’s documented memory and size constraints.

For whole-site work, crawl scope and rendering needs affect the amount of work and storage required. A focused scope avoids collecting unrelated pages; validating a sample helps catch missing dynamic content before relying on the archive. Keep a separate copy of important source URLs and crawl settings so you can understand what the archive covers.

ScreenshotNeo as a complementary capture option

ScreenshotNeo is a website screenshot API and MCP server, not a website crawler or a replacement for an archival collection. Try it first when the immediate need is a clean screenshot of a rendered page—for example, a visual record of a selected page. Its API returns PNG, JPEG, WebP, or PDF, and its cookie-banner, popup, and chat-widget cleanup is designed for cleaner shots. The following request captures one URL; it does not crawl a site. See the ScreenshotNeo API documentation for configuration details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does saving a complete page mean the whole website was archived?

No. A complete rendered page is a different scope from every page on a site. Use a crawler when you need link discovery across a defined site boundary.

Which WebScrapBook format should I start with?

Start with folder output if reliability and broad usability matter most. Choose a consolidated format when its portability benefits fit your use and its limitations are acceptable.

Can either extension preserve the site exactly as it behaves online?

The reviewed sources do not promise preservation of every resource, login state, interaction, or future behavior. Inspect captures that matter to your work.

Can ScreenshotNeo replace a crawler for archiving?

No. ScreenshotNeo captures a requested page as an image or PDF; it does not discover and crawl every page on a website.