ScreenshotNeo

BlogHow-to

How to Search and Browse Archived Pages in ArchiveBox

Find archived pages with ArchiveBox’s CLI, web UI, static index, or API, then open a snapshot and locate the matching text.

By the ScreenshotNeo team4 October 20267 min read

Search an ArchiveBox collection from the command line, web UI, REST API, or a generated static index. Search can match snapshot metadata such as URL, title, timestamp, and tags, as well as archived content through the configured backend. Results identify matching snapshots; they do not promise to highlight the matching paragraph. Open a result’s details page, then use browser find or an external file-search tool to locate the text within the saved files.

ArchiveBox search behavior depends on its version and configured backend. Treat the commands and endpoint below as documented examples, and check your installed version and configuration before changing backend settings. ArchiveBox’s search guide, usage guide, and configuration guide describe the available approaches.

Method Use it when What to expect
CLI You have shell access to the ArchiveBox data directory. Quick searches and scripts; available behavior depends on version and backend.
Web UI You want to search and open results in a browser. Search the index or snapshot list, then open a snapshot’s details.
Static HTML index You generated an index and want to browse it without using the built-in web server. The index can be searched and sorted in a browser; its file icon opens details.
REST API You need to integrate listing/search into another tool. The documented list endpoint accepts a search filter; confirm the response and authentication requirements for your release.
External file search You need to inspect saved files directly or search beyond the configured backend’s capabilities. Results depend on which formats were captured and the tool’s supported file types.

2. Search from the CLI

The search guide documents this command form:

archivebox list --filter-type=search 'text to search'

Replace the quoted phrase with a distinctive word, URL fragment, page title, or text passage. Quote phrases so the shell passes them as one argument. Check the syntax supported by your installed release before relying on a command in a script:

archivebox list --help

The CLI returns matching snapshots, not necessarily the exact text location. If a search is unexpectedly empty, try a shorter distinctive term and check whether your configured backend indexes archived content or only metadata.

3. Search in the web UI or static index

Web UI

  1. Open your ArchiveBox instance and go to its searchable index or snapshot list.
  2. Enter a URL, title, tag, or a phrase expected to appear in archived content.
  3. Review the matching snapshots. Compare their URL and timestamp when the same page was saved more than once.
  4. Open the matching snapshot’s details page, then open an available saved output.

Static HTML index

If you have access to the ArchiveBox data directory, the usage guide documents generating a static index with:

archivebox list --html --with-headers > ./index.html

Open index.html in a browser and use its search and sorting controls. Select the file icon for a result to open its details page. The static index is a browseable listing; it does not replace the archived output files.

4. Search through the REST API

The documented list endpoint uses filter_type=search:

curl -G 'http://localhost:8000/api/v1/list' \
  --data-urlencode 'filter_type=search' \
  --data-urlencode 'q=text to search'

Use the host and port where your ArchiveBox instance is reachable. The dossier identifies the endpoint pattern as /api/v1/list?filter_type=search; query parameter details, authentication, and response shape can vary by release, so consult the API exposed by your installation. Do not assume that the API returns a highlighted passage: its purpose here is to find matching snapshots.

5. Open a snapshot and find the matching passage

A search result is the route to a saved snapshot. ArchiveBox may store multiple output formats, and the available files depend on what was captured and which archiving methods were configured.

  1. Open the result’s details page and confirm its URL and capture timestamp.
  2. Choose an available saved output, such as an HTML page or PDF, if present.
  3. Use the browser’s find command (Ctrl+F on Windows/Linux or Command+F on macOS) to search within that output.
  4. If the match is not visible in the browser, use an external search tool on the relevant snapshot files. Match the search tool to the file types you need to inspect.

A snapshot may not contain every representation of the original page. For example, a PDF or other binary file is searchable only if it was captured and your chosen search backend or external tool can inspect it.

6. Pick a search backend

ArchiveBox documents ripgrep, Sonic, and SQLite FTS5 search options. The selected backend affects search coverage, dependencies, indexing, and operational work. The guides describe differing defaults in different contexts, so inspect your installed version and current configuration instead of assuming one universal default.

Backend Useful when Trade-offs
ripgrep You want a filesystem scan without a separate search index or indexer, especially for a smaller collection. Search can slow as the collection grows. The guide notes it does not search binary files such as PDFs, ebooks, or compressed archives. It can also provide regex search.
Sonic You need indexed search or broader content support described by the search guide. Requires an additional dependency and background worker, plus index configuration and maintenance.
SQLite FTS5 You want a full-text option based on SQLite. The search guide describes it as experimental and notes that it uses an index database and update step. Verify current support before depending on it.

Compare backends against your actual needs: collection size and filesystem speed, required file types, query features such as regex, index storage and refresh needs, and whether you can operate an extra service or worker. Any rough size guidance in project documentation is operational guidance, not a guaranteed performance threshold.

Check configuration before changing it

The configuration guide lists ripgrep, sqlite, and sonic as possible engine values, and says the selected engine is used by the UI and CLI. A documented configuration command for SQLite is:

archivebox config --set SEARCH_BACKEND_ENGINE=sqlite

Use that only if SQLite is the intended backend for your installed version. Read the current search setup instructions for backend-specific dependencies and index-update steps, then verify the effective configuration. Plugin settings and setup requirements can differ by backend.

7. Troubleshooting

Symptom Likely cause What to do
No results for text you can see in a saved page The selected backend may not index that output or file type, or the content index may not be current. Try a metadata term or a shorter phrase. Check backend setup and refresh/index requirements; search the saved file externally if needed.
PDF or ebook content is missing from ripgrep results The search guide says ripgrep does not search binary formats such as PDFs and ebooks. Use a backend with the required content support, if available for your version, or a file-search tool that can inspect that format.
The CLI rejects --filter-type or returns different output CLI options can change across releases. Run archivebox list --help and follow the documentation matching your installed version.
The UI and CLI seem to return different results They may be using different configuration, versions, or search paths. Check the active backend and configuration for the instance and CLI environment; the configuration guide says the selected engine is used by both UI and CLI.
The API returns an error or unexpected data Host, endpoint, query parameters, authentication, or response details may differ by release or deployment. Use the API path and parameters documented for your running instance. Confirm the service address and inspect its API documentation.
A snapshot result does not show the exact matching text Search finds matching snapshots, not a highlighted location. Open the details page and use browser find or an external tool against the saved output.
A details page has no expected file That format may not have been captured, or its archiving method may not have been enabled. Review the files actually listed for the snapshot and the capture configuration used for it.

8. Performance, reliability, and cost considerations

  • Performance: Filesystem scanning avoids maintaining a separate index, but the search guide says scans become slower as collections grow. Indexed approaches add index storage and background work; measure against your archive and workload rather than treating approximate project guidance as a benchmark.
  • Reliability: Searchability depends on what was captured, whether the relevant output exists, and whether the configured backend can read it. Keep the ArchiveBox data and any backend index maintained according to the instructions for your version.
  • Cost: The research material gives no universal hosting or operating cost. A filesystem scan avoids a separate indexing service; indexed backends may require additional resources and operational components. Actual cost depends on your deployment.
  • Version awareness: Search defaults and setup instructions can evolve. Record your ArchiveBox version and verify the active backend before troubleshooting or documenting a workflow.

Or skip the browser setup

If your task is to capture a current page rather than search an existing ArchiveBox collection, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. For example, this cURL request saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

Sign up free for ScreenshotNeo.

FAQ

Does ArchiveBox search show the matching paragraph?

No. Search returns matching snapshots. Open one and use browser find or an external file-search tool to locate a passage.

Can I search a static index without running the ArchiveBox web UI?

Yes. Generate the HTML index, open it in a browser, and use its search and sorting controls. You still need the saved outputs to inspect a snapshot.

Which backend should I use for PDFs?

Do not assume ripgrep will search PDF contents; the search guide says it does not search binary formats such as PDFs. Check the current backend documentation for supported file types or use a suitable external tool.

Why do instructions say different backends are the default?

ArchiveBox documentation can reflect different versions or contexts. Check the installed version and active configuration before relying on a stated default.