ArchiveBox Fails to Capture JavaScript Pages: Troubleshooting Guide
Find out why ArchiveBox misses JavaScript content, how to check Chrome and extractor setup, and which capture settings and outputs to inspect.
When ArchiveBox misses content from a JavaScript page, first check that ArchiveBox can resolve its Chrome runtime and required extractor packages. Then inspect the selected capture plugins, logs, timeouts, page loading behavior, and whether ArchiveBox skipped the URL because it was already indexed. A “capture” can produce several different artifacts, so check which one is missing: rendered DOM, SingleFile HTML, screenshot, PDF, or Wget copy.
Use this sequence to narrow down the cause before changing settings:
- Record
archivebox versionand the error from stdout or the Web UI. - Install or resolve Chrome and relevant Node-based extractors through ArchiveBox.
- Confirm the needed browser plugin is enabled for the snapshot.
- Use log evidence to decide whether a timeout adjustment is relevant.
- Check whether the page needs more time, scrolling, login, or another interaction.
- Force a fresh capture if the URL was previously archived.
- Compare the available output formats to locate the failing stage.
1. Identify exactly what failed
Start by separating “no new snapshot,” “no rendered artifact,” and “an artifact exists but is incomplete.” They point to different causes.
| Symptom | First check |
|---|---|
| No new snapshot or visible processing | Whether the URL is already indexed and only-new behavior skipped it. |
| Chrome extractor error or missing browser artifacts | Chrome resolution, installation, and the selected provider shown by archivebox version. |
| SingleFile, Readability, or Node-related error | ArchiveBox-managed Node and JavaScript extractor packages. |
| Snapshot exists, but content is blank or partial | Which artifact is incomplete and whether the live page needs more loading time, scrolling, authentication, or interaction. |
| All browser captures fail | Normal URL reachability, access restrictions, browser setup, and logs. |
ArchiveBox uses Chrome to run page JavaScript and capture rendered content, but different capture methods save different kinds of output. A successful screenshot does not by itself prove that SingleFile or a Wget clone will also be complete. The official troubleshooting guidance recommends combining methods because not all sites archive effectively with every method. ArchiveBox project site
2. Record the environment and error
From the ArchiveBox data directory, collect the version output and the relevant stdout or Web UI error. Note:
- ArchiveBox version and the selected Chrome provider, version, and projected path.
- Operating system or container context.
- The target URL, redacting any private query values or credentials.
- Which outputs were produced and which are missing or incomplete.
- Whether this URL has been archived before.
- The plugin selection and any timeout messages in the capture logs.
archivebox version
This creates a baseline before you change configuration. Browser and extractor resolution can vary by installed release, so use the version output and documentation matching your installation when interpreting configuration details.
3. Resolve Chrome through ArchiveBox
If logs show a missing browser or Chrome extractor errors, run ArchiveBox’s installer from the data directory, then inspect the result:
archivebox install chrome
archivebox version
ArchiveBox resolves a compatible host browser or a managed build through abxpkg. The version output reports which provider and browser version were selected and the projected path. Use that information to confirm the browser ArchiveBox will use. Avoid substituting an unrelated browser path unless documentation for your installed release directs you to do so.
4. Install Node-based JavaScript extractors when needed
If the error names Node, SingleFile, or Readability, install the packages through ArchiveBox’s dependency manager:
archivebox install node singlefile readability
archivebox version
These JavaScript extractor packages are managed through abxpkg; the documented setup is not a separate global npm installation. Check the resulting version output and recapture after resolving the dependency error.
5. Confirm the relevant plugin is enabled
Chrome is required for browser-backed outputs such as DOM capture, SingleFile, screenshots, and other Chrome plugins. Inspect the capture’s selected plugin list and enabled settings before assuming JavaScript execution itself failed.
ArchiveBox configuration can be set with archivebox config, ArchiveBox.conf, or environment variables. The plugin whitelist can also be used for a targeted capture. Check the settings and plugin availability documented for your installed release; marketplace entries and configuration can change over time.
6. Change timeouts only when logs point to a timeout
TIMEOUT caps the time for one extractor invocation per snapshot. Plugins can also have their own <PLUGIN>_TIMEOUT settings. If the logs show that a slow page or particular extractor exceeded its limit, raise the relevant limit and recapture.
The cited configuration guidance gives 30–3000 seconds as a recommended range and warns that values below 5 seconds can cause Chrome hangs and broader failures. Verify these limits against your installed version. A timeout increase will not fix missing dependencies, an unselected plugin, a blocked request, a login requirement, or a URL that is unreachable.
7. Check delayed, scroll-driven, or interactive content
Inspect the live page in a normal browser and determine when the missing content appears:
- After a short delay: Check whether the capture method supports an appropriate wait setting, and use timeout logs to tell a slow load from a failed one.
- After scrolling: For scroll-driven lists, the marketplace documents an infinite-scroll expansion plugin. Its presence is an option to try, not a guarantee for every site.
- After a known phrase appears: The screenshot plugin documents a wait-for-text option. This applies to that capture path and should not be assumed to configure every extractor.
- After clicking or interacting: Determine whether simple page rendering is sufficient. A page that requires a user action may need a method that supports the interaction.
- Only when signed in: Confirm access using an appropriate browser session. ArchiveBox describes importing a Chrome profile or using a dedicated persona for private content.
If the target itself does not load in an ordinary browser, diagnose reachability or access first. Do not treat an authentication or automation block as a JavaScript timing problem.
8. Force a fresh capture if the URL was already archived
ArchiveBox may skip a URL that is already indexed under its only-new behavior. Use the supported recapture option:
archivebox add --no-only-new URL
Replace URL with the target address. Do not move or delete the archive/ tree to work around deduplication; force a fresh capture through the command instead.
9. Compare outputs to locate the failure
ArchiveBox can produce several artifacts, and each answers a different question:
| Output | What it helps you inspect |
|---|---|
| Rendered DOM HTML | The DOM state captured from the browser after scripts ran. |
| SingleFile HTML | A self-contained HTML artifact; useful for checking that method independently. |
| Screenshot | A visual record of what the browser rendered at capture time. |
| A visual document of the rendered page. | |
| Wget clone | A downloaded copy of the page and its resources, complementary to browser rendering. |
If the screenshot contains the content but a DOM artifact is absent, that suggests a failure in a different part of the capture pipeline than if every Chrome-backed output fails. Treat that as a diagnostic inference, then confirm it against plugin logs. Completeness, portability, and replay behavior can differ by method; a self-contained file, a folder of downloaded resources, and a visual document are not interchangeable.
10. Troubleshooting common errors
| Error or symptom | Likely cause | Action |
|---|---|---|
| Chrome executable or browser not found | ArchiveBox has not resolved or installed a usable browser. | Run archivebox install chrome, then check provider, version, and path with archivebox version. |
| SingleFile, Readability, or Node package missing | ArchiveBox-managed extractor dependencies are absent. | Run archivebox install node singlefile readability, then inspect archivebox version. |
| No new snapshot appears | The URL may already be indexed and skipped. | Recapture with archivebox add --no-only-new URL. |
| One extractor times out | The page or extractor exceeded its configured time limit. | Read the logs, identify the extractor, and adjust its timeout only if the evidence supports it. |
| Page is blank in every artifact | Possible reachability, access, browser setup, or site restriction issue. | Check the URL in a normal browser, inspect errors, confirm authentication if required, then verify runtime and plugin setup. |
| Content is missing only from one artifact | Capture methods have different behavior and outputs. | Compare DOM, SingleFile, screenshot, PDF, and Wget results; consult the relevant plugin log. |
| Content appears only after scrolling or interaction | The page depends on behavior beyond its initial render. | Inspect the page behavior and try a suitable documented plugin or method; do not assume it will work for every site. |
11. Performance, reliability, and cost considerations
Browser-based capture requires starting or using a browser and waiting for the page and extractor to complete. A slow page, delayed content, or an overly short timeout can affect reliability; increasing a limit also means an extractor can occupy resources longer. Use the smallest timeout adjustment supported by the evidence, and capture only the outputs needed for your archival goal when your configuration permits that choice.
There is no cited benchmark or failure rate for JavaScript capture reliability. Site behavior varies, and the ArchiveBox documentation explicitly cautions that not every site works equally well with every method. Preserve more than one suitable artifact when the page matters and one format alone does not meet your replay or portability needs.
12. Keep capture and replay security separate
A page being captured and an archived file being replayed are separate concerns. ArchiveBox warns that dangerous full-replay modes can run archived JavaScript on the same origin as the admin UI and should not be exposed on a public hostname. Do not weaken replay security settings as a shortcut for a capture failure. Diagnose the capture path and keep replay settings appropriate to your deployment.
13. Escalate with reproducible evidence
If the URL is reachable, dependencies are present, and the failure continues, follow the official troubleshooting guidance and report enough detail to reproduce the issue:
- ArchiveBox version and relevant dependency/provider output.
- Operating system or container context.
- Redacted target URL.
- Relevant stdout or Web UI errors and timeout messages.
- Selected plugins and enabled settings.
- Which artifacts were produced, missing, or incomplete.
- Whether the URL was previously archived and whether a forced recapture changed the result.
Redact credentials, private query parameters, and personal content before sharing logs or URLs.
Or skip the browser setup
For a screenshot or PDF without configuring ArchiveBox’s browser and extractors, ScreenshotNeo provides a website screenshot API and MCP server. Its capture process accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. AI agents can use its MCP server with Claude, Cursor, or another MCP client.
One GET request returns a screenshot or PDF. See the ScreenshotNeo API documentation for options and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo offers 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 screenshots. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month, no card required.
FAQ
Does ArchiveBox run JavaScript during capture?
Yes. ArchiveBox uses Chrome during archiving to run page JavaScript and capture the rendered page. A missing artifact can still be caused by setup, plugin selection, loading behavior, access, or the specific output method.
Should I install SingleFile globally with npm?
The documented ArchiveBox setup installs Node and JavaScript extractor packages through archivebox install, using its managed dependency resolution.
Will increasing the global timeout fix a blank page?
Only if logs show that a timeout is the cause. A blank page can also point to reachability, access, browser setup, or plugin selection.
Can I make archived pages replay exactly like the live site?
Not necessarily. Capture formats preserve different things, and replay depends on the saved resources and security settings. Choose outputs based on whether you need a visual record, portable HTML, or a downloaded clone.


