ArchiveBox Chromium Timeout Error: How to Increase the Page Load Timeout
Fix an ArchiveBox Chromium timeout by setting CHROME_TIMEOUT, choosing the right configuration scope, and diagnosing failures a longer timeout cannot fix.
To give ArchiveBox’s Chrome extractor more time, set CHROME_TIMEOUT in seconds. For example, run archivebox config --set CHROME_TIMEOUT=300 to save a five-minute Chrome-specific timeout in the collection. Use TIMEOUT when you want to raise the per-extractor limit generally, and check CRAWL_TIMEOUT separately if the entire crawl is stopping early. These settings have different scopes; raising one does not automatically raise the others.
Choose the timeout that matches the failure
ArchiveBox has more than one timeout setting. Identify whether Chrome itself is being stopped, another extractor is timing out, or the crawl as a whole is reaching its wall-clock limit.
| Setting | What it limits | When to change it |
|---|---|---|
CHROME_TIMEOUT |
Runtime allowed for the Chrome extractor on one snapshot. | A Chrome output such as a screenshot, PDF, or DOM capture is timing out. |
TIMEOUT |
Runtime allowed for one extractor invocation on one snapshot; the documented default is 60 seconds. | Several extractors need more time, or you want a shared default where no extractor-specific override applies. |
CRAWL_TIMEOUT |
Total wall-clock time for a crawl, including snapshots, extractors, retries, and discovery passes. | The complete crawl ends before its work is finished. This does not extend an individual Chrome invocation. |
An extractor-specific timeout takes precedence over the shared TIMEOUT for that extractor. If Chrome is the only failing extractor, change CHROME_TIMEOUT first. The ArchiveBox configuration reference describes TIMEOUT as a per-extractor, per-snapshot runtime limit, gives 60 seconds as the default, and recommends a range of 30 to 3000 seconds; it also warns against values below five seconds. Treat that range as project guidance, not a universal ideal for every site or deployment. ArchiveBox configuration reference
Set CHROME_TIMEOUT
Persist the value with the CLI
archivebox config --set CHROME_TIMEOUT=300
This stores the Chrome-specific setting in the ArchiveBox collection configuration. Choose a value in seconds that allows the pages you archive to finish loading while keeping the runtime acceptable for your collection.
Check the effective value
archivebox config --get CHROME_TIMEOUT
archivebox config
The first command retrieves the specific setting; the second displays configuration. Check the effective value after changing it, especially if the command runs in a container or under an environment that may supply overrides. The configuration documentation covers supported configuration methods.
Set it in ArchiveBox.conf
In the data directory’s existing ArchiveBox.conf, add the value under its existing [ARCHIVING_CONFIG] section:
[ARCHIVING_CONFIG]
CHROME_TIMEOUT=300
If the section already exists, add only the setting to that section rather than creating a duplicate. Keep the value as a number of seconds.
Use an environment variable for one command
env CHROME_TIMEOUT=300 archivebox add 'https://example.com'
This sets the value for that command invocation. It is useful when you want to try a longer timeout for one archive operation without making it the collection’s saved default. ArchiveBox documents CLI, configuration-file, and environment configuration, including Docker workflows. Configuration methods · ArchiveBox project
When to change TIMEOUT or CRAWL_TIMEOUT
If multiple extractors need a longer per-snapshot allowance, set the shared limit:
archivebox config --set TIMEOUT=120
For a command-level default, set TIMEOUT in the environment in the same way as CHROME_TIMEOUT. A Chrome-specific override takes precedence for Chrome, so changing only TIMEOUT may not help if CHROME_TIMEOUT is already set to a lower value.
If the whole crawl is being cut off, inspect CRAWL_TIMEOUT. That limit applies to the total crawl, not just Chrome. Increasing CHROME_TIMEOUT may let an individual snapshot run longer, but the crawl can still stop when its overall time limit is reached. Check the effective configuration and adjust the setting whose scope matches the observed cutoff.
Diagnose the failure before raising the limit again
- Identify the output that failed. Note whether the missing or incomplete result is Chrome DOM HTML, a screenshot, a PDF, or another Chrome-produced output.
- Check the configured values. Inspect
CHROME_TIMEOUTandTIMEOUT, then consider whether an environment or collection configuration is affecting the command. - Change the Chrome-specific limit and retry. If the error indicates Chrome reached its per-extractor limit, increase
CHROME_TIMEOUTto a suitable number of seconds. - Reassess if the result is still wrong. A longer timeout cannot guarantee that a site renders successfully or that every output is generated. Investigate rendering behavior, browser issues, site behavior, and output paths rather than continually increasing the timeout.
A historical ArchiveBox issue reported screenshots and PDFs with a persistent loading spinner alongside missing HTML-related outputs. It is an example of why the symptom may involve rendering or output generation, not proof of a current general defect. ArchiveBox issue tracker
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Chrome still times out at the old limit. | The setting was not saved where this collection runs, an environment value differs, or the effective value was not checked. | Run archivebox config --get CHROME_TIMEOUT in the same collection and execution environment. Set the value through the CLI, the collection’s ArchiveBox.conf, or the command environment as appropriate. |
Changing TIMEOUT does not change Chrome’s limit. |
CHROME_TIMEOUT is an extractor-specific override. |
Set CHROME_TIMEOUT explicitly and inspect both values. |
| Some extractors succeed, but the whole crawl ends early. | The crawl may have reached CRAWL_TIMEOUT. |
Check the crawl-level limit separately; per-extractor settings do not extend total crawl wall-clock time. |
| The command runs longer, but the page remains on a spinner or output is missing. | The site may not finish rendering, or the failure may concern browser behavior or output generation rather than the configured timeout. | Record which output is incomplete and investigate that failure specifically. A timeout increase only changes how long the extractor may run. |
| A configuration edit has no effect in Docker. | The command may use a different collection, configuration file, or environment than expected. | Check the effective value inside the same Docker workflow and confirm that it is using the intended ArchiveBox data directory. The documented configuration mechanisms also apply to Docker. |
Performance, reliability, and cost considerations
A larger per-extractor limit gives slow pages more time, but a timed-out task that is allowed to run longer can also occupy a worker longer. If many pages are slow, total crawl duration can grow; the overall crawl limit remains a separate constraint. Select a value based on the pages you need to capture and the runtime your deployment can allow, then check whether the resulting output is complete.
Timeout changes improve the chance that a slow capture has enough time to finish; they do not guarantee successful rendering or output creation. Keep the scope narrow when the problem is isolated to Chrome, and use the shared or crawl-wide setting only when the observed failure matches that scope.
Or skip the browser setup
If your goal is to capture a webpage rather than operate an archival browser workflow, ScreenshotNeo provides a website screenshot API and MCP server. See the API documentation for its request options. This cURL example saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Equivalent Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with response headers identifying the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The Free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots. Create a free account and get 1,000 screenshots a month with no card.
FAQ
Does CHROME_TIMEOUT override TIMEOUT?
Yes. It is the Chrome-specific extractor override; TIMEOUT is the shared per-extractor limit.
Is CRAWL_TIMEOUT another name for the page load timeout?
No. It limits the complete crawl’s wall-clock runtime, while CHROME_TIMEOUT applies to Chrome on an individual snapshot.
Will setting the timeout to 300 seconds fix every Chromium error?
No. It helps only when Chrome is being stopped by its configured runtime limit. It cannot ensure that a page renders or that an output is generated.
Can I use the same settings with Docker?
Yes. ArchiveBox documents CLI, configuration-file, and environment methods for Docker workflows as well. Verify the effective setting in the collection and environment that run the capture.


