ArchiveBox screenshots are blank: common causes and fixes
Find out why an ArchiveBox screenshot is blank, then check the image, Chrome setup, target page, viewport, and extraction logs in order.
A blank ArchiveBox screenshot is a symptom, not a diagnosis. First open the generated screenshot.png and check whether it is missing, empty, a valid uniformly blank image, or a real capture of a login, consent, error, or bot-check page. Then inspect the Chrome or Chromium runtime, the target page’s response, screenshot resolution, and the complete extraction log. These checks help locate which stage failed without assuming that Chromium itself is broken.
1. Identify what “blank” means
ArchiveBox uses headless Chrome to create screenshot PNGs. Before changing configuration, inspect the artifact itself and note its file size and pixel dimensions.
| What you find | What it suggests | Next check |
|---|---|---|
No screenshot.png |
The screenshot extractor may not have produced an artifact, or extraction may have failed or timed out. | Read the full extraction log for browser launch, dependency, or timeout errors. |
| Zero-byte or unreadable file | The output may have been interrupted or not written successfully. | Retry once and compare the new output and log. |
| Valid image, uniformly blank | The browser may have rendered an empty page, or the expected content may not have appeared before capture. | Open the URL in a normal browser and compare page state; inspect site blocking and timing clues. |
| Login, consent, error, or challenge page | The image may accurately show what the browser received. | Check authentication, consent state, site access, and bot checks. |
| Content appears but is cropped or outside the view | The viewport or screenshot resolution may not match the page region you expected. | Inspect RESOLUTION and SCREENSHOT_RESOLUTION. |
Keep the image and relevant extraction output before retrying. The distinction between a missing file and a valid image of an unexpected page is often the fastest way to narrow the problem.
2. Verify the Chrome or Chromium runtime
ArchiveBox’s current installation guidance checks for a compatible host browser and can install a managed Chromium build when needed. Use the resolver workflow and inspect which browser, version, and path ArchiveBox reports:
archivebox install chrome
archivebox version
If you intend to use a host browser, the current Chromium installation documentation describes setting CHROME_BINARY and letting the installer validate and project it. Follow that workflow rather than bypassing browser resolution with an arbitrary executable path. Also check that the required dependencies are installed and that the target links can be visited normally.
See the ArchiveBox Chromium installation guide and its troubleshooting guide.
3. Check what the target site returned
Open the same URL in a normal browser and compare the visible page with the capture. Look for a login wall, consent dialog, access-denied response, bot challenge, or content that loads only after interaction. An empty or blocked page can be faithfully captured as a blank-looking image.
Some sites block automated requests or requests that do not resemble a normal browser; ArchiveBox’s configuration guidance notes that this can result in 403 responses or empty responses and suggests trying a current browser user-agent. Treat this as a site-specific diagnostic, not a guaranteed fix. Do not change identity or access controls to bypass restrictions you are not authorized to pass.
For a page that requires authentication, check whether the Chrome-based extractor has the same authentication state. ArchiveBox documents using personas for authentication; Chrome extractors read profile state from the persona’s Chrome user-data directory. A normal browser session on your workstation does not by itself establish that the extractor has that session.
Reference: ArchiveBox configuration.
4. Inspect screenshot resolution and viewport
ArchiveBox documents RESOLUTION as a width-and-height viewport setting for screenshot and PDF capture, as well as the extractor-specific SCREENSHOT_RESOLUTION. Check the configured values if the screenshot dimensions or visible region are unexpected. The documentation does not establish that an ordinary resolution value causes a blank image, so treat this as a check rather than a presumed cause.
Compare the actual image dimensions with the expected viewport. If the file dimensions are plausible but its content is absent, continue investigating page rendering and timing rather than treating resolution as the diagnosis.
5. Retry the capture and compare evidence
If the URL already exists in the archive, the troubleshooting guide documents this command to intentionally capture it again:
archivebox add --no-only-new https://example.com/
Replace the example URL with the affected URL. Keep the original output and compare the repeated capture and full log. A one-off timeout and a repeatable failure on one URL point to different areas than a failure across every URL.
A successful PDF or DOM capture alongside a blank screenshot narrows the investigation to differences in capture method or browser rendering, but does not prove a specific cause. ArchiveBox notes that sites do not archive equally well with every method; combining methods such as wget, PDFs, and screenshots can help preserve content when one output is incomplete. Avoid manually moving or deleting archive data to force a recapture: the troubleshooting guide warns that this can desynchronize database and snapshot files.
Historical reports describe Chromium timeouts and repeated Docker capture failures on older releases. They are useful examples of errors to look for in logs, not evidence of a current general ArchiveBox defect: issue 1125 and issue 1181.
6. Use a diagnostic matrix
| Comparison | If the results differ | What to investigate |
|---|---|---|
| One URL versus several unrelated URLs | Only one site fails | Site response, authentication, bot checks, consent, or page-specific rendering. |
| Fresh capture versus repeat capture | Only one attempt fails | Intermittent timeout or load timing; preserve both logs. |
| Logged-in browser versus extractor | Extractor shows login | Persona and Chrome profile authentication state. |
| Screenshot versus PDF or DOM output | Other output contains content | Screenshot-specific browser rendering, viewport, or timing; this comparison narrows but does not identify the cause. |
| Missing artifact versus valid blank image | Different artifact state | For missing output, focus on extraction and browser launch logs; for a valid image, focus on what rendered. |
Common errors and fixes
| Symptom or log clue | Likely area | Next action |
|---|---|---|
| Chrome or Chromium executable not found, or browser launch fails | Browser resolution, installation, or dependencies | Run archivebox install chrome, inspect archivebox version, and follow the current install guide. |
| Timeout while producing screenshot | Slow page load, browser startup, or environment-specific delay | Save the complete log, repeat once, and compare other URLs and extractor outputs. Do not assume an old issue report describes your current version. |
| 403 or empty response | Site-specific automated request blocking or access restrictions | Check the URL in a normal browser and inspect the actual response and page state. A current user-agent is one documented diagnostic to consider, not a universal remedy. |
| Login or consent screen captured | Missing extractor authentication or consent state | Check the persona and Chrome profile used by the extractor and compare with the expected page. |
| Wrong size or cropped content | Viewport or screenshot-resolution configuration | Inspect RESOLUTION and SCREENSHOT_RESOLUTION and compare with output dimensions. |
| Repeat capture is skipped | URL already exists and only-new behavior applies | Use archivebox add --no-only-new URL as documented; do not move or delete snapshot files manually. |
Or skip the browser setup
If you need a screenshot of a public page without maintaining a local browser capture setup, ScreenshotNeo is a website screenshot API and MCP server. Its one-call endpoint returns an image or PDF; see the API documentation for options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents use screenshot, page-information, and PDF-capture tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
What to include when asking for help
If the checks do not identify the cause, share enough detail for someone to distinguish browser setup from a target-specific problem:
- Output from
archivebox version, including the reported browser provider, version, and path. - Deployment type and operating system.
- The target URL, with sensitive query values redacted.
- Whether
screenshot.pngis absent, zero-byte, uniformly blank, or shows a page; include file size and dimensions. - The complete relevant extraction log, including timeout or browser launch messages.
- Whether PDF or DOM capture succeeds, and whether the issue repeats across other URLs.
The official ArchiveBox troubleshooting guide asks users with substantial problems to share affected URLs and errors. Without installation details and logs, there is no evidence-based way to name a single exact fix.
FAQ
Does a blank screenshot prove Chrome is not installed?
No. A valid image can show an empty, blocked, or unauthenticated page. Check the artifact and page state, then verify the browser runtime and logs.
Should I change the user-agent whenever a capture is blank?
No. ArchiveBox documents user-agent behavior as one possible site-specific issue. First confirm what the site returned and whether the browser can access the page normally.
Can a working PDF prove the screenshot extractor is configured correctly?
No. It shows that another capture path produced output, which helps narrow the investigation but does not establish why the screenshot differs.
Where should I start if every screenshot is missing?
Start with archivebox install chrome, archivebox version, dependencies, and the full extraction log. Those checks address the shared browser runtime before investigating individual sites.


