ArchiveBox Cannot Download YouTube Pages: How to Troubleshoot Extractor Errors
Diagnose ArchiveBox YouTube extractor errors by checking the selected tool, dependencies, URL access, and logs before retrying or preserving page context another way.
When ArchiveBox cannot download a YouTube page, first determine whether adding the URL failed or whether the media extractor failed during capture. Inspect the ArchiveBox output for the exact URL and failed step, check which extractor and version ArchiveBox resolved, and test whether the URL opens in a normal browser. The error text matters: there is no single fix for every YouTube extraction failure.
A failed media download does not necessarily mean ArchiveBox could not preserve any page context. A screenshot or HTML snapshot may still be possible, but neither should be described as a downloaded, playable video.
1. Identify exactly what failed
Start with the ArchiveBox stdout or Web UI output. Locate the URL and the extractor step that produced the error. Distinguish among:
- URL ingestion or parsing failure: ArchiveBox did not successfully add or interpret the submitted URL.
- Media extraction failure: the capture reached the media extractor, such as
yt-dlporyoutube-dl, but that step failed. - Other capture-method failure: a screenshot, browser, or other archive method failed independently of media extraction.
Do not infer the cause from the phrase “cannot download” alone. Preserve the complete error, including any lines immediately before and after it.
2. Check ArchiveBox and its resolved dependencies
ArchiveBox documents yt-dlp or youtube-dl among its media extraction dependencies. Its current documentation describes dependency resolution through abxpkg; the install and version commands below help establish what ArchiveBox can use in the current deployment. See the ArchiveBox troubleshooting guide and ArchiveBox documentation.
archivebox version
archivebox install
Run these in the same environment where ArchiveBox runs. A host-installed executable may differ from the dependency available inside a container. For the supported bare-metal setup, the troubleshooting guide also recommends checking the environment manager and installed tool:
uv --version
uv tool list
archivebox version
Use the installation path appropriate to your actual deployment. After installation or dependency resolution, inspect archivebox version again and confirm which media extractor is present and selected. Do not assume a particular extractor is active just because it is installed on the host.
3. Verify the URL and ordinary access
Open the exact target URL in a normal browser. Confirm that it is the intended video or page and that it is currently reachable from the machine or network running ArchiveBox. This is a useful first check, but browser access alone does not prove that an extractor can retrieve the media.
The yt-dlp FAQ explains that support cannot always be determined from a site’s appearance on a support list: URL schemes and site behavior can change. The practical diagnostic is to run the relevant extractor and inspect its output. Do not treat a URL as supported or unsupported without that evidence.
4. Capture diagnostic details before changing things
Record enough information to make the failure reproducible. yt-dlp recommends including verbose version output when reporting issues; its documented form is yt-dlp -Uv followed by the rest of the command. For example, once you know the extractor command ArchiveBox is using, run the equivalent command with verbose output in the same environment and preserve the full result.
yt-dlp -Uv "https://www.youtube.com/watch?v=VIDEO_ID"
Use the actual URL that failed. If ArchiveBox invokes youtube-dl rather than yt-dlp, record that fact and the version shown by the environment; do not silently substitute one tool and then report its output as the original failure.
Collect these details:
- ArchiveBox version and deployment type (for example, container or bare metal).
- The exact URL and whether it opens in a browser from the ArchiveBox host.
- The extractor ArchiveBox actually selected and its version.
- The complete ArchiveBox error output and, where applicable, verbose extractor output.
- What changed immediately before the error began, such as an upgrade or dependency reinstall.
5. Match the error to a supported next step
Use the observed error to decide what to investigate. Avoid applying a generic cookie, user-agent, IP, or version workaround without evidence that it addresses this failure.
| What you observe | What to check next | What not to assume |
|---|---|---|
| ArchiveBox reports a missing dependency or cannot find an extractor | Run archivebox version and archivebox install in the ArchiveBox runtime; verify the container or environment has the resolved dependency. |
That installing a tool on the host also installs it inside a container. |
| The URL does not open in a normal browser from the relevant network | Confirm the URL is correct and investigate ordinary access or network reachability first. | That the media extractor is the cause when the page itself is inaccessible. |
| The extractor returns an error for a URL that opens in a browser | Retain verbose output, confirm the selected extractor and version, and consult current yt-dlp guidance for that exact error. | That browser access guarantees media extraction will work. |
| A 403 or access-related error appears | Read the current yt-dlp FAQ guidance for the exact error and environment. Its discussion includes site-specific cookie and IP considerations. | That exporting cookies is a universal YouTube fix. |
| An error resembles “Unable to extract uploader id” | Check which extractor is being invoked and its version, then compare the complete output with current tool guidance. | That this historical error identifies the cause of every current YouTube failure. |
A historical ArchiveBox issue from 2023 recorded “Unable to extract uploader id” in one Docker Compose setup using youtube-dl. It is a useful example of why the selected extractor and its version matter, but it is not evidence that all current YouTube errors share that cause. See the historical issue.
6. Retry an already indexed URL intentionally
ArchiveBox normally avoids downloading a URL that is already indexed. After correcting a suspected dependency or configuration issue, use its documented retry form to request another attempt:
archivebox add --no-only-new "https://www.youtube.com/watch?v=VIDEO_ID"
Replace the example with the exact URL. This asks ArchiveBox to attempt the capture again; it does not guarantee that YouTube media extraction will succeed. Avoid moving or deleting parts of the archive tree as a retry workaround: ArchiveBox warns that separating database state from snapshot files can cause problems.
7. Preserve page context if media extraction still fails
ArchiveBox supports multiple archive methods, and its troubleshooting guidance recommends combining methods for sites that are not captured effectively by every method. A screenshot or HTML snapshot can preserve page context when available, while media extraction remains a separate outcome. These captures are not a substitute for a downloaded playable video.
If what you need is a clean visual record of the page, a screenshot service can capture that page without requiring you to maintain browser automation. ScreenshotNeo is a website screenshot API and MCP server; its capture flow accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Each cleanup step can be turned off. That preserves a visual page capture, not the YouTube video file.
Or skip the browser setup
For a screenshot of the page rather than a media download, ScreenshotNeo accepts one GET request. The example targets the YouTube video page; replace it with the page URL you want to capture. Get an API key and see the ScreenshotNeo API documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.youtube.com/watch?v=VIDEO_ID -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://www.youtube.com/watch?v=VIDEO_ID",
},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://www.youtube.com/watch?v=VIDEO_ID',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
ScreenshotNeo removes cookie banners, popups, and chat widgets before the screenshot. Bot checks, blank pages, and failed loads are never billed, and the response identifies the page verdict and billing status in headers. Its MCP server gives AI agents screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card.
Performance, reliability, and cost considerations
- Keep the failure local to the failing step. A media extractor error does not establish that all ArchiveBox capture methods failed. Inspect each method’s result before deciding what was preserved.
- Retry after a meaningful change. Repeatedly adding an already indexed URL without the documented retry option may not run a fresh capture. Retry after checking or changing the suspected dependency or access condition.
- Preserve logs. Full output and tool versions make it possible to distinguish an outdated or missing dependency from URL access or extractor behavior.
- Budget for the required outcome. If you need playable media, a page screenshot is not equivalent. If a visual record is sufficient, ScreenshotNeo bills only clean shots; bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers say which outcome occurred. Its published monthly plans are Free: 1,000 shots; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.
Troubleshooting checklist
- Find the exact ArchiveBox output line and identify whether ingestion, media extraction, or another capture method failed.
- Run
archivebox version; check the resolved extractor and its version in the same runtime. - Run
archivebox installif dependency resolution or installation is indicated. - Open the exact URL in a browser from the machine or network running ArchiveBox.
- Capture complete logs and use verbose extractor output when appropriate.
- Consult current yt-dlp guidance for the observed error; do not apply cookies or other workarounds generically.
- If the URL is already indexed and you have addressed a suspected cause, retry with
archivebox add --no-only-new URL. - Check whether another archive method preserved useful page context, and label it accurately as a page capture rather than a video download.
FAQ
Does a YouTube extractor error mean the whole ArchiveBox capture failed?
No. Check the output of each archive method. Media extraction may fail while another method preserves page context.
Can I tell whether yt-dlp supports a URL from a support list?
Not reliably in every case. URL formats and site behavior can change; run the extractor and inspect its output for the exact URL.
Will retrying guarantee that the video downloads?
No. --no-only-new requests a new attempt for an already indexed URL. Success still depends on the URL, access, and extractor behavior.
Can ScreenshotNeo download the video for ArchiveBox?
No. ScreenshotNeo returns a screenshot or PDF of a page. It is useful when a visual record is the desired outcome, not when you need the playable media file.


