How to Archive Pages Behind a Login with ArchiveBox
Use an imported browser profile and ArchiveBox persona to capture authenticated pages, then protect the cookies and private data stored in your archive.
To archive a page that requires a login with ArchiveBox, import a dedicated logged-in Chromium browser profile as a persona, then add the page URL with that persona selected. This supplies authentication to Chromium-based extractors and, when cookie export succeeds, to extractors that read a Netscape-format cookies.txt. Treat the profile and resulting archive as sensitive: captures can retain personal data and reusable session tokens.
ArchiveBox documents authenticated archiving as an advanced use case. The commands below follow its documented workflow; use a separate browser profile and, where possible, a burner account rather than your everyday credentials. See the ArchiveBox usage guide, Chromium setup documentation, and security overview.
Recommended workflow: import a browser profile as a persona
- Create a dedicated browser profile. In Chrome, Chromium, Brave, or Edge, make a profile used only for archiving. Sign into the target site in that profile. A dedicated or burner account limits the impact if archived session data is exposed.
- Import the profile into an ArchiveBox persona. Close the browser first so its profile data is no longer being actively changed. Run the command from your ArchiveBox installation/context:
archivebox persona create --import=chrome personal
The documentation describes importing profiles from Chrome, Chromium, Brave, and Edge. If needed, specify a browser profile such as Default or Profile 1 according to the persona command’s options in your installed ArchiveBox version.
- Add the authenticated URL using that persona:
archivebox add --persona=personal 'https://members.example.com/'
Replace the URL with the page you can access in the dedicated browser profile. Keep the URL quoted, especially if it contains shell characters such as &. The persona name must match the one created above.
- Review the resulting snapshot. Verify the page content and each output type you care about. A successful browser screenshot does not prove that a cookie-based extractor also authenticated successfully.
What the persona supplies to extractors
An imported persona keeps a Chromium profile and an exported cookie file together. Chromium-based extractors can reuse the profile; tools such as wget, curl, and yt-dlp can use the exported cookies.txt. ArchiveBox names Screenshot, PDF, DOM, and SingleFile among Chrome-based extractors, and wget, Mercury, and media among examples that do not use Chrome.
| Capture method family | Authentication source | What to verify |
|---|---|---|
| Chromium-based (for example, Screenshot, PDF, DOM, SingleFile) | Imported browser profile | Check the rendered page or saved file for the expected signed-in content. |
| Cookie-file consumers (for example, wget and media extractors) | Exported cookies.txt |
Check that cookie export exists and that the output contains authenticated content. |
This split means authentication is not automatically shared with every extractor just because one output worked. If automatic cookie extraction fails, the usage documentation describes placing a Netscape-format cookies.txt in the persona directory. Protect that file like a password: it may contain live session credentials.
Alternative: sign in to an ArchiveBox-managed Chrome profile
You can sign into a Chrome profile managed by ArchiveBox and use that profile for Chrome-based captures. This route does not itself create a cookies.txt for extractors that do not use Chrome. Choose it only when the Chrome-based outputs are sufficient, and inspect each desired output after capture.
| Approach | Coverage | Consideration |
|---|---|---|
| Import host profile into persona | Chromium profile plus exported cookies, if extraction succeeds | Recommended when both browser and cookie-based extractor families need login state. |
| Log into ArchiveBox-managed Chrome | Chrome-based extractors | Does not create the cookie file required by non-Chromium extractors. |
Protect private captures and session data
Authenticated snapshots may contain personal information, private tokens, or session cookies. ArchiveBox warns: “Future viewers of your archive may be able to use any reflected archived session tokens to log in as you.” Use a separate account, keep the archive directory private, and avoid publicly serving authenticated captures unless access is restricted and you understand the exposure.
For private captures, ArchiveBox’s security overview demonstrates disabling the Archive.org integration before creating the persona and adding the private URL. Its documented visibility controls include PUBLIC_INDEX=False, PERMISSIONS=private, and PUBLIC_ADD_VIEW=False. The authentication documentation describes PERMISSIONS as per-snapshot and says custom permissions for non-admin users and groups are not currently supported. Check the documentation for your version before changing configuration.
Visibility settings control access through the application; they do not remove sensitive data from the files stored on disk. Restrict access to the archive directory and backups too. Consider whether the site owner’s terms and your organization’s policies permit archiving the content.
Or skip the browser setup
If you only need a clean screenshot of a page you can access, ScreenshotNeo is a website screenshot API and MCP server. It does not import your ArchiveBox browser session or archive authenticated content that requires those cookies. For a page the service can reach, make one GET request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers say the page verdict and whether the request was billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Capture shows a login page | The browser profile was not signed in, the wrong profile was imported, or the login expired. | Sign in again in the dedicated profile, confirm the correct profile name, re-import if needed, and add with the matching persona. |
| Chrome screenshot works, but wget or media output is logged out | Those extractors use cookies rather than the Chromium profile, or cookie export failed. | Check for the persona’s Netscape-format cookies.txt; use the documented manual placement workflow if automatic extraction fails. |
| No authenticated output appears | The URL may redirect, require an interactive step, or the session may have expired. | Open the exact URL in the dedicated browser profile, complete any required access step, then retry and inspect the captured output. |
| Persona command cannot find the browser profile | The browser/profile selection does not match the installed profile location or name. | Confirm the supported browser and profile name (for example, Default or Profile 1) and consult the usage guide for the installed version. |
| A private capture is visible through the web interface | Public visibility or add-view settings may still allow access. | Review PUBLIC_INDEX, PERMISSIONS, and PUBLIC_ADD_VIEW settings, then verify access behavior. Restrict filesystem and backup access as well. |
| Archived session remains usable longer than expected | The snapshot preserved a session token. | Revoke the account session at the source site, remove exposed copies where possible, and treat all archive replicas as sensitive. |
Reliability, performance, and storage considerations
- Session freshness: Sessions expire and sites can require reauthentication. Confirm the page is signed in immediately before importing and capturing.
- Extractor-specific results: Browser and cookie consumers authenticate differently. Check each output needed for the archive rather than relying on a single success signal.
- Profile consistency: Import a dedicated profile that is not concurrently being changed by another browser process. If the import is stale, refresh the login and import again.
- Private-data footprint: The archive can retain the captured page and credentials or tokens reflected in it. Account for archive copies, backups, and any access-controlled hosting when deciding what to capture.
- Cost: The supplied ArchiveBox documentation does not establish a per-capture price or storage estimate. Plan based on your own installation and storage requirements; do not assume visibility settings encrypt or erase files.
FAQ
Can I use my everyday browser profile?
The workflow can import a browser profile, but a dedicated profile and burner credentials reduce exposure of unrelated browsing data and account access.
Does importing a persona make every extractor authenticated?
No. Chromium-based extractors use the profile, while some other extractors need the exported cookie file. Verify the outputs you need.
Does setting a private permission make the archive files harmless?
No. Application access controls do not remove personal data or session tokens from stored snapshots and backups.
Can ScreenshotNeo capture the same private page using my ArchiveBox login?
No. ScreenshotNeo’s request example is for a URL the service can access; it does not use the imported ArchiveBox persona described here.


