How to Archive GST Portal Pages with ArchiveBox for Record Keeping
Save GST portal pages with ArchiveBox, preserve official downloads separately, and build a restorable recordkeeping workflow that accounts for login limits.
Use ArchiveBox to keep a dated snapshot of a GST portal page, but download the actual return data and official documents separately. A snapshot can preserve useful context about what a page displayed; it does not establish that the underlying return data was captured, prove authenticity, or by itself satisfy recordkeeping duties.
For login-protected pages, sign in through an authorized browser session and inspect the resulting snapshot. ArchiveBox documents browser-cookie and persona support for some authenticated sites, but the GST Portal uses CAPTCHA and OTP controls, so successful capture of every protected screen is not guaranteed. Do not try to bypass those controls.
What to preserve, and why
Keep two complementary kinds of material:
- Official records: downloaded returns, certificates, notices, statements, and other files relevant to the tax period. Keep these as the primary records.
- Page snapshots: ArchiveBox copies of relevant portal pages that can provide context, such as the page URL and what was visible when you captured it.
ArchiveBox is self-hosted software that saves URLs into a collection and can produce formats such as HTML, PDF, screenshots, and WARC. The collection consists of files that you must store, protect, back up, and be able to restore. See the ArchiveBox project and its usage documentation.
CBIC’s Accounts and Records Rules address electronic records, digital-signature authentication, electronic backup that can be restored within a reasonable period, production in readable form, and audit-trail information. The rules state that electronic records may be maintained when authenticated by digital signature. Whether a particular set of files and processes meets a taxpayer’s obligations depends on the circumstances; an ArchiveBox snapshot alone is not shown by these sources to satisfy them. Consult the current rules and a qualified adviser for questions about your obligations. CBIC GST Rules (PDF)
Retention context: portal availability and record retention differ
A GST Portal advisory dated September 24, 2024 stated that return data available for taxpayer viewing would be retained for seven years and advised taxpayers to download relevant data for future reference. The copy retrieved for this guide is hosted by a state GST department. Treat this as the statement in that advisory, not a guarantee of present portal behavior; check the current official advisory and portal guidance before relying on it. GSTN archival advisory PDF
A GST Council 2019 flyer summarizes the specified accounts and related documents as being preserved for 72 months (six years), calculated from the due date of the annual return for the relevant year. It also describes longer retention for certain proceedings or investigations. This general summary is not a personalized deadline; check the current law and your circumstances. GST Council publication
Prepare ArchiveBox
Install ArchiveBox using its current project instructions. The commands below assume the ArchiveBox CLI is installed and available as archivebox. Start with a dedicated, access-controlled data directory; do not put authenticated archives in a public web directory.
mkdir -p ~/gst-records/archivebox
cd ~/gst-records/archivebox
archivebox init
archivebox version
CLI options and installation details can vary by ArchiveBox version and installation method. Check archivebox help and the project documentation if a command differs in your environment.
Step-by-step GST archiving workflow
- Sign in normally. Use an account you are authorized to access. Complete CAPTCHA, OTP, or any other portal authentication as required. A previously trusted browser or device may have a different OTP flow, but do not assume that authentication will be skipped.
- Download the underlying records. Save applicable filed returns, acknowledgements, certificates, notices, and supporting data using the GST Portal’s available download functions. Give files clear period-based names, for example
FY2024-25_GSTR-3B_2025-04_filed-return.pdf. Use the actual file extension and contents rather than renaming a file to imply a format it does not have. - Record the relevant page URLs. Copy the page address after navigating to the relevant portal screen. Avoid including passwords, OTPs, or session tokens in URLs or notes.
- Add URLs to ArchiveBox. For public pages, add the URL directly. For an authenticated page, use the optional persona workflow below and expect that the portal may still require interactive authentication.
- Inspect every result. Open the archived output and check whether it contains meaningful page content for the relevant date and period. A URL entry or screenshot of a login challenge is not a captured return.
- Keep an index and backup. Record the capture date, original URL, tax period, record type, and location of the corresponding official download. Make a separate backup and periodically confirm it can be restored and read.
Archive a public URL
archivebox add 'https://example.gov.in/public-information'
Replace the example with the actual page URL. To add several URLs from a text file, put one URL per line and pipe the file into ArchiveBox:
cat > gst-urls.txt <<'EOF'
https://example.gov.in/public-information
https://example.gov.in/another-public-page
EOF
archivebox add < gst-urls.txt
The example domains are placeholders. Use only URLs you are authorized to access and retain.
Optional: use an authorized browser persona
ArchiveBox documents importing cookies and browser state into a persona for some logged-in-site workflows. Its documentation describes Chrome, Chromium, Brave, and Edge profile import. This is a way to provide an existing authorized session to compatible extractors; it does not solve CAPTCHA, OTP, expired-session, or portal-specific capture limitations. Use a dedicated browser profile for archiving where practical, and protect the resulting files as sensitive data.
# Run from the ArchiveBox data directory; import a browser profile you are authorized to use.
archivebox persona create --import=chrome gst-archive
# Add a URL using that persona.
archivebox add --persona=gst-archive 'https://example.gov.in/authorized-page'
See ArchiveBox’s persona usage notes and its security overview. Imported cookies and browser profiles can contain session tokens and personal information. Snapshots may expose secrets reflected in page content or response metadata. Keep the archive private, restrict filesystem access, and do not share it or publish it without reviewing and sanitizing sensitive content.
Check the saved output
ArchiveBox can produce several kinds of output, depending on available extractors and configuration. A screenshot or PDF is a visual rendering; HTML or a single-file capture may preserve page structure; WARC is a web-archive format. These formats are not interchangeable and none automatically proves that a page is an authenticated official record.
- Open the snapshot from the ArchiveBox interface or the collection’s documented layout.
- Confirm that the page shows the expected tax period and content rather than a sign-in form, CAPTCHA, OTP prompt, error, or empty shell.
- Compare any displayed values to the separately downloaded official file.
- Keep the original portal URL and capture date in your index. Preserve the downloaded source file unchanged.
Organize, secure, and back up the archive
A practical directory structure separates official downloads from page snapshots and keeps a human-readable index:
gst-records/
official-downloads/
FY2024-25/
returns/
notices/
certificates/
archivebox/
data/ # ArchiveBox collection
index.csv # period, type, original URL, capture date, file path
backup-log.txt
This is an organizational example, not an ArchiveBox-mandated layout. Choose storage appropriate to your volume and access needs. Keep another copy on separate storage, such as an external drive, and maintain a routine that tests restoration. CBIC’s rules call for proper electronic backup that can be restored within a reasonable period. A backup that has never been checked may not be usable when needed.
Restrict access to the archive and backups. Authenticated snapshots may contain personal or business information, cookies, or session-related material. Encrypt storage where appropriate, keep the backup separately, and review retention and deletion practices with the people responsible for your records.
ArchiveBox snapshots and official downloads compared
| Question | Official portal download | ArchiveBox snapshot |
|---|---|---|
| What does it preserve? | The downloaded return or document itself, subject to the portal’s download options. | A capture of a URL using available extractors; output can include rendered or source-oriented formats. |
| Does it require a login? | Use the portal’s normal authorized login and download flow. | Public URLs may not. Authenticated pages may require a compatible browser persona and a valid session; capture is not guaranteed. |
| What is it useful for? | Preserving the underlying file for retrieval and review. | Providing context about what a web page displayed at capture time. |
| What must you verify? | That the correct period and document were downloaded and remain readable. | That meaningful page content was captured, the URL and date are recorded, and the archive remains private and restorable. |
Use both where useful, but keep official downloads and other required records central. A snapshot does not turn a login page or a partial rendering into the underlying return data.
Or skip the browser setup
If you need a rendered screenshot of a public page for context, ScreenshotNeo can return an image with one GET request. This does not replace GST downloads or ArchiveBox’s broader archiving workflow. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
- Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing status in headers.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents. - The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. All features are on every plan.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| The snapshot contains a login page or CAPTCHA. | The page required interactive authentication or the session was not available to the extractor. | Sign in through the authorized portal flow, then retry with a compatible persona if appropriate. Do not bypass CAPTCHA or OTP. If the protected content still is not captured, retain the official download and document the limitation. |
| The capture is blank or missing data. | The page may rely on client-side rendering, an expired session, delayed content, or a failed load. | Check the page in the authorized browser, verify the downloaded source record, and inspect ArchiveBox’s capture output and logs. Do not treat an empty snapshot as evidence of page contents. |
archivebox command is not found. |
The CLI is not installed or is not on the current shell’s PATH; a Docker-based installation may require running commands through its documented container setup. | Follow the current installation instructions and confirm with archivebox version or the relevant container command. |
| Persona import fails or browser state is absent. | The browser is unsupported, the wrong profile was selected, or ArchiveBox cannot access the browser’s profile data. | Check the documented supported browsers and profile selection options. Use an authorized profile and review ArchiveBox’s current persona instructions. |
| Files expose private information. | Authenticated page content, cookies, or session data may be preserved in snapshot files or metadata. | Restrict access, do not publish the archive, rotate compromised credentials if needed, and review the project’s security guidance before sharing any material. |
| A saved URL appears to be missing or duplicated. | ArchiveBox settings may skip existing URLs or an import may not have supplied the intended URL. | Check the collection and current CLI help. ArchiveBox documents an explicit re-snapshot option, --no-only-new, for intentionally capturing a URL again. |
| A backup exists but cannot be restored. | The backup was not verified, is incomplete, or lacks required files. | Run periodic restore checks into a separate location and confirm both the ArchiveBox collection and official downloads are readable. |
Performance, reliability, and cost considerations
- Capture time: Browser-rendered pages and pages that load content dynamically can take longer than simple static pages. Authentication prompts and portal delays can also prevent a useful capture. Avoid promising a fixed capture time.
- Reliability: A successful command means a URL was processed, not necessarily that the intended authenticated content was preserved. Inspect each important result and keep source downloads.
- Storage: HTML, screenshots, PDFs, media, and WARC data have different storage footprints. Keep only the outputs that fit your retention and retrieval workflow, while ensuring that required material remains available.
- Maintenance: Self-hosting means you manage software updates, access controls, storage, backups, and recovery. Keep the archive independent from the machine whose failure it is meant to survive.
- Cost: ArchiveBox is self-hosted; practical costs depend on your compute, storage, backup, and maintenance choices. The dossier does not establish a universal cost or benchmark.
FAQ
Can ArchiveBox save GST returns before the portal removes them?
It may capture a page if the relevant content is accessible to its extractors, but download the official return data directly from the portal while available. A page snapshot is not a substitute for the downloaded record.
Can ArchiveBox archive a GST page that requires login?
ArchiveBox documents personas and browser-cookie import for some authenticated sites. The GST Portal’s CAPTCHA and OTP controls mean compatibility and successful capture cannot be promised. Use an authorized session and inspect the output.
Does an ArchiveBox snapshot meet GST recordkeeping requirements?
The cited sources do not establish that a snapshot alone meets those requirements. Preserve the actual records, maintain restorable backups and relevant audit-trail information, and confirm obligations against current rules and your circumstances.
Should I keep both a screenshot and the downloaded return?
When useful, yes: the download preserves the document, while a snapshot can add context about a page. Keep the source file and its period labels clear so the two are not confused.


