ScreenshotNeo

BlogHow-to

How to Archive Indian Government Websites with ArchiveBox

Preserve Indian government pages with ArchiveBox: define scope, capture useful formats, record provenance, and keep a separate backup.

By the ScreenshotNeo team4 October 20268 min read

ArchiveBox is a self-hosted tool for saving supplied URLs in formats such as WARC, HTML, PDF, and screenshots. To make a useful reference collection of Indian government pages, first define what you need to preserve, check the particular site’s rules, capture a small set of URLs, record where and when each capture was made, and keep a separate copy of the collection.

A personal ArchiveBox collection is a research or reference copy. It does not become an official government archive, establish that every page or embedded resource was captured, or by itself demonstrate compliance with public-records obligations.

1. Define the collection’s purpose and scope

Before importing URLs, decide what question the collection is meant to answer and write down the boundaries. A focused list is easier to review and preserve than an open-ended crawl.

  • Identify the source: department or agency, official domain, and relevant subdomains.
  • List the material: specific pages, tenders, recruitment notices, announcements, press releases, reports, or downloadable documents.
  • Set the time range: note the period covered and capture date. For time-sensitive material, record the notice’s publication, validity, and expiry dates when available.
  • Include language versions: list English and other relevant versions separately; do not assume one page represents every version.
  • Choose a capture boundary: decide whether you are saving only listed pages or following links, and document that decision.

Prefer official domains and retain each exact URL. Government website guidance treats a site’s URL as an authenticity indicator, but a URL alone does not prove that a local copy is complete or authoritative.

2. Check the specific website’s rules

Review the site’s copyright, terms of use, privacy notice, content archival policy, and access directives such as robots.txt before automating requests. Rules and technical restrictions can vary by site. General government website guidance is not blanket permission to crawl every site or reuse its material.

Keep capture proportionate to the purpose. Avoid bypassing login controls, access restrictions, or anti-abuse measures. If the intended use has legal, institutional, or publication consequences, consult the responsible records or legal team rather than treating a personal capture as authorization.

3. Install ArchiveBox and make a small pilot

ArchiveBox is self-hosted and supports URL imports. Follow the current installation instructions and operating guidance in the ArchiveBox documentation, since installation requirements can depend on the environment and change over time. This article does not assume a particular operating system or installation method.

  1. Set up a local ArchiveBox instance using the current official instructions.
  2. Create a short input list containing a few representative URLs, including a normal page and, if relevant, a notice or document.
  3. Import that list through the method documented for your installation.
  4. Review the resulting captures and note which formats were produced and which resources did not load.
  5. Only after reviewing the pilot, expand the URL list or establish recurring imports.

ArchiveBox offers command-line, web-interface, REST API, Python API, and filesystem ways to work with collections. Choose one workflow that is maintainable for your project. For recurring work, its documentation describes importing from sources such as bookmarks or feeds; schedule and review those imports according to your needs.

4. Choose formats for preservation and access

ArchiveBox can produce original HTML, CSS, and JavaScript, single-file HTML, PNG screenshots, PDF, WARC, page titles, article text, favicons, and headers. Outputs depend on the page and capture process, so check what exists for each item rather than assuming every format is present.

Representation Useful for Keep in mind
WARC A portable web-archive representation for preserving captured web responses. It is not the same as a rendered, easy-to-read page; retain a convenient copy too when useful.
Original HTML/CSS/JavaScript Inspecting captured source files and resources. Dynamic behavior or externally hosted dependencies may not be preserved or replay correctly.
Single-file HTML Opening a consolidated page for convenient reference. It may not preserve all behavior or every resource from the live site.
PDF or PNG screenshot Quick visual reference or sharing a fixed rendering. These are rendered views, not substitutes for the underlying source or a WARC.
Title, article text, favicon, headers Search, cataloging, and contextual clues about a capture. Use them as supporting metadata, not proof that the full page was preserved.

For notices or pages whose later interpretation matters, retain the original capture formats available and add a convenient rendered format where appropriate. Record the source URL and capture date in a catalog or collection notes. Also record the scope, import method, and any known capture gaps.

5. Preserve dates, context, and a second copy

For each important item, keep a simple catalog entry with the exact URL, page title, capture date and time, department or site, document or notice date if shown, language, and a brief note about why it was included. This provenance is practical guidance for making a collection understandable later; it does not certify authenticity.

Keep a duplicate of the ArchiveBox collection somewhere separate from the machine running it. An offline external hard drive is one possible option. Periodically check that the backup is readable and that representative archived files can still be opened. A second copy helps with device failure, but it does not prove the capture is complete or satisfy an institution’s records policy.

6. Understand the institutional boundary

The Government of India’s GIGW policy template says departments must set out how long older material stays online, when it moves to offline archives, and whether or when it can be deleted. GIGW guidance also addresses archival handling for expired announcements, tenders, recruitment notices, news, and press releases. These are institutional website-management responsibilities, not a grant of permission to crawl any site.

The National Archives of India describes its role as custodian of Government of India records of enduring value, with responsibilities under the Public Records Act, 1993 and Public Records Rules, 1997 for records in central-government ministries, departments, and public-sector undertakings. That context does not mean NAI accepts arbitrary website crawls or that a personal ArchiveBox collection meets statutory requirements.

Describe your collection precisely: it preserves material retrieved by a particular process at a particular time. Without additional evidence, do not claim it is complete, an official archive, legally admissible, or proof of the authenticity of every embedded item.

7. Troubleshooting ArchiveBox captures

Symptom Possible cause What to do
A URL produces no useful capture The page may be unavailable, restrict automated access, require authentication, or depend on a redirect or script. Open the URL normally, check the site’s access rules and the capture logs, and record the failure. Do not bypass access controls.
The screenshot or PDF is blank or incomplete Content may load dynamically, require interaction, or depend on resources the capture did not retrieve. Compare the available output formats, inspect logs and captured files, and document the limitation. Do not treat one rendered output as a complete record.
Images or linked resources are missing Resources may be hosted on another domain, blocked, or loaded only after interaction. Review the site’s rules and the capture’s recorded resources. Add relevant allowed URLs explicitly when appropriate, then recapture and note what changed.
A page changes between captures The live site was updated, or the capture happened at a different time or state. Keep both captures with their dates and URLs. For notices, preserve the notice’s own dates and context in your catalog.
Recurring imports miss newly relevant pages The source feed or bookmark set may not include them, or the import scope may be too narrow. Review the source and scope, add known URLs, and periodically audit the collection against the pages you intend to preserve.
A local copy will not replay like the live site Pages may rely on scripts, authentication, or external services that were not captured or are no longer available. Use static outputs such as PDF or screenshot for visual reference alongside source files and WARC where available. Record replay limitations.

8. Performance, reliability, and storage

There is no universal capture-success rate or storage requirement established here. Collection size and completion time depend on the number of URLs, page complexity, resources, and the formats produced. Start with a small pilot and inspect both its outputs and storage use before increasing the scope.

  • Keep imports bounded: use an explicit URL list or well-defined recurring source, and review changes to that source.
  • Expect variable results: dynamic pages, third-party dependencies, authentication, and access restrictions can affect what is captured.
  • Plan for recovery: keep a second copy and periodically confirm it can be read.
  • Track gaps: log failed or partial captures so later readers do not mistake missing material for content that never existed.

Or skip the browser setup

If you only need a clean screenshot of a page rather than a self-hosted preservation collection, ScreenshotNeo provides a website screenshot API. It is not an ArchiveBox replacement or an official records archive. One GET request can return a PNG, JPEG, WebP, or PDF; see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.india.gov.in -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://www.india.gov.in"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://www.india.gov.in',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie and consent banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include X-Page-Verdict and X-Billed headers.
  • An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
  • The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.

FAQ

Can ArchiveBox save a whole government website automatically?

ArchiveBox accepts URLs and supports recurring imports, but this does not guarantee a complete site crawl. Define a boundary, check site rules, and review what was actually captured.

Should I keep WARC if I already have a PDF?

They serve different purposes: a PDF is convenient to read, while WARC is a web-archive format. Keep both when each is useful to your preservation goal.

Does a personal capture satisfy a government records requirement?

No such conclusion follows from using ArchiveBox. Institutional retention and official records responsibilities belong to the relevant department and applicable policies.

Can ScreenshotNeo replace ArchiveBox?

No. ScreenshotNeo returns screenshot or PDF captures through an API, while ArchiveBox is a self-hosted collection tool with multiple preservation-oriented outputs. Choose according to whether you need a quick rendered capture or a locally managed collection.