How to Capture Screenshots of Indian Government Websites with Apify
Use Apify’s Website Content Crawler to capture selected government pages, find their screenshots, and use them as evidence in a GIGW-informed review.
To capture screenshots of Indian government websites with Apify, run its Website Content Crawler with the crawler type set to playwright:firefox and saveScreenshots enabled. The crawler stores screenshots in the run’s default key-value store and adds a screenshotUrl field to dataset items for pages with screenshots. Screenshot capture can slow a run and increase storage costs, so start with a small, deliberate page scope.
What Apify captures, and where the screenshots go
The Website Content Crawler is the documented fit for a crawl that extracts page content and can save screenshots. Screenshot saving is off by default and is available only with the playwright:firefox crawler type. When enabled, screenshots are stored in the run’s default key-value store; dataset output references an image with screenshotUrl. Check the dataset after the run: a screenshot link is not guaranteed for every URL or every kind of page.
A screenshot is a record of one rendered page state at a particular time. Keep each screenshot’s page URL and capture date in your review notes. It can support visual inspection, but cannot establish backend behavior, ongoing compliance, or a site’s complete accessibility or security status.
Capture a focused set of pages in Apify Console
- Open the official Website Content Crawler Actor in Apify Console.
- Add the government website’s official URL as a Start URL. For a review of a particular service or section, seed that section or add a short list of specific page URLs.
- Set the crawler type to
playwright:firefox. Screenshot saving does not work with other crawler types. - Turn on
saveScreenshots. - Set scope controls before starting: use include and exclude URL patterns,
maxCrawlDepth, andmaxCrawlPages. Sitemap discovery can help find pages beyond the seed pages, but a sitemap crawl may take longer. - Start the Actor and wait for its run to finish. Open the run’s Output tab and inspect the dataset.
- Find each page’s
screenshotUrland open it to view the stored image. Keep the page URL beside the image when exporting results or preparing review notes.
For a one-off visual check, explicit Start URLs give tight control over scope. For a wider inventory, sitemap discovery can surface additional pages. In either case, use page and depth limits to keep the run bounded.
Automate an Apify run with cURL
Apify also documents starting Actors through its API. This example starts the Website Content Crawler with a single seed URL, Firefox/Playwright, screenshots enabled, and a page cap. Replace the token and target URL. The API response includes run information; use the returned run ID to retrieve the default dataset items after the run completes.
curl -X POST "https://api.apify.com/v2/acts/apify~website-content-crawler/runs?token=YOUR_APIFY_TOKEN" \\
-H "Content-Type: application/json" \\
-d '{
"startUrls": [{"url": "https://www.india.gov.in/"}],
"crawlerType": "playwright:firefox",
"saveScreenshots": true,
"maxCrawlPages": 10,
"maxCrawlDepth": 1
}'
After the run finishes, retrieve its default dataset items using the run ID from the response:
curl "https://api.apify.com/v2/actor-runs/RUN_ID/dataset/items?token=YOUR_APIFY_TOKEN&format=json"
Inspect the returned records for screenshotUrl. A run request and a dataset read are separate steps; the dataset may not contain final output until the Actor has completed. Apify’s default dataset API documentation describes retrieving the default dataset items for a run.
Using screenshots in a GIGW-informed review
Guidelines for Indian Government Websites and Apps (GIGW) 3.0 covers government websites and applications at central, state, district, and local levels. It addresses usability, user-centricity, universal accessibility, and security. Screenshots can help document selected visible conditions, such as whether ownership information and a visible update or review date appear on important pages, or whether a page’s visible content is laid out as expected.
Some checks require more than an image. Page titles and language attributes require inspecting page markup; accessible table markup requires accessibility or source inspection; print behavior requires testing the print experience. Other technical, manual, and security checks also need their own evaluation. GIGW supports quality assessment and STQC certification processes; an Apify screenshot run does not certify a website or prove complete GIGW conformity. Consult current official guidance and assess updated cybersecurity guidance where relevant.
Scope and responsible operation
Before crawling a specific government site, check its published access conditions and begin with a small run. The crawler offers a respectRobotsTxtFile option, but its documented behavior does not identify a specific user agent and does not support the crawl-delay directive. Do not treat that setting as complete robots-policy or rate-limit compliance.
- Use only the URLs needed for the review, and set page and depth caps.
- Use include and exclude patterns to constrain discovery to the intended section.
- Keep a record of capture date, page URL, and any relevant run settings.
- Review the output for missing screenshots and confirm that each image corresponds to the intended page.
Or skip the browser setup
If you need a screenshot without configuring a crawler, ScreenshotNeo takes a screenshot with one GET request. See the API documentation for its options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.india.gov.in/ -o shot.webp
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
No screenshotUrl in output |
saveScreenshots is off, or the crawler type is not playwright:firefox. A page may also fail to produce a screenshot record. |
Confirm both settings in the run input, then inspect the run log and the specific dataset item. |
| Run takes longer after enabling screenshots | Browser rendering and screenshot storage add work to the crawl. | Reduce Start URLs, page count, or depth; restrict URL patterns and avoid broad discovery for a small review. |
| Dataset has fewer pages than expected | The crawl is bounded by URL patterns and page/depth caps; discovery may not reach every page. | Review the input limits and patterns. Add deliberate Start URLs or use sitemap discovery if broader coverage is needed. |
| Screenshot link does not open | The URL may be missing, copied incorrectly, or the run output may not contain the expected record. | Copy the complete screenshotUrl from the dataset item and verify the run completed. |
| API request is rejected | The token may be absent or invalid, or the request body may not be valid JSON. | Use a valid Apify token, preserve the JSON content type, and check the API response for the specific error. |
| A page looks incomplete or different from a normal visit | The captured state may depend on page loading, interaction, access conditions, or the time of capture. | Open the target page independently, record the capture time, and treat the screenshot as evidence of that single observed state. |
Performance, reliability, and cost considerations
Apify’s Actor documentation warns that saving screenshots reduces performance and increases storage costs. Keep the crawl limited to the pages relevant to the review, especially when capturing a large domain. The research does not establish a fixed capture speed, storage price, or success rate; these depend on the run and target pages.
For repeatable reviews, invoke the Actor through the API or a client library and retrieve the run’s dataset programmatically. For manual reviews, Console provides a direct way to configure a run and inspect its output. In both cases, verify the resulting dataset and screenshot links rather than assuming every discovered page yielded an image.
FAQ
Does Apify save one screenshot for every URL?
The Actor documents screenshot links in dataset output, but a link should be confirmed per item. Check the output for screenshotUrl rather than assuming every requested URL produced an image.
Can a screenshot prove that a site complies with GIGW?
No. It can support selected visible checks. GIGW includes checks requiring markup inspection, interaction, print testing, manual evaluation, and other technical review.
Should I crawl the entire government domain?
Only if the review requires that scope. For a targeted inspection, explicit Start URLs and page/depth limits make the run easier to control and review.


