ScreenshotNeo

BlogHow-to

How to Batch Convert Indian Government Website Pages to PDF with CloudConvert

Create one CloudConvert website capture per page URL, export the PDFs, and check filenames, rendering, access, and current plan limits.

By the ScreenshotNeo team4 October 20267 min read

CloudConvert can capture a website as a PDF. To batch convert a list of Indian government page URLs, automate one capture-website task per URL, then export the resulting PDFs. The reviewed documentation describes a URL for each capture task; it does not establish a dedicated operation that accepts a list of website URLs in one capture task. Treat batching as a repeatable job-building pattern, not as a documented URL-list feature. CloudConvert’s capture-website documentation shows the capture and export task pattern.

1. Prepare and map the page URLs

  1. Collect the exact public page URLs in a CSV, spreadsheet, or text file. Use the individual pages you need, rather than assuming a homepage or search-results page will produce the same archive.
  2. For each URL, choose a stable output filename. An index plus a short page slug, such as 001-department-notice.pdf, helps avoid collisions and preserves ordering.
  3. Keep a source-to-output mapping. Record at least the source URL, intended filename, job or task identifier, and final download status. This makes failures and duplicate names easier to identify.
  4. Check that you are authorized to process the content. A page being publicly reachable does not establish that a particular third-party capture is permitted or that the page will be accessible from CloudConvert.

2. Create one website capture task per URL

In a CloudConvert job, create a capture-website task for each page and set the output format to PDF. Give every task a unique name. The following JSON illustrates how the documented per-page task shape can be composed for two URLs. It is an explanatory example, not a verified complete API request: confirm accepted fields, job size, and the exact request wrapper in CloudConvert’s live Job Builder or API documentation before submitting it.

{
  "tasks": {
    "capture-page-001": {
      "operation": "capture-website",
      "url": "https://example.gov.in/page-one",
      "output_format": "pdf",
      "filename": "001-page-one.pdf"
    },
    "capture-page-002": {
      "operation": "capture-website",
      "url": "https://example.gov.in/page-two",
      "output_format": "pdf",
      "filename": "002-page-two.pdf"
    },
    "export-pdfs": {
      "operation": "export/url",
      "input": ["capture-page-001", "capture-page-002"],
      "archive_multiple_files": true
    }
  }
}

The structure follows the documented capture-plus-export model: capture tasks produce PDFs, and an export/url task exports their results. The docs reviewed do not verify this exact combined two-page payload as a tested recipe or specify an unlimited number of captures per job. For a production batch, use the live API schema or Job Builder to validate the fields and generated job.

3. Automate a larger batch

Generate the capture task definitions from your URL list, assigning unique task names and filenames to every row. Submit jobs in manageable groups if the list is large, then track each URL through capture, export, download, and inspection. The official CloudConvert CLI documents batch processing for file conversions, including multiple inputs and globs; that example does not establish a batch-of-URLs feature for website capture. Its command documentation describes website capture as taking a URL. See the CloudConvert CLI documentation and use it only in ways its current command syntax supports.

There is not enough verified API request detail in the research for a complete, runnable CLI, Python, or Node.js submission script here. In particular, do not infer an endpoint, authentication format, or CLI flag from the task JSON above. Build and validate the exact request using CloudConvert’s current API documentation, then generate one capture task for each input URL. For small batches, repeated capture jobs are another practical option.

4. Handle dynamic pages and PDF layout

Some pages rely on client-side scripts or take time to render. CloudConvert’s HTML-to-PDF API documentation describes URL or HTML input, custom authorization headers, and waiting for a CSS selector. A selector wait can help when a reliable element appears only after the desired content has loaded. Choose a selector that marks the actual content, not a generic element that may appear before the page is ready. The capture operation also documents layout and header/footer settings. Review the HTML-to-PDF API and capture options for the fields currently available.

Preview representative output before processing a large set. A successful task does not guarantee that every interactive or script-rendered element appears as intended. Check page breaks, headers and footers, text clipping, and whether important content is missing.

5. Export and retrieve the PDFs

Use export/url to retrieve the capture results. For individual retrieval, handle the outputs as separate files. If you prefer a single download, the export documentation describes archiving multiple files into one ZIP. It also documents exporting to object storage.

CloudConvert documents export URLs as available for 24 hours. Download the files promptly, or use a supported object-storage destination when you need a retained archive. Keep the mapping between each source URL and its downloaded filename even when the files arrive in a ZIP. See CloudConvert’s export-files documentation.

6. Check the batch before relying on it

  • Confirm every input URL has a corresponding task result and output filename.
  • Open each PDF or apply a suitable automated document check to identify blank, incomplete, or error-page captures.
  • Check representative pages from different site sections, especially where layouts or scripts differ.
  • Look for clipped text, missing images or dynamic content, unexpected page breaks, duplicate names, and incorrect ordering.
  • Retry only failed or defective pages after diagnosing the cause; retain the original URL and task result in your tracking record.
  • Download temporary export links within their documented 24-hour availability period.

Limits, cost, privacy, and reliability

CloudConvert’s pricing page displayed a free tier with 10 credits per day, five concurrent tasks, a 1 GB maximum file size, and a five-minute maximum processing time when reviewed on October 3, 2026. These are current displayed plan limits, not permanent guarantees or a promise that a particular batch will fit. Check the live pricing page for current limits and costs before planning a larger run. The available research does not establish a maximum URL count per job or performance benchmarks for a given batch size.

For reliability, use manageable groups, unique names, a URL-to-file manifest, and a retry process for failed pages. Do not assume that public pages will all be reachable from a third-party service: redirects, rate limits, browser state, or script rendering may affect individual URLs. The compatibility and permitted capture status of any specific Indian government domain have not been verified here.

Review authorization and privacy before submitting credentials, personal records, or restricted content to a third-party processor. CloudConvert’s terms place responsibility for submitted data on the user. Its privacy policy identifies Lunaweb GmbH in Germany as data controller and is dated April 17, 2026. The terms also describe conditions for API use, including a requirement that API use be an integral component of software adding significant value and substantial functionality; review the current terms if you are building a public conversion product.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API can return a PDF as well as PNG, JPEG, or WebP. One GET request captures one URL; it is useful when you want a direct capture call rather than configuring a browser yourself. It does not replace the multi-page CloudConvert workflow above when your required output is a batch of PDFs.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

See the ScreenshotNeo API documentation for the current PDF parameter and options. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; and its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

Troubleshooting

Symptom Likely cause What to do
A page does not produce a PDF The URL redirects, is inaccessible to the capture service, or the task configuration is invalid. Check the exact URL and task result, then validate the operation fields against CloudConvert’s current documentation. Do not assume all government sites are reachable.
The PDF is blank or misses content The page may depend on scripts or may not have finished rendering when captured. Use an appropriate documented selector wait where available, and preview the result. Choose a selector that appears when the desired content is ready.
Text or page sections are clipped The page layout or print rendering may not fit the chosen PDF layout. Review the available layout and header/footer settings, then capture a sample again and inspect page breaks.
Some files are missing from the ZIP A capture task may have failed or was not included in the export task inputs. Compare export inputs with the capture task manifest and inspect each task result. Re-export after resolving failed captures.
An export link no longer works CloudConvert documents export URLs as temporary, available for 24 hours. Run or retrieve the export again if possible, and download promptly. For retained archives, consider a supported object-storage export.
The batch is slow or exceeds account limits Concurrency, processing-time, credit, or file-size limits may constrain the job. Check the live pricing limits and process manageable groups. The reviewed docs do not provide a guaranteed throughput figure or universal batch size.
The same output filename is overwritten or ambiguous Multiple tasks use a duplicate filename or the source-to-output mapping is missing. Use unique indexed or slug-based filenames and maintain a URL-to-filename manifest.

FAQ

Can CloudConvert batch convert a list of website URLs to PDF?

The documented pattern is one website capture task per URL, with results exported afterward. The reviewed documentation does not establish a single capture operation that accepts a URL list.

Can I combine the resulting PDFs into one download?

The export documentation describes archiving multiple output files into a ZIP. That gives you one archive download; it does not mean the PDFs have been merged into one PDF document.

Will every Indian government page work?

That is not established by the documentation reviewed. Check each specific page’s accessibility, rendering behavior, and applicable site requirements.

CloudConvert documents export/url links as available for 24 hours. Download them within that window or use a supported storage export for retention.

Does the CLI’s batch feature mean I can pass it a list of URLs?

No such URL-list behavior is established by the reviewed CLI documentation. Its batch examples concern file conversions; website capture is described as taking a URL.