ScreenshotNeo

BlogHTML to image & PDF

How to Batch Convert Website URLs to PDF with CloudConvert

CloudConvert documents a capture task for one URL at a time. Learn how to automate that workflow for multiple pages, handle results, and understand its limits.

By the ScreenshotNeo team4 October 20269 min read

CloudConvert’s documented website-to-PDF workflow captures one URL per capture-website task, then uses an export/url task to retrieve the PDF. To process multiple URLs, you can submit a job containing a capture and export task for each URL. That multi-URL pattern follows CloudConvert’s named-task job model, but its reviewed documentation does not establish a dedicated bulk-URL feature, a maximum batch size, or concurrency guarantees. Treat it as an API implementation approach to validate against your account and current API limits.

CloudConvert describes the operation as: “Convert a website to PDF or capture a screenshot of a website (PNG, JPG).” Its HTML-to-PDF service is described as headless-Chrome based. Those are vendor-documented capabilities; rendering quality depends on the target site and has not been independently tested here. CloudConvert’s capture documentation is the source for the workflow below.

1. What the CloudConvert workflow does

A CloudConvert job is made of named tasks. For one page, the documented pattern is:

  1. Create a job with a capture-website task, setting the page URL and PDF output format.
  2. Add an export/url task that takes the capture task’s output as input.
  3. Wait for the job to finish, then download the PDF from the export task’s result.

For several pages, repeat those task pairs within a job: one uniquely named capture task and one export task per URL. This is an inference from the documented per-task URL field and named-task job structure, not a published CloudConvert bulk-conversion promise. The documentation reviewed here does not say how many tasks a job may contain or whether tasks run concurrently.

Need CloudConvert approach
One PDF per website URL One capture task with output_format: "pdf", then an export task
Several PDFs Repeat named capture/export task pairs in a job; confirm current task and account constraints
Completion notification Use a webhook, which CloudConvert recommends for asynchronous jobs
Long-term file storage Configure output storage such as S3 or Azure instead of relying on a temporary download URL

2. Create a job for one URL

Start with the official single-URL shape. Replace YOUR_API_KEY and the example URL with your values. Keep the API key on a trusted server; do not place it in browser-side JavaScript or a public repository.

curl --request POST \
  --url https://api.cloudconvert.com/v2/jobs \
  --header 'Authorization: Bearer YOUR_API_KEY' \
  --header 'Content-Type: application/json' \
  --data '{
    "tasks": {
      "capture_page": {
        "operation": "capture-website",
        "url": "https://example.com",
        "output_format": "pdf"
      },
      "export_pdf": {
        "operation": "export/url",
        "input": "capture_page"
      }
    }
  }'

The response includes the job and its task data. Once the job is finished, read the export task’s result to get the download URL. For production, handle job completion asynchronously rather than assuming the PDF is ready immediately.

3. Extend the task pattern to multiple URLs

For a multi-page batch, generate unique task names and connect each export task to its matching capture task. The following JSON shows the intended structure for two URLs. It is an implementation illustration based on the documented task model; the reviewed documentation does not confirm a dedicated batch endpoint, a permitted task count, or that this exact multi-capture payload is accepted by every account configuration. Check current CloudConvert API constraints before sending large jobs.

{
  "tasks": {
    "capture_001": {
      "operation": "capture-website",
      "url": "https://example.com/first",
      "output_format": "pdf"
    },
    "export_001": {
      "operation": "export/url",
      "input": "capture_001"
    },
    "capture_002": {
      "operation": "capture-website",
      "url": "https://example.com/second",
      "output_format": "pdf"
    },
    "export_002": {
      "operation": "export/url",
      "input": "capture_002"
    }
  }
}

A batch builder should validate each input URL, assign stable task names, and keep a map from input URL to export task. Avoid deriving task names directly from URLs: URLs can contain characters unsuitable for identifiers, be very long, or collide after normalization.

Batch-building checklist

  1. Normalize and validate the input list. Decide how to handle duplicate URLs and redirects.
  2. Assign each URL a deterministic, unique task name such as capture_001.
  3. Create a capture task with operation, url, and output_format.
  4. Create an export task whose input points to that capture task.
  5. Persist the job ID, input URL, and task names so a worker can resume result handling.
  6. Use webhooks for completion, or query the job endpoint and apply a bounded polling schedule.
  7. Download each exported file promptly or route outputs to configured storage.
  8. Record per-URL success or failure; a batch should not be treated as all-or-nothing unless your application explicitly enforces that policy.

4. PDF rendering options and page behavior

CloudConvert’s product documentation describes URL or HTML input, custom authorization headers for protected pages, waiting for a CSS selector, and layout controls including page size, margins, zoom, headers, and footers. These are CloudConvert’s stated capabilities, and the available sources do not independently verify rendering fidelity for a particular site.

Wait for content to appear

Some pages render content after JavaScript runs or after an API request completes. A selector wait can help when a known element indicates that the page is ready. Pick an element tied to the content you need in the PDF, not a generic page shell that appears before the content loads.

Protected pages

For a page that requires authorization, CloudConvert says custom headers can be supplied. Treat those credentials as secrets: use the minimum access needed, avoid logging header values, and do not use a capture service to access content you are not authorized to retrieve.

Page size, margins, zoom, and headers/footers

Set layout options to match the document’s purpose. Wider margins can prevent clipped content; zoom can trade readability for fewer pages. The capture documentation describes header/footer templates supplied through import tasks, with display_header_footer enabled and sufficient top or bottom margin space. Documented template classes include date, title, url, pageNumber, and totalPages. Consult the operation reference for the exact current parameter schema before adding these settings to a job.

5. Handle asynchronous jobs and downloaded files

CloudConvert recommends webhooks to notify your application when a job completes. Polling the job endpoint is another option, but use a delay that grows between requests and stop after a defined deadline. The Jobs API reference lists statuses including waiting, processing, finished, and error; your handler should account for all of them.

Two 24-hour lifecycle details matter and refer to different things:

  • CloudConvert’s quickstart says an export/url download link is valid for 24 hours.
  • The Jobs reference says jobs are automatically deleted 24 hours after they end.

Download or move output to your own storage within the link’s validity period. CloudConvert’s quickstart also describes routing output to storage such as Amazon S3 or Azure so your application does not have to rely on manually downloading temporary links.

Webhook handling

  • Verify incoming webhook requests according to CloudConvert’s current webhook guidance.
  • Make the handler idempotent: retries or duplicate notifications should not store the same PDF twice.
  • On completion, inspect each export task and record a result per URL.
  • On error, save the task error details for diagnosis and retry only when appropriate.
  • Return promptly from the webhook handler; perform file downloads in a worker if they may take time.

6. Troubleshooting common problems

Symptom Likely cause What to do
Job creation is rejected Malformed JSON, missing required capture fields, invalid task names or unsupported settings Validate the JSON, check the current capture operation schema, and ensure every export task references an existing capture task.
Job stays in waiting or processing The capture is still queued or the website takes time to render Use a webhook or bounded polling with backoff. Set an appropriate selector wait when the page has a reliable ready marker.
Job ends with error The site may be unreachable, access may be denied, or a task configuration may be invalid Inspect the individual task’s error details. Confirm the URL is reachable from the service, and check required headers or options.
PDF is blank or missing late content Client-rendered content had not appeared when capture began Wait for a content-specific CSS selector and confirm the selector exists on every URL in the batch.
Protected page shows an access error The page requires authentication or custom request headers Supply the required authorized headers through the documented mechanism; do not expose secrets in client code or logs.
Export link no longer works The temporary URL has expired CloudConvert documents a 24-hour validity window. Download promptly or configure storage for the output.
Some PDFs are absent from a batch A task failed, the result collector skipped it, or task names were mismatched Track every capture/export pair and inspect every task status independently instead of assuming the job produced all expected files.
Batch request exceeds an account or API constraint The reviewed sources do not specify a universal task-count limit Check current API and account limits. If needed, divide the input into smaller jobs and control how many jobs your worker submits at once.

7. Performance, reliability, and cost

Capture time depends on each target page, its assets, JavaScript behavior, and any wait condition. The reviewed documentation does not publish a throughput guarantee or a maximum batch size, so avoid promising a fixed completion time. For larger workloads, start with small jobs, measure your own pages, and tune job submission and retries around observed task outcomes.

For reliability, persist job and task identifiers before handing work to a background worker, make webhook processing idempotent, keep per-URL status, and retry transient failures selectively. Do not retry every failed capture indefinitely; an invalid URL or persistent access denial will not be fixed by repeated attempts.

CloudConvert’s HTML-to-PDF API page displays a starting price of $0.008 per file and directs readers to its full pricing information and calculator. Treat this only as a displayed starting figure, not a guaranteed quote for a particular plan, batch, or date. Confirm current pricing and any account-specific limits before estimating production costs. Review the CloudConvert API page for current details.

Or skip the browser setup

If you need clean website screenshots alongside PDF capture workflows, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. For a PDF, set the output format using the API’s documented parameters; the example below uses the product’s supplied one-call screenshot request.

See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted like a visitor would, and known consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers say the page verdict and whether it was billed.
  • An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
  • The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo to get 1,000 free screenshots a month with no card.

Frequently asked questions

Does CloudConvert have a documented bulk website-URL converter?

The reviewed documentation demonstrates one URL per capture task. It does not establish a separate bulk-URL interface or promise a particular batch size.

Does one job produce one combined PDF or one PDF per URL?

The documented capture operation produces an output for a capture task. The multi-URL pattern described here pairs each URL with its own export task, so plan to handle separate PDFs unless you add a separate PDF-merging step.

Can I capture a page that requires login?

CloudConvert’s product page describes custom authorization headers for protected URLs. Whether a site permits automated access depends on its authentication and access controls.

How long do I have to retrieve the result?

The quickstart says export URLs are valid for 24 hours. Jobs are also documented as being deleted 24 hours after they end. Download or route files to storage promptly.

Will every website render identically to a normal browser?

No universal fidelity guarantee is established by the reviewed sources. The service describes a headless-Chrome-based workflow, but site scripts, access controls, and dynamic content can affect the resulting PDF.