ScreenshotNeo

BlogHow-to

CloudConvert API Tutorial: Convert Web Pages to PDF in Node.js

Convert a public web page to PDF with CloudConvert API v2 in Node.js, from job creation through export, with production notes and a simpler alternative.

By the ScreenshotNeo team4 October 20269 min read

Use CloudConvert API v2’s capture-website task to render a public page URL as a PDF, then connect an export/url task to retrieve the result. In Node.js, create a job with those two tasks and wait for completion in a simple script. For production, process completion asynchronously with a webhook and store or deliver the resulting file through your application.

1. Set up a CloudConvert API client

Create a CloudConvert API key and keep it on the server. Job creation requires the task.write scope. The official API v2 base URL is https://api.cloudconvert.com/v2, and CloudConvert provides a Node.js SDK. See the API documentation and quickstart.

Install the SDK in your Node.js project:

npm install cloudconvert

Set the API key in the server environment, for example:

export CLOUDCONVERT_API_KEY="your-api-key"

Do not put this key in browser JavaScript or expose it in a public repository. Use a server-side environment variable or a secrets manager.

2. Create a website-to-PDF job

This complete example creates a capture task for a public URL, exports its PDF, waits for completion, and prints the resulting file URL. The task parameters and SDK method follow CloudConvert’s documented operation and quickstart examples. Check the current SDK documentation for import or runtime differences in the version you install.

import CloudConvert from 'cloudconvert';

const apiKey = process.env.CLOUDCONVERT_API_KEY;
if (!apiKey) {
  throw new Error('Set CLOUDCONVERT_API_KEY before running this script.');
}

const cloudConvert = new CloudConvert(apiKey);

const job = await cloudConvert.jobs.create({
  tasks: {
    'capture-page': {
      operation: 'capture-website',
      url: 'https://example.com',
      output_format: 'pdf'
    },
    'export-pdf': {
      operation: 'export/url',
      input: 'capture-page'
    }
  }
});

const completedJob = await cloudConvert.jobs.wait(job.id);
const exportTask = completedJob.tasks.find(
  (task) => task.name === 'export-pdf'
);

if (!exportTask || exportTask.status !== ' finished') {
  throw new Error(`PDF export did not finish. Job: ${job.id}`);
}

for (const file of exportTask.result?.files ?? []) {
  console.log(`${file.filename}: ${file.url}`);
}

In the status check above, use the API’s returned task status as documented by the SDK version in your project. A failed task should be surfaced to your application with its task error details rather than treated as a successful export. The export result contains file information and a downloadable URL. Treat that URL as output from the job, and download or move the file to your own storage if your application needs durable access.

CloudConvert describes the operation as converting a website to PDF or capturing a website screenshot in PNG or JPG. For this tutorial, output_format: "pdf" requests the PDF output. See the Capture Website operation reference.

3. Download the exported PDF

You can download the export URL from your Node.js server. This example writes the bytes to a local file and checks HTTP status before saving:

import { writeFile } from 'node:fs/promises';

async function saveExportedPdf(fileUrl, destination = 'page.pdf') {
  const response = await fetch(fileUrl);
  if (!response.ok) {
    throw new Error(`PDF download failed: HTTP ${response.status}`);
  }

  const bytes = new Uint8Array(await response.arrayBuffer());
  await writeFile(destination, bytes);
  return destination;
}

Call saveExportedPdf(file.url) for an exported file. In a web application, you can instead stream the file to the client or write it to an object-storage integration. Avoid assuming that a job’s export URL is a permanent archive; use storage you control if you need long-term availability.

4. Equivalent cURL request

The API workflow is a job creation request with a capture task and an export task. Replace the key and target URL. Keep the authorization header on the server.

curl -X POST "https://api.cloudconvert.com/v2/jobs" \
  -H "Authorization: Bearer $CLOUDCONVERT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tasks": {
      "capture-page": {
        "operation": "capture-website",
        "url": "https://example.com",
        "output_format": "pdf"
      },
      "export-pdf": {
        "operation": "export/url",
        "input": "capture-page"
      }
    }
  }'

Use a key with the task.write scope for job creation. The response identifies the created job; poll or wait for its completion according to your application flow, then read the export task’s result. See the Jobs reference for job behavior.

5. Python equivalent

Although the main example uses Node.js, the same API v2 job structure can be submitted with Python’s standard HTTP client. Install requests if it is not already available:

python -m pip install requests
import os
import requests

api_key = os.environ["CLOUDCONVERT_API_KEY"]

response = requests.post(
    "https://api.cloudconvert.com/v2/jobs",
    headers={
        "Authorization": f"Bearer {api_key}",
        "Content-Type": "application/json",
    },
    json={
        "tasks": {
            "capture-page": {
                "operation": "capture-website",
                "url": "https://example.com",
                "output_format": "pdf",
            },
            "export-pdf": {
                "operation": "export/url",
                "input": "capture-page",
            },
        }
    },
    timeout=30,
)
response.raise_for_status()
job = response.json()["data"]
print("Created job:", job["id"])

This submits the job; it does not wait for the conversion or download the result. For production, use CloudConvert’s webhook completion flow or implement polling with a bounded timeout and backoff. Do not hold an HTTP request open indefinitely while a remote conversion runs.

6. Capture options and input choices

The capture-website task supports the required url and output_format parameters. Its operation reference also lists optional engine selection, engine version, filename, and timeout parameters. Consult the live operation reference for accepted values and current behavior before setting them.

Choice When to use it What to account for
url input Capture a page available at a public URL. The URL must be reachable by CloudConvert’s service. Authentication, robots, access controls, and client-side rendering can affect whether the page is available and complete.
HTML file input Convert supplied HTML rather than a live page URL. CloudConvert’s HTML-to-PDF offering accepts a URL or HTML file; use the capture-website operation when the input is specifically a website URL.
Engine and engine version Select a rendering engine configuration when your workflow requires one. Confirm currently supported values and rendering guarantees in the operation documentation. Do not assume identical output across engines or versions.
Filename Set the output name for downstream handling. Use a safe filename and do not rely on the extension alone to verify file content.
Timeout Allow a page more time when its load or rendering takes longer. A higher timeout can increase job duration and does not fix a page that is inaccessible or stuck.
export/url Make the output file available through a downloadable URL. Download or store the output if you need to keep it beyond the job workflow.

CloudConvert also documents storage integrations for workflows that should write results to storage. Choose export handling based on how your application serves or archives PDFs.

7. Use asynchronous jobs and webhooks in production

Jobs are asynchronous by default, and CloudConvert recommends webhooks for completion notification. Its Jobs reference notes that synchronous requests may be unsuitable for long-running jobs. A production integration should generally:

  1. Create the job and save its ID with your application’s request or record.
  2. Return promptly to your caller with a pending state or job identifier.
  3. Receive the configured completion webhook and verify it using the current CloudConvert webhook guidance.
  4. Fetch or inspect the job result, check that the export task completed, then download or store the PDF.
  5. Record failure details and provide a retry path that does not accidentally create unlimited duplicate jobs.

Keep webhook handling idempotent: the same completion notification should not trigger duplicate customer-visible work. Set application-level limits for concurrent jobs, retry counts, and how long a job may remain pending. Use polling only when it fits your use case, with a bounded schedule rather than rapid repeated requests.

8. Performance, reliability, and cost

  • Performance: conversion time depends on the remote page and job. Pages with slow responses or extensive client-side work can take longer. The operation’s timeout option can adjust the task limit, but it cannot make an unavailable page render.
  • Reliability: use asynchronous completion notifications for long-running work, inspect task status before using an output, and preserve the job ID and errors for support and retries. Store important PDFs in storage you control.
  • Cost: CloudConvert’s HTML-to-PDF page advertised a starting price of $0.008 per file when researched. This is a vendor price figure, can change, and should not be treated as a guaranteed price for every configuration. Check the live HTML-to-PDF page and pricing before estimating production spend.
  • Operational cost: include your own download, storage, webhook, and retry work in the total cost of a conversion pipeline.

9. Troubleshooting

Symptom Likely cause What to do
Unauthorized or forbidden API response The key is absent, invalid, or lacks the required job-creation scope. Confirm the server environment variable and use an API key with task.write scope. Never print the secret into logs.
Job creation succeeds but no PDF URL appears The job or export task is still running, or the export task failed. Wait for job completion or handle the webhook; inspect the named export task’s status and error before reading result.files.
Capture task fails to load the page The URL is malformed, unreachable from the service, redirects unexpectedly, or requires access the capture process does not have. Check the public URL and redirects, confirm the page can be reached externally, and consult the task error details. Do not assume a local-only URL is accessible remotely.
PDF is blank or incomplete The page may not have finished rendering, depend on unavailable resources, or require a particular rendering configuration. Review the operation’s current engine, version, and timeout options. Check the page’s external dependencies and try a suitable timeout without assuming it resolves blocked content.
Download returns an HTTP error The export URL may be stale, inaccessible, or copied incorrectly. Use the URL from the completed export result, check the response status, and download promptly or move output to durable storage.
Node.js import or top-level await error The project’s module configuration differs from the example. Use the import style supported by the installed SDK and configure the file as an ES module, or wrap the flow in an async function in a CommonJS project.
Request times out while waiting The conversion takes longer than the caller’s request budget. Switch to asynchronous job handling and a webhook, or use bounded polling outside the user-facing request.

10. Or skip the browser setup

If your goal is simply to capture a URL as a PDF, ScreenshotNeo offers a direct API call. See the ScreenshotNeo API documentation for its parameters and response behavior.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. It also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Learn more at ScreenshotNeo.

Create a free account and get 1,000 screenshots a month with no card.

11. Frequently asked questions

Can I convert a page that is only available on my laptop?

A remote capture service needs to reach the URL from its own environment. A localhost address on your machine is not generally the same resource as localhost from a remote service; make the page reachable through an appropriate accessible URL or provide HTML input when that fits the workflow.

Does exporting a URL mean the PDF is stored permanently?

No permanent retention is established by the export URL alone. Download the file or configure a storage integration for the retention your application requires.

Can I use this workflow for screenshots too?

The documented capture operation supports screenshot output formats as well as PDF. Set the output format and follow the current operation reference for supported values and task options.

References