How to Use PDFCrowd from Google Sheets to Convert a Column of URLs to PDFs
Use Google Apps Script and PDFCrowd to convert a spreadsheet column of URLs into Drive PDFs, with per-row status, safe retries, and resumable batches.
To convert a column of URLs into PDFs, use a bound Google Apps Script project to read the sheet, send each valid URL to PDFCrowd’s HTML-to-PDF HTTP API, save successful response bytes as PDF files in Google Drive, and write each result back to its row. The example below processes a bounded batch per run, records failures, and skips rows already marked successful.
PDFCrowd’s documented endpoint is https://api.pdfcrowd.com/convert/24.04/. The request uses HTTP Basic authentication with your PDFCrowd username and API key and sends the source address in a form field named url. PDFCrowd’s servers must be able to reach the source page. See the PDFCrowd HTTP API guide and its parameter reference.
1. Prepare the spreadsheet
Use one row per source URL. Put headers in row 1 and keep the input, output, and status columns together. This example expects:
| Column | Header | Purpose |
|---|---|---|
| A | URL | Source page to convert |
| B | PDF link | Drive URL written after success |
| C | Status | Processing result or error summary |
Create a destination folder in Drive and copy its folder ID from the folder URL. The ID is the portion after /folders/. Decide how reruns should work: the script below skips rows marked SUCCESS; clear that status to intentionally create a new PDF for the row.
2. Store credentials and configure Apps Script
- In the spreadsheet, open Extensions → Apps Script.
- In Apps Script project settings, add script properties named
PDFCROWD_USERNAME,PDFCROWD_API_KEY, andPDF_FOLDER_ID. Store the account username, API key, and destination folder ID as their values. - Paste the code below into the project and save it. Run
convertUrlBatchonce from the editor and approve the requested spreadsheet, Drive, and external request permissions.
Keep the API key out of spreadsheet cells and source URLs. Script properties reduce accidental exposure in the sheet, but editors of the Apps Script project may be able to access project configuration or change code. Restrict project and spreadsheet access to people who should be able to use the credentials.
For projects with an explicit OAuth scope list in appsscript.json, include https://www.googleapis.com/auth/script.external_request along with the required Sheets and Drive scopes. Google documents UrlFetchApp for HTTP requests and responses in its UrlFetchApp reference and URL Fetch Service guide. Apps Script’s spreadsheet service is documented in the SpreadsheetApp reference.
3. Run this Apps Script
const PDFCROWD_ENDPOINT = 'https://api.pdfcrowd.com/convert/24.04/';
const BATCH_SIZE = 10; // Conservative starting point; tune to current quotas and account limits.
const START_ROW = 2;
const URL_COLUMN = 1;
const PDF_LINK_COLUMN = 2;
const STATUS_COLUMN = 3;
function convertUrlBatch() {
const properties = PropertiesService.getScriptProperties();
const username = properties.getProperty('PDFCROWD_USERNAME');
const apiKey = properties.getProperty('PDFCROWD_API_KEY');
const folderId = properties.getProperty('PDF_FOLDER_ID');
if (!username || !apiKey || !folderId) {
throw new Error('Set PDFCROWD_USERNAME, PDFCROWD_API_KEY, and PDF_FOLDER_ID in Script Properties.');
}
const sheet = SpreadsheetApp.getActiveSpreadsheet().getActiveSheet();
const lastRow = sheet.getLastRow();
if (lastRow < START_ROW) return;
// Read the URL and tracking columns together in one spreadsheet operation.
const rowCount = lastRow - START_ROW + 1;
const rows = sheet.getRange(START_ROW, URL_COLUMN, rowCount, 3).getValues();
const folder = DriveApp.getFolderById(folderId);
const auth = Utilities.base64Encode(username + ':' + apiKey);
let processed = 0;
for (let i = 0; i < rows.length && processed < BATCH_SIZE; i++) {
const sheetRow = START_ROW + i;
const source = String(rows[i][0] || '').trim();
const existingStatus = String(rows[i][2] || '').trim();
if (!source || existingStatus === 'SUCCESS') continue;
if (!isHttpUrl_(source)) {
writeRowResult_(sheet, sheetRow, '', 'ERROR: URL must start with http:// or https://');
processed++;
continue;
}
try {
const response = UrlFetchApp.fetch(PDFCROWD_ENDPOINT, {
method: 'post',
contentType: 'application/x-www-form-urlencoded',
payload: { url: source },
headers: { Authorization: 'Basic ' + auth },
muteHttpExceptions: true
});
const code = response.getResponseCode();
if (code < 200 || code >= 300) {
const detail = response.getContentText().slice(0, 500).replace(/\s+/g, ' ');
writeRowResult_(sheet, sheetRow, '', 'ERROR: HTTP ' + code + (detail ? ' — ' + detail : ''));
processed++;
continue;
}
// A successful conversion is binary PDF content. Save the blob, never the text body.
const blob = response.getBlob().setContentType('application/pdf');
const filename = safePdfFilename_(source, sheetRow);
blob.setName(filename);
const file = folder.createFile(blob);
writeRowResult_(sheet, sheetRow, file.getUrl(), 'SUCCESS');
} catch (error) {
writeRowResult_(sheet, sheetRow, '', 'ERROR: ' + String(error).slice(0, 500));
}
processed++;
}
}
function isHttpUrl_(value) {
return /^https?:\/\//i.test(value);
}
function safePdfFilename_(source, rowNumber) {
let host = 'page';
try {
host = new URL(source).hostname.replace(/[^a-z0-9.-]/gi, '_') || host;
} catch (ignored) {}
return host + '-row-' + rowNumber + '.pdf';
}
function writeRowResult_(sheet, rowNumber, link, status) {
sheet.getRange(rowNumber, PDF_LINK_COLUMN, 1, 2).setValues([[link, status]]);
}
The payload object is encoded by Apps Script as form data. The endpoint returns the PDF as response bytes on success. The code checks the HTTP status before creating a Drive file so an API error response is not saved with a .pdf extension. The response status, content, and blob methods are described in Google’s URL Fetch documentation.
4. Process the column in resumable batches
Run convertUrlBatch again to process the next eligible rows. It scans from the top, ignores blank rows and rows marked SUCCESS, and handles up to BATCH_SIZE eligible rows in each execution. Failed rows remain eligible on a later run, so fix the cause before retrying or they will fail again.
- Start with a small batch and adjust based on observed execution duration, current Apps Script execution and URL Fetch quotas, and your PDFCrowd account limits.
- For a large sheet, consider adding a persisted cursor or a time-driven trigger. A cursor avoids scanning completed rows on every run; a trigger can schedule follow-up batches. Keep status writes so you can recover from an interrupted execution.
- Successful rows are skipped. To regenerate a PDF, clear the row’s status; choose whether to retain or clear its existing link before running.
- Files are named from a sanitized source hostname and row number. If rows may be reordered or hostnames repeat, use a stable record ID column in your own naming scheme.
This is deliberately a conservative starting configuration, not a quota guarantee. Google and PDFCrowd limits can depend on the current account, project, and service policies; check their current documentation before scheduling high-volume runs.
5. cURL, Python, and Node.js request equivalents
These examples show the same authenticated form POST for one URL. They save the response only after a successful status code. They do not implement the spreadsheet loop or Drive storage; use the Apps Script example for those parts.
cURL
curl --fail-with-body --user 'PDFCROWD_USERNAME:PDFCROWD_API_KEY' \
--data-urlencode 'url=https://example.com/' \
'https://api.pdfcrowd.com/convert/24.04/' \
--output example.pdf
Python
import os
import requests
endpoint = 'https://api.pdfcrowd.com/convert/24.04/'
response = requests.post(
endpoint,
data={'url': 'https://example.com/'},
auth=(os.environ['PDFCROWD_USERNAME'], os.environ['PDFCROWD_API_KEY']),
timeout=120,
)
if not response.ok:
raise RuntimeError(f'PDFCrowd HTTP {response.status_code}: {response.text[:500]}')
with open('example.pdf', 'wb') as pdf_file:
pdf_file.write(response.content)
Node.js
const username = process.env.PDFCROWD_USERNAME;
const apiKey = process.env.PDFCROWD_API_KEY;
if (!username || !apiKey) throw new Error('Set PDFCROWD_USERNAME and PDFCROWD_API_KEY');
const body = new URLSearchParams({ url: 'https://example.com/' });
const auth = Buffer.from(`${username}:${apiKey}`).toString('base64');
const response = await fetch('https://api.pdfcrowd.com/convert/24.04/', {
method: 'POST',
headers: {
Authorization: `Basic ${auth}`,
'Content-Type': 'application/x-www-form-urlencoded',
},
body,
});
if (!response.ok) {
throw new Error(`PDFCrowd HTTP ${response.status}: ${(await response.text()).slice(0, 500)}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('example.pdf', bytes));
Do not put credentials directly into scripts that will be committed or shared. Use environment variables or a protected secret store appropriate to your runtime.
6. URL access and conversion options
The API converts a web page URL, not a local file path. PDFCrowd fetches the page and its resources from its servers. Use a fully qualified http:// or https:// URL that is reachable from the public internet. A page on localhost, an intranet-only host, or a page that requires a login may not be accessible. The PDFCrowd API key authenticates your API request; it does not sign you into the source website.
For conversion controls, consult the versioned API’s current parameter reference. The required workflow here uses the url form field; optional conversion settings should be added as documented form fields. Confirm the currently supported parameter names and values before relying on a setting in a batch. The endpoint’s 24.04 segment is the documented API version used in this guide; check PDFCrowd’s current docs before starting a new production integration or changing versions.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Apps Script asks for authorization or reports a permission error | The script has not been authorized, or an explicit OAuth scope list omits external requests or required Sheets/Drive access. | Run the function from the editor and grant the required scopes. Add https://www.googleapis.com/auth/script.external_request when managing explicit scopes. |
| HTTP 401 or 403 from PDFCrowd | Incorrect account username/API key, revoked credentials, or an account-level restriction. | Check the username and API key in Script Properties and confirm account access with PDFCrowd. Do not add the API key to the url field. |
| HTTP 400 or another conversion error | Malformed input, unsupported request parameters, or a source URL PDFCrowd cannot process. | Inspect the status and truncated error text in the row. Check the API parameter reference; test the URL independently and correct the input or settings. |
| The source works in your browser but conversion fails | The page may be private, login-protected, blocked, or otherwise unreachable from PDFCrowd’s servers. | Check public reachability and any source-site access requirements. API credentials do not provide source-site credentials. |
| A Drive file contains an error page or is not a valid PDF | Error response bytes were saved as a file. | Use status checking before file creation, as in the script. Remove the bad file, fix the request, clear the row status, and rerun. |
| Duplicate PDFs appear | A row was reset or marked unsuccessful after a file was created but before the sheet link was written. | Check the destination folder before retrying. Establish whether reruns should replace, retain, or create another file; the sample creates a new file for each successful retry. |
| Execution times out or stops partway | The batch is too large or a conversion takes too long for the current execution limit. | Reduce BATCH_SIZE, rerun to continue, and consider a cursor or scheduled batches. Do not assume a fixed row count is safe across accounts. |
| Invalid URL status | The cell is blank, misspelled, or lacks an HTTP(S) scheme. | Enter a complete URL beginning with http:// or https://. The script skips blanks. |
For more detailed diagnosis, PDFCrowd documents optional JSON error formatting and debug information in its API documentation. Keep logs useful but avoid writing secrets or sensitive source content into status cells.
8. Performance, reliability, and cost considerations
- Reduce spreadsheet chatter. The script reads the data range once and writes only the result cells for processed rows. For very large sheets, collect row updates and write them in groups where practical.
- Bound each run. Conversion time varies with source pages and resources. Small batches make failures and timeouts easier to recover from. Determine batch size from current Google quotas, execution limits, and your PDFCrowd account limits.
- Retry deliberately. A timeout or transient server issue may justify a later retry. Invalid URLs, authentication failures, and inaccessible pages need correction first. Avoid immediate blind retry loops that consume execution time and API allowance.
- Keep a durable row record. Write the Drive link and success marker after file creation. If execution stops between creating the file and writing the row, inspect the destination folder before retrying to prevent duplicates.
- Plan storage and service costs. Each successful conversion creates a Drive file and uses PDFCrowd service capacity under the applicable account plan. This dossier does not establish current quotas, prices, or a safe batch size; check the current plan and service limits before processing a large column.
Or skip the browser setup
If you need a screenshot or PDF of a page without maintaining a browser capture setup, ScreenshotNeo is a website screenshot API and MCP server for developers. Its API accepts one GET request with a URL and can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free and get 1,000 screenshots a month with no card.
FAQ
Can I convert a URL that is only available on my computer?
No. The source page must be reachable from PDFCrowd’s servers. A local localhost address is not reachable by the service.
Does the script make one combined PDF?
No. It creates one PDF per eligible URL and records each file’s Drive link on the corresponding row.
Can I safely share this spreadsheet?
Only share it with people authorized to see the URLs and use the bound script. Keep the API key in protected project configuration, and restrict who can edit the Apps Script project.
Does a failed row stop the rest of the batch?
The sample records an error and moves on to the next row. Correct the cause and rerun; failed rows are not marked successful and remain eligible.


