How to Schedule Recurring Webpage PDFs with PDFShift and Cron
Schedule PDFShift webpage conversions with cron, securely manage API keys, save PDFs durably, and handle time zones, retries, and failures.
Direct answer: Cron schedules a script, and the script calls PDFShift’s conversion API at each scheduled time. The API request uses POST https://api.pdfshift.io/v3/convert/pdf, sends the API key in the X-API-Key header, and supplies a webpage URL or HTML in the JSON source field. Check the response, save the PDF to durable storage, and report failures through logs or alerts. [PDFShift API documentation]
This guide uses Bash, curl, and cron on a Linux host. It also includes a Python implementation and the direct PDFShift API request in Node.js. Use the variant that fits your environment; the schedule and operational requirements are the same.
1. Decide how the scheduled capture should work
Before adding a cron entry, make four choices:
- Schedule and time zone: decide whether the run follows the host’s local clock or an explicitly configured time zone. Daylight-saving changes can skip a local time or run it twice, depending on the transition.
- Input: submit a URL when PDFShift should fetch the page. Submit raw HTML when your process already has the markup or needs to avoid having PDFShift fetch the source page.
- Output: accept the PDF directly in the response, or request a named file and receive a temporary download URL. Download temporary files promptly and copy them to storage you control.
- Failure policy: decide how many retries are appropriate, where logs go, how failures are surfaced, and what should happen if a run is repeated.
PDFShift’s integration guide says a filename option changes the response to JSON containing a download URL, which remains available for two days. The direct response avoids that extra download step. [PDFShift Make integration guide]
2. Store the PDFShift key safely
PDFShift documents X-API-Key as the API key header; its help article says this replaced the older Basic Auth method on 2025-05-06. Keep the key out of source control, public client-side code, and logs. Set it in the environment of the account that runs the job, or use the host’s secrets facility. [PDFShift authentication guidance]
For a user crontab, one simple setup is to put the key in a separate environment file readable only by the scheduled account:
# Create this file outside the repository, for example /home/runner/.config/pdfshift.env
PDFSHIFT_API_KEY=replace_with_your_key
Restrict access to the file using your host’s file permission tools. The script below reads it without printing the secret. Cron has a limited environment, so explicitly load required variables and use absolute paths.
3. Create a Bash and curl capture script
Save this as /absolute/path/to/capture-page.sh. Set PAGE_URL and OUTPUT_DIR for your job. This example asks PDFShift to return the PDF response directly, writes to a temporary file first, then moves it into place only after a successful HTTP response.
#!/bin/sh
set -eu
ENV_FILE="/home/runner/.config/pdfshift.env"
OUTPUT_DIR="/home/runner/pdf-archive"
PAGE_URL="https://example.com/report"
if [ ! -r "$ENV_FILE" ]; then
echo "Cannot read PDFShift environment file: $ENV_FILE" >&2
exit 1
fi
. "$ENV_FILE"
if [ -z "${PDFSHIFT_API_KEY:-}" ]; then
echo "PDFSHIFT_API_KEY is missing" >&2
exit 1
fi
mkdir -p "$OUTPUT_DIR"
# Use UTC in the filename so it remains unambiguous across host time zones.
STAMP=$(date -u +%Y%m%dT%H%M%SZ)
FINAL="$OUTPUT_DIR/report-$STAMP.pdf"
TEMP="$OUTPUT_DIR/.report-$STAMP.pdf.part"
cleanup() {
rm -f "$TEMP"
}
trap cleanup EXIT HUP INT TERM
HTTP_STATUS=$(curl --silent --show-error \
--output "$TEMP" \
--write-out '%{http_code}' \
--request POST 'https://api.pdfshift.io/v3/convert/pdf' \
--header "X-API-Key: $PDFSHIFT_API_KEY" \
--header 'Content-Type: application/json' \
--data "{\"source\":\"$PAGE_URL\"}" \
--max-time 180)
case "$HTTP_STATUS" in
2[0-9][0-9]) ;;
*)
echo "PDFShift returned HTTP $HTTP_STATUS" >&2
# Preserve a short response excerpt for diagnosis, without exposing the key.
head -c 2000 "$TEMP" >&2 || true
exit 1
;;
esac
if [ ! -s "$TEMP" ]; then
echo "PDFShift returned an empty response" >&2
exit 1
fi
mv "$TEMP" "$FINAL"
echo "Saved PDF: $FINAL"
The JSON string construction above is suitable for a fixed, controlled URL without quote characters. For arbitrary URLs, build JSON with a JSON-aware tool such as Python rather than interpolating input into JSON. For stricter validation, check the saved response’s PDF signature before moving it into the archive.
Make the script executable and run it once as the same account cron will use. Use the absolute path to curl and to the script if the host’s PATH is uncertain. Do not add the key to the crontab command itself.
Schedule weekdays at 08:00
Open the scheduled account’s crontab with crontab -e and add this illustrative entry:
0 8 * * 1-5 /absolute/path/to/capture-page.sh >> /absolute/path/to/capture-page.log 2>&1
Linux user crontab lines have five time and date fields followed by a command; cron checks entries every minute. Confirm syntax and time-zone behavior for the cron implementation on your host. [Linux crontab manual]
4. Make the schedule and output reliable
Use explicit paths and permissions
Cron typically starts with a small environment and, for the Linux implementation described in its manual, runs commands through /bin/sh by default. Ensure the scheduled account can read the script and secret file and write to the archive and log directories. Set any necessary environment variables inside the script or crontab. [Linux crontab environment notes]
Choose a time-zone policy
Many cron implementations use the machine’s local time. The Linux manual documents CRON_TZ for setting a crontab time zone and warns that daylight-saving transitions can cause a nonexistent local time not to run and a repeated local time to run twice. Verify that your cron supports the setting before relying on it. If a missed or duplicate report matters, use a UTC schedule or make the job safe to rerun and record the intended capture date separately from the execution time. [Linux crontab time-zone and daylight-saving notes]
Make reruns safe
Timestamped output prevents a rerun from overwriting an earlier PDF, but it can create duplicates. If there should be one report per intended date, derive the final filename from that scheduled date and decide whether a rerun should replace, skip, or version the existing file. For concurrent or manually retried runs, use a lock or another single-run mechanism supported by your host.
Alert on failure
The example exits nonzero on a failed request, so the log records the failure. Configure an alerting mechanism appropriate to your host. The cron manual describes MAILTO for command output, but delivery depends on mail being configured. Do not treat an attempted conversion as success until the PDF reaches its intended destination. [Linux crontab manual]
5. Python implementation
This version uses the requests package. Install it in the Python environment used by the scheduled account, then save the script with a fixed URL and output directory. The key is read from the process environment.
import os
from pathlib import Path
from datetime import datetime, timezone
import requests
API_URL = "https://api.pdfshift.io/v3/convert/pdf"
PAGE_URL = "https://example.com/report"
OUTPUT_DIR = Path("/home/runner/pdf-archive")
api_key = os.environ.get("PDFSHIFT_API_KEY")
if not api_key:
raise RuntimeError("PDFSHIFT_API_KEY is missing")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
final_path = OUTPUT_DIR / f"report-{stamp}.pdf"
temp_path = OUTPUT_DIR / f".report-{stamp}.pdf.part"
try:
response = requests.post(
API_URL,
headers={"X-API-Key": api_key},
json={"source": PAGE_URL},
timeout=(15, 180),
)
response.raise_for_status()
if not response.content:
raise RuntimeError("PDFShift returned an empty response")
temp_path.write_bytes(response.content)
temp_path.replace(final_path)
print(f"Saved PDF: {final_path}")
except requests.RequestException as exc:
# Do not print request headers or secrets.
detail = getattr(exc.response, "text", "") if getattr(exc, "response", None) is not None else ""
raise RuntimeError(f"PDFShift request failed: {exc}; response: {detail[:1000]}") from exc
finally:
if temp_path.exists():
temp_path.unlink()
Run it through an absolute interpreter path in cron, or activate the intended virtual environment in a wrapper script. Do not assume an interactive shell’s environment is available to cron.
6. Node.js API request
This example uses built-in fetch on a Node.js runtime that provides it. It writes the direct PDF response after checking the HTTP status.
import { mkdir, writeFile, rename, unlink } from 'node:fs/promises';
const apiKey = process.env.PDFSHIFT_API_KEY;
if (!apiKey) throw new Error('PDFSHIFT_API_KEY is missing');
const pageUrl = 'https://example.com/report';
const outputDir = '/home/runner/pdf-archive';
const stamp = new Date().toISOString().replace(/[-:.]/g, '');
const finalPath = `${outputDir}/report-${stamp}.pdf`;
const tempPath = `${outputDir}/.report-${stamp}.pdf.part`;
await mkdir(outputDir, { recursive: true });
try {
const response = await fetch('https://api.pdfshift.io/v3/convert/pdf', {
method: 'POST',
headers: {
'X-API-Key': apiKey,
'Content-Type': 'application/json',
},
body: JSON.stringify({ source: pageUrl }),
signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
const detail = (await response.text()).slice(0, 1000);
throw new Error(`PDFShift returned HTTP ${response.status}: ${detail}`);
}
const pdf = Buffer.from(await response.arrayBuffer());
if (pdf.length === 0) throw new Error('PDFShift returned an empty response');
await writeFile(tempPath, pdf, { flag: 'wx' });
await rename(tempPath, finalPath);
console.log(`Saved PDF: ${finalPath}`);
} finally {
await unlink(tempPath).catch(() => {});
}
Supply PDFSHIFT_API_KEY through the service manager or secret store that launches the process. Do not put it in a checked-in file or print it as part of error diagnostics.
7. URL, raw HTML, authentication, and PDF options
URL source
Send {"source":"https://example.com/report"} when PDFShift should retrieve and render the page. The page must be reachable by the conversion service. A page available only on a private network or behind an interactive sign-in flow may not be reachable this way.
Raw HTML source
Send HTML in the source field when your application already has the markup. PDFShift’s Node guide recommends raw HTML when available because it avoids PDFShift fetching the source page itself; it also recommends inline styles and scripts to reduce external resource requests. That is vendor guidance, not a guarantee of faster or identical rendering for every page. [PDFShift Node guide]
curl --silent --show-error \
--request POST 'https://api.pdfshift.io/v3/convert/pdf' \
--header "X-API-Key: $PDFSHIFT_API_KEY" \
--header 'Content-Type: application/json' \
--data-binary '{"source":"<html><body><h1>Scheduled report</h1></body></html>"}' \
--output report.pdf
For generated HTML or values containing quotes, create the JSON with a JSON encoder rather than hand-escaping it in shell. External stylesheets, fonts, and images still need to be reachable unless embedded or inlined.
Authenticated pages
PDFShift’s integration guide documents authentication, cookies, and custom HTTP headers as ways to handle pages that require access. Use narrowly scoped credentials, keep them in a protected secret store, and avoid logging sensitive page contents. Do not assume a browser-only sign-in flow can be reproduced by a server-side conversion request. [PDFShift integration guide]
Direct PDF or temporary URL
The examples above save the direct PDF response. If you use the documented filename option to receive JSON with a download URL, the URL is temporary and the guide says it remains available for two days. Download it as part of the same scheduled workflow and move it to storage whose retention and access controls you manage. [PDFShift integration guide]
PDF layout settings
PDFShift supports conversion options in its API. Consult its current API reference for exact option names and accepted values before adding page size, margins, orientation, headers, footers, or other layout settings. Keep the scheduled job’s options explicit so output changes are reviewable. [PDFShift API documentation]
8. cURL reference request
Here is the underlying request independent of the cron wrapper. It writes the API response to report.pdf; production scripts should also check the status and handle the file atomically as shown above.
curl --fail-with-body --silent --show-error \
--request POST 'https://api.pdfshift.io/v3/convert/pdf' \
--header "X-API-Key: $PDFSHIFT_API_KEY" \
--header 'Content-Type: application/json' \
--data '{"source":"https://example.com/report"}' \
--output report.pdf
9. Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Job works interactively but not in cron | Cron has a limited environment, different working directory, or different shell. | Use absolute paths, load required variables explicitly, set the working directory in the script, and test as the cron account. |
| Missing or invalid API key | The secret file was not readable, the variable was not exported/loaded, or the header was omitted. | Check file permissions and variable presence without printing its value; send the key as X-API-Key. |
| HTTP error from conversion endpoint | Authentication, request shape, source URL access, or a conversion problem. | Record the status and a bounded response excerpt without headers or secrets. Verify the endpoint, JSON content type, source, and page reachability. |
| PDF is empty, truncated, or invalid | Response handling saved an error body, the process was interrupted, or output was written in place. | Check HTTP status before accepting output, write to a temporary file, confirm it is nonempty, then rename it into place. |
| Page content or assets are missing | Resources require authentication, are blocked from external fetching, or load dynamically. | Check page and asset accessibility; consider raw HTML with inline resources where appropriate, or use documented authentication, cookies, or headers. |
| No output file after requesting a filename | The response is JSON with a temporary URL rather than PDF bytes. | Parse the JSON, download the URL during the job, and persist the downloaded file before the URL expires. |
| Run is missing or duplicated near a clock change | The scheduled local time does not exist or occurs twice during daylight-saving transition. | Choose a deliberate time zone, consider UTC, and make output naming and rerun behavior idempotent. |
| Two overlapping runs overwrite or duplicate output | A slow conversion overlaps the next schedule or a retry runs concurrently. | Use a lock, unique job identity, or date-based idempotency rule; do not assume every run finishes within one interval. |
10. Performance, reliability, and cost
- Conversion time: rendering depends on the page, its assets, and the conversion service. Set a request timeout long enough for expected pages, but finite so a stuck job can fail and alert.
- Retries: retry transient network errors and server failures with a small bounded backoff. Avoid retrying authentication or malformed-request errors unchanged. Since conversion may have succeeded while the client lost the response, use deterministic output identity or record the intended run before retrying.
- Storage: account for PDF size, retention, backups, and access controls. A temporary vendor URL is not an archive. Verify that the final file reached the destination before reporting success.
- Scheduling: if the cron period is shorter than the slowest expected conversion, prevent overlapping runs. Add monitoring for both missed runs and failed deliveries.
- Cost: check PDFShift’s current plan, conversion limits, and billing terms for your expected schedule. The research available for this guide does not establish current pricing, so no price or conversion allowance is stated here.
Or skip the browser setup
If the deliverable can be an image of the page rather than a PDF, ScreenshotNeo provides a one-request website screenshot API. It returns PNG, JPEG, or WebP, so it is an alternative for recurring visual snapshots rather than a PDF converter. See the ScreenshotNeo API docs for options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/report -o shot.webp
Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, no card required.
FAQ
Does cron convert the webpage to PDF?
No. Cron triggers the command on a schedule. The script makes the PDFShift API request and handles the returned file.
Can I keep PDFs generated through a temporary download URL?
Yes, but download the file while the URL is valid and copy it to storage you control. PDFShift’s integration guide says its generated URL remains available for two days.
Can I use this workflow for image screenshots?
Yes, if an image is the right output. ScreenshotNeo offers an image screenshot endpoint; it is not a substitute for PDFShift when the required artifact is a PDF.


