How to Automate ShrinkTheWeb Captures with a Cron Job
Schedule screenshot captures with cron, make failures visible, and verify the current ShrinkTheWeb API contract before relying on it.
Cron schedules the work; a script performs the capture. To automate ShrinkTheWeb captures, first confirm its current API endpoint, authentication method, request parameters, and response format in current vendor documentation. The available references do not verify those details, so do not copy old Drupal settings into a new API client as if they were a current contract.
The reliable pattern is: write a script that makes the documented capture request, validates and saves the result, run it manually as the account that will own the crontab, then schedule it with absolute paths and logs. Cron does not render the page or check that the resulting image is usable.
1. Confirm the current ShrinkTheWeb API contract
Before writing the request, obtain current official ShrinkTheWeb instructions and verify that the service and your account are available. The research for this guide did not establish a current endpoint, authentication scheme, request syntax, response format, limits, or account terms.
Historical Drupal integration guides from 2019 describe profile Access and Secret keys, caching duration, and options such as width, full-length captures, maximum height, viewport dimensions, post-load delay, and image quality. These are historical integration details, not a current API specification. Some options were account-tier dependent. The Drupal project was later marked as appearing unsupported; that does not establish whether the ShrinkTheWeb service itself is available today. [Drupal project status]
From current vendor documentation, record:
- The HTTPS endpoint and HTTP method.
- How credentials are sent, and whether they are access keys, signed parameters, headers, or another mechanism.
- Required URL and capture parameters, including supported dimensions or full-page behavior.
- Whether the response is image bytes, JSON containing a URL, or another format.
- How errors, rate limits, timeouts, and account limits are reported.
- Any retention, caching, or usage terms that affect your schedule.
Do not put credentials in a URL that might be logged unless the vendor specifically requires that method and you have accounted for log exposure. Prefer a restricted environment file or secret manager, and keep it readable only by the job account.
2. Write a capture script around the documented request
The capture script should own the target URL, approved capture options, output destination, timeouts, and failure handling. The example below is deliberately an implementation template: replace the marked request construction and response handling with the exact current ShrinkTheWeb contract. It is not runnable as a ShrinkTheWeb API client until those vendor-specific details are filled in.
#!/usr/bin/env python3
"""Template: fill in request details from current ShrinkTheWeb documentation."""
import os
import sys
from pathlib import Path
import requests
API_ENDPOINT = os.environ.get("SHRINKTHEWEB_ENDPOINT")
ACCESS_KEY = os.environ.get("SHRINKTHEWEB_ACCESS_KEY")
SECRET_KEY = os.environ.get("SHRINKTHEWEB_SECRET_KEY")
TARGET_URL = "https://example.com/"
OUTPUT = Path("/var/lib/site-captures/example-com.png")
if not API_ENDPOINT or not ACCESS_KEY:
raise SystemExit("Set SHRINKTHEWEB_ENDPOINT and SHRINKTHEWEB_ACCESS_KEY")
# Replace this with the current documented auth and parameter names.
params = {
"url": TARGET_URL,
"access_key": ACCESS_KEY,
}
# Add a secret only as directed by the vendor's current documentation.
if SECRET_KEY:
params["secret_key"] = SECRET_KEY
try:
response = requests.get(API_ENDPOINT, params=params, timeout=(10, 90))
response.raise_for_status()
except requests.RequestException as exc:
print(f"Capture request failed: {exc}", file=sys.stderr)
raise SystemExit(1)
# Replace with the documented response validation. If the API returns JSON
# containing an image URL, fetch that URL and validate its response instead.
content_type = response.headers.get("Content-Type", "")
if not content_type.startswith("image/"):
print(f"Expected image response, got {content_type!r}", file=sys.stderr)
raise SystemExit(1)
OUTPUT.parent.mkdir(parents=True, exist_ok=True)
tmp = OUTPUT.with_suffix(OUTPUT.suffix + ".tmp")
tmp.write_bytes(response.content)
tmp.replace(OUTPUT)
print(f"Saved {OUTPUT} ({len(response.content)} bytes)")
The temporary-file-then-rename pattern prevents a consumer from seeing a partially written image. Validate the actual format or image dimensions too if the API can return an error image or an unexpected payload with an image content type.
3. Run it manually as the cron account
- Install the runtime and dependencies where the scheduled account can access them. Use a virtual environment or a packaged script if that is how your server manages Python dependencies.
- Create the output directory and ensure the job account can write to it.
- Provide credentials to that account through a protected environment file or secret manager. Interactive shell variables may not be present in cron.
- Run the script using absolute paths as the same user that will own the crontab.
- Check the exit status, logs, file timestamp, file size, image format, and whether the page content is actually present.
For example, after adapting the template and setting the documented endpoint and credentials in the account environment:
/usr/bin/python3 /opt/site-captures/capture.py
4. Add the cron schedule
Edit the crontab for the same operating-system account used in the manual run:
crontab -e
A five-field cron schedule has this shape:
# minute hour day-of-month month day-of-week
15 6 * * * /usr/bin/python3 /opt/site-captures/capture.py >> /var/log/site-captures/capture.log 2>&1
This example runs daily at 06:15 in the cron host’s local time. Choose a frequency that matches how often the page changes and how many captures your account allows. Cron implementations vary; consult the host’s manual for timezone, daylight-saving, and syntax behavior.
For multiple pages, use a configuration file or a loop in the script, log each target separately, and return a nonzero exit code if any required capture fails. Avoid launching overlapping copies if a slow request could still be running when the next schedule starts. A lock file or a single worker can prevent duplicate runs.
5. Configure output, logs, and failure visibility
- Use absolute paths. Cron may start with a minimal environment and an unexpected working directory.
- Set explicit timeouts. Use connect and read timeouts appropriate to the vendor’s documented response behavior; a request should not hang indefinitely.
- Preserve useful errors. Send standard output and standard error to a log, and rotate the log so it cannot grow without limit.
- Return failure status. Treat HTTP errors, malformed responses, authentication failures, and invalid image output as failures rather than silently saving them.
- Alert on repeated failures. Use your existing monitoring or job runner to surface nonzero exits and missing output. Cron itself is a trigger, not an alerting or retry service.
- Make writes safe. Write to a temporary path and rename only after validation. Keep prior good captures if a new request fails.
- Protect credentials and captures. Restrict file permissions and avoid printing secrets or full credential-bearing request URLs in logs.
6. Capture settings and edge cases
Use only settings documented for the current API and account. Historical integrations suggest dimensions, full-page capture, maximum height, viewport, delay, quality, and caching may be relevant concepts, but their current names, availability, and behavior must be checked with the vendor.
- Slow or dynamic pages: A page may need a documented post-load wait. Too little wait can miss content; too much increases job duration and the chance that runs overlap.
- Long pages: Full-page output can be much larger and slower than a viewport capture. Confirm size and height limits.
- Authentication and private pages: Confirm whether the service supports the required cookies or headers. Do not assume it can capture pages behind a login.
- Redirects and URL encoding: Ensure the request library encodes the target URL correctly and that the vendor permits the final destination.
- Transient failures: If retrying, use a small bounded retry count with backoff for transient network errors or documented retryable status codes. Avoid retrying invalid credentials or unsupported parameters.
- Repeated schedules: A run lasting longer than its interval can cause duplicate work. Use locking or a queue and decide whether missed runs should be skipped or caught up.
- Output naming: Use stable names for the latest capture, or timestamped names if history matters. Define a retention policy for old files.
7. Troubleshooting
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Works in a terminal but not in cron | Different PATH, working directory, environment, user, or permissions. | Use absolute executable and file paths; run under the crontab owner; load secrets explicitly; log stdout and stderr. |
| No output file appears | The request failed, the script exited early, or the destination is not writable. | Inspect the exit code and log, verify directory ownership, and confirm the API response before writing. |
| Authentication or parameter error | Old integration instructions or guessed parameter names do not match the current API. | Check current vendor documentation and account settings. Do not rely on 2019 Drupal examples for current syntax. |
| Saved file is JSON or HTML | The API returned an error or a URL/status payload rather than image bytes. | Inspect status, content type, and response structure; implement the documented response flow and validate before replacing the image. |
| Image is blank or incomplete | Page load timing, blocked resources, redirects, or a capture limitation. | Check the target in a browser, then adjust only supported wait or capture settings and inspect the returned result. |
| Requests time out or runs overlap | Slow network, slow page rendering, or a schedule interval shorter than job duration. | Set bounded timeouts, choose a less frequent schedule, add a lock, and use bounded retries only for transient errors. |
| Output is unexpectedly large | Full-page capture or high dimensions/quality. | Use the smallest documented dimensions and quality that meet the use case; confirm current service limits. |
| Capture stops working after a service change | Endpoint, authentication, limits, or account terms changed. | Recheck current official documentation and monitor failures rather than assuming the old integration remains supported. |
8. Performance, reliability, and cost
Capture time depends on the target page, network, selected capture settings, and service behavior; no current ShrinkTheWeb timing or pricing figures were verified for this guide. Avoid setting a schedule more frequently than the application needs. For many URLs, account for total run duration, service limits, output storage, and concurrency before choosing an interval.
Reliability comes from validating every response, retaining the previous good artifact when a run fails, logging enough to diagnose failures, and alerting when expected outputs are missing. Set retries based on documented error behavior rather than retrying every failure indiscriminately.
For cost planning, confirm current account tiers, included usage, overage rules, and whether options such as full-page capture affect usage with ShrinkTheWeb directly. The historical Drupal docs note account-tier-dependent options, but do not establish present-day pricing or limits.
9. Or skip the browser setup
ScreenshotNeo is a screenshot API and MCP server. A single GET request captures a URL; use the current API documentation for the complete parameter set and error handling.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/ \
-o shot.webp
Use the same request from Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Or from Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
With ScreenshotNeo, cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
FAQ
Does cron take the screenshot?
No. Cron starts the script on a schedule. The script calls a capture mechanism, checks the result, and stores or forwards it.
Can I assume the old Drupal module describes today’s API?
No. Those integration guides are from 2019 and do not verify current ShrinkTheWeb endpoints, credentials, parameters, or account terms.
Does an unsupported Drupal module mean ShrinkTheWeb is unavailable?
No. The project status applies to that integration project, not necessarily the underlying service. Confirm the service directly with current vendor materials.
What should I do if the API returns a link instead of image bytes?
Follow the current documented response flow: validate the API result, fetch the returned image URL if required, then validate and save that image.
Sources
- Drupal.org ShrinkTheWeb project page for the integration project’s status.
- Historical Drupal integration guides dated March 4, 2019 describe old settings and credential fields; they are not used here as a current API contract.


