ScreenshotNeo

BlogHow-to

How to Schedule Daily Website Screenshots with PagePeeker

PagePeeker offers automated snapshots through its premium service, but its public API docs do not describe a daily scheduler. Here are the documented options and a careful API-based workaround.

By the ScreenshotNeo team4 October 202610 min read

Short answer: PagePeeker says its premium site-snapshot service can capture websites on an automated schedule or on demand. Its public site describes weekly or more frequent snapshots, and advises against capturing more often than once per day. However, its public API documentation does not show a daily recurrence setting or self-serve scheduling instructions. To arrange daily captures as a managed PagePeeker service, contact PagePeeker or check your premium account and confirm the cadence and delivery details for your account.

If you want to build your own recurring job, you can schedule calls to PagePeeker’s V2 thumbnail API with an external scheduler. That is an API-caller workflow, not a documented PagePeeker daily scheduler, and PagePeeker does not publicly establish that it provides the same archive or retention as its managed snapshot service.

1. Choose how to schedule the captures

Managed premium site snapshots

PagePeeker describes site snapshots as a premium service that can run automatically or on demand. Its recommendation is to capture once per week or more often, but no more frequently than once per day. The public materials do not specify the exact setup steps, supported time zones, delivery method, archive retention, or whether customers can configure the schedule themselves. Ask PagePeeker to confirm those details, the daily cadence, and current plan terms before relying on it.

Your own scheduled API caller

The V2 API documents thumbnail requests, a premium refresh=1 parameter, and a readiness-check sequence. It does not document a recurring schedule parameter. You can run a script once a day with cron, a CI scheduler, or another job runner, but confirm with PagePeeker how cache behavior and API usage apply to your account. Do not assume these API calls create the same managed snapshot archive.

2. Request a PagePeeker V2 thumbnail

The documented V2 direct-link pattern accepts size, optional code, refresh, wait, and url. The exact options available depend on the account. PagePeeker identifies V2 as its current recommended API; V1 is discontinued and redirects to V2. Use the V2 endpoint directly.

The examples below show the generation request. For a fresh capture, PagePeeker documents premium refresh=1. Do not use it unless your account supports it. PagePeeker also documents requesting generation, polling readiness, and then fetching the image; see the readiness section below if you need to avoid treating a pending result as a finished screenshot.

cURL

curl -L --get 'http://api.pagepeeker.com/v2/thumbs.php' \
  --data-urlencode 'size=m' \
  --data-urlencode 'code=YOUR_PAGEPEEKER_CODE' \
  --data-urlencode 'refresh=1' \
  --data-urlencode 'url=https://example.com/' \
  --output screenshot.jpg

Omit code if your account does not require it, and omit refresh=1 unless your account supports premium refresh. URL-encoding the complete target URL matters when it includes its own query string or special characters.

Python

import os
from pathlib import Path
from urllib.parse import urlencode

import requests

endpoint = "http://api.pagepeeker.com/v2/thumbs.php"
params = {
    "size": "m",
    "url": "https://example.com/",
}
code = os.environ.get("PAGEPEEKER_CODE")
if code:
    params["code"] = code

# Set this only when the account supports premium refresh.
if os.environ.get("PAGEPEEKER_REFRESH") == "1":
    params["refresh"] = "1"

response = requests.get(endpoint, params=params, timeout=90)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if not content_type.startswith("image/"):
    raise RuntimeError(
        f"Expected an image, got {content_type}: {response.text[:500]}"
    )
Path("screenshot.jpg").write_bytes(response.content)

Install the dependency with python -m pip install requests. Keep account codes in environment variables or a secret store, not in a public script or client-side page.

Node.js

const endpoint = new URL("http://api.pagepeeker.com/v2/thumbs.php");
endpoint.searchParams.set("size", "m");
endpoint.searchParams.set("url", "https://example.com/");
if (process.env.PAGEPEEKER_CODE) {
  endpoint.searchParams.set("code", process.env.PAGEPEEKER_CODE);
}
if (process.env.PAGEPEEKER_REFRESH === "1") {
  endpoint.searchParams.set("refresh", "1");
}

const response = await fetch(endpoint);
if (!response.ok) {
  throw new Error(`PagePeeker returned HTTP ${response.status}`);
}
const contentType = response.headers.get("content-type") || "";
if (!contentType.startsWith("image/")) {
  throw new Error(`Expected an image, got ${contentType}: ${await response.text()}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await import("node:fs/promises").then(({ writeFile }) =>
  writeFile("screenshot.jpg", bytes)
);

This uses the Node.js built-in fetch available in current Node releases. Set PAGEPEEKER_CODE only if your account uses a code, and set PAGEPEEKER_REFRESH=1 only if premium refresh is enabled.

3. Run a capture once per day

On a Unix-like host with cron, save a script that performs the API request, make it executable, then schedule it. This example runs daily at 02:15 in the host’s configured time zone:

15 2 * * * /usr/bin/env bash /opt/snapshots/capture-daily.sh >> /var/log/pagepeeker-daily.log 2>&1

Make sure the script uses an absolute output path and has access to required environment variables. If the host uses UTC but you expect local time, account for that in the schedule. Cron syntax and time-zone behavior vary by scheduler, so verify the job’s configured time zone and inspect its execution logs.

A recurring task should handle failures deliberately: record the target URL, scheduled time, response status, and whether a valid image was saved; use bounded retries with backoff for temporary network errors; and alert or surface missed runs. Avoid blindly retrying every error, since API calls and readiness checks can count toward usage.

4. Use the readiness-check sequence for asynchronous captures

PagePeeker documents a readiness endpoint at /v2/thumbs_ready.php. Its documented flow is to request thumbnail generation, check readiness, and then retrieve the image. The readiness response contains Error and IsReady. PagePeeker’s public docs do not prescribe a polling interval or retry policy, so choose a finite wait budget for your job and stop if the response reports an error.

from time import sleep
from urllib.parse import urlencode

import requests

base = "http://api.pagepeeker.com"
site_url = "https://example.com/"
params = {"size": "m", "url": site_url}
if code := os.environ.get("PAGEPEEKER_CODE"):
    params["code"] = code

# For premium accounts that support it, request a fresh render first.
if os.environ.get("PAGEPEEKER_REFRESH") == "1":
    params["refresh"] = "1"

# Start/request the thumbnail. Check the response according to your account's
# API behavior; do not assume the body is an image until it is ready.
requests.get(f"{base}/v2/thumbs.php", params=params, timeout=90).raise_for_status()

ready_params = {key: value for key, value in params.items() if key != "refresh"}
ready_url = f"{base}/v2/thumbs_ready.php?{urlencode(ready_params)}"
for attempt in range(10):
    ready_response = requests.get(ready_url, timeout=30)
    ready_response.raise_for_status()
    state = ready_response.json()
    if state.get("Error"):
        raise RuntimeError(f"PagePeeker reported a readiness error: {state}")
    if state.get("IsReady"):
        image_response = requests.get(
            f"{base}/v2/thumbs.php", params=ready_params, timeout=90
        )
        image_response.raise_for_status()
        if not image_response.headers.get("Content-Type", "").startswith("image/"):
            raise RuntimeError("Thumbnail endpoint did not return an image")
        Path("screenshot.jpg").write_bytes(image_response.content)
        break
    if attempt < 9:
        sleep(3)
else:
    raise TimeoutError("Thumbnail did not become ready within the polling budget")

This implementation illustrates a finite readiness loop; verify the exact request behavior for your account before deploying. Each readiness check is an API call, so polling more often increases call usage.

5. Understand options, freshness, and usage

Option or behavior What to know
size Part of the documented V2 thumbnail request. Confirm the size values available to your account in PagePeeker’s API documentation.
url The target page. Encode the entire URL as a parameter, especially when it contains a query string.
code Optional in the documented endpoint pattern. Keep any account code private; do not expose it in public client-side code.
refresh=1 Premium option to force regeneration. A daily job without refresh may get a cached thumbnail, depending on cache and account behavior.
wait An optional documented parameter; availability may depend on account tier. Confirm its account-specific behavior with PagePeeker.
Readiness endpoint Checks whether a thumbnail is ready, but each check counts as an API call according to PagePeeker’s FAQ.
Cache PagePeeker says generated thumbnails are cached for several days. A repeated API request is not by itself proof of a newly rendered page.

PagePeeker distinguishes a render—the robot fetching the site and creating a thumbnail—from an API call. Its FAQ says displaying an existing thumbnail, creating an uncached one, checking availability, and any API request each add one API call. A job that refreshes, polls, retrieves, retries, or checks status may therefore use multiple calls per target per day. Build the expected call count from the actual sequence and confirm how your plan treats it.

Its public pricing page has listed Basic at $5.99/month for 100,000 API calls and seven-day caching, Advanced at $39.99/month for 1,000,000 calls and five-day caching, and Premium at custom pricing with customizable caching. Pricing and plan terms can change; confirm current details directly. These API plan listings do not establish that scheduled premium snapshots are included.

6. Troubleshooting daily screenshot jobs

Symptom Likely cause What to do
The screenshot looks old The API may have served a cached thumbnail; PagePeeker says generated images are cached for several days. Check whether your account supports premium refresh=1. Confirm cache behavior and cadence with PagePeeker.
The saved file is not an image The response may be an error, pending result, or placeholder. Check HTTP status and Content-Type; use the documented readiness flow and inspect its Error field.
URL with query parameters fails Characters in the target URL were interpreted as API parameters. URL-encode the complete target URL, using a query-parameter encoder as in the examples.
The schedule runs at the wrong hour The host scheduler’s time zone differs from the intended one, or daylight-saving rules shifted the local time. Confirm the scheduler time zone and choose the expected UTC or local-time policy.
Usage is higher than expected Thumbnail retrievals, readiness checks, retries, and other API calls can all count. Count each request in the workflow, reduce unnecessary polling, and keep retries bounded.
The managed daily schedule is not visible in the account Public materials do not establish that schedule setup is self-serve. Contact PagePeeker and confirm setup, delivery, retention, time zone, and terms for your premium account.
A cron run silently fails Cron has a limited environment, relative paths, or missing credentials. Use absolute paths, provide secrets through the job environment, log output, and alert on nonzero exit status.

7. Performance, reliability, and cost considerations

  • Choose a cadence that matches the change rate. PagePeeker recommends weekly or more often and says not to capture more than once daily. Daily jobs create 30 or 31 scheduled runs per month per URL, before retries or polling.
  • Budget API calls per workflow. A refresh request, readiness checks, image retrieval, and retry calls can each affect usage. Estimate these separately from the number of actual renders.
  • Keep polling bounded. Use a maximum number of readiness checks and a clear failure state, rather than waiting forever. Excessively frequent checks add API calls without guaranteeing a faster render.
  • Make the job observable. Record run time, target URL, HTTP outcome, readiness status, and output path. Alert on repeated failures so missed captures do not go unnoticed.
  • Confirm archive and retention needs. The public API materials do not establish that an external scheduled API caller receives the same archive or retention as the managed premium snapshot service.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its one-call API can return a screenshot or PDF, and its options include caching with a chosen TTL, async jobs with signed webhooks, bulk capture, custom waits, and full-page capture. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed, and the response identifies the page verdict and billing status. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

Frequently asked questions

Can PagePeeker take a screenshot every day?

PagePeeker says its premium site-snapshot service supports automated schedules, and its guidance says no more frequently than once per day. Contact the provider to confirm daily setup for your account.

Does the PagePeeker API support scheduled screenshots?

The public V2 API documentation describes thumbnail requests, premium refresh, and readiness checks, but not a recurring schedule parameter. An external scheduler can call the API, subject to account and cache behavior.

How often should I take website snapshots?

PagePeeker recommends weekly or more frequent snapshots and advises against more than one capture per day. The right cadence depends on how quickly the pages you track change.

Does an API-based job provide the same archive as the managed service?

That equivalence is not established in PagePeeker’s public materials. Confirm retention and archive behavior with PagePeeker before choosing the API workflow.

Sources