How to Bulk Capture Screenshots from a CSV of URLs with PagePeeker
Use Python to read URLs from a CSV and save PagePeeker thumbnails, with practical guidance on encoding, readiness, usage, retries, and alternatives.
Direct answer: PagePeeker’s documented V2 API accepts one page URL per thumbnail request; its reviewed documentation does not describe CSV import or a native bulk endpoint. To capture a CSV, run a script that reads each row, URL-encodes its URL, calls the V2 endpoint, and saves the response. The example below uses Python and includes validation, bounded concurrency, timeouts, retries, response checks, and a results log.
What this workflow does
This is a client-side loop over PagePeeker’s individual URL endpoint, not a built-in PagePeeker bulk feature. Each valid CSV row produces its own request. The documented V2 pattern is http://{entrypoint}.pagepeeker.com/v2/thumbs.php?size={size}&code={code}&refresh={refresh}&wait={wait}&url={url}. PagePeeker documents free and api entrypoints. The code parameter is optional in the documentation and recommended for server-side calls; do not expose it in public client-side code because others could use your account quota. See the PagePeeker API documentation.
The standard documented sizes are:
| Size | Dimensions | Typical use |
|---|---|---|
t |
90 × 68 | Very small previews |
s |
120 × 90 | Compact lists |
m |
200 × 150 | Small cards |
l |
400 × 300 | Larger previews |
x |
480 × 360 | Largest standard thumbnail |
Premium accounts may offer additional sizes. Choose based on the display dimensions and bandwidth budget. These are thumbnails, not full-page images. PagePeeker describes full-page capture and further output controls as premium features; check its full-page capture information and current account terms if you need content below the fold.
Prepare the CSV
Use a header named url, with one absolute HTTP or HTTPS URL per row. For example, save this as urls.csv:
url
https://example.com/
https://www.python.org/
https://www.wikipedia.org/
The script below reads UTF-8 CSV files, including files saved with a UTF-8 byte-order mark. It rejects missing or malformed URL values and records the row number and original URL in a CSV results file. It creates filenames from row numbers so duplicate URLs do not overwrite one another.
Runnable Python script
Save as capture_csv.py. Install the sole dependency with python -m pip install requests, then run python capture_csv.py urls.csv. Set the PagePeeker entrypoint and optional account code through environment variables so credentials are not committed with the script.
import csv
import os
import re
import sys
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from urllib.parse import urlparse
import requests
API_ENTRYPOINT = os.environ.get("PAGEPEEKER_ENTRYPOINT", "free")
PAGEPEEKER_CODE = os.environ.get("PAGEPEEKER_CODE")
SIZE = os.environ.get("PAGEPEEKER_SIZE", "x")
OUT_DIR = Path("screenshots")
RESULTS_PATH = Path("capture-results.csv")
TIMEOUT_SECONDS = 60
MAX_WORKERS = 2
MAX_ATTEMPTS = 3
VALID_SIZES = {"t", "s", "m", "l", "x"}
def valid_page_url(value):
parsed = urlparse(value)
return parsed.scheme in {"http", "https"} and bool(parsed.netloc)
def capture(row_number, page_url):
params = {"size": SIZE, "url": page_url}
if PAGEPEEKER_CODE:
params["code"] = PAGEPEEKER_CODE
endpoint = f"http://{API_ENTRYPOINT}.pagepeeker.com/v2/thumbs.php"
output_path = OUT_DIR / f"capture-{row_number:05}.img"
last_error = "request failed"
for attempt in range(1, MAX_ATTEMPTS + 1):
try:
response = requests.get(endpoint, params=params, timeout=TIMEOUT_SECONDS)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "").lower()
if not content_type.startswith("image/"):
raise ValueError(f"expected image response, got Content-Type {content_type!r}")
if not response.content:
raise ValueError("empty response body")
# Keep the returned bytes as-is. The endpoint's response headers and
# image format determine the actual format; do not assume PNG.
output_path.write_bytes(response.content)
return row_number, page_url, "saved", str(output_path), ""
except (requests.RequestException, OSError, ValueError) as exc:
last_error = str(exc)
if attempt < MAX_ATTEMPTS:
time.sleep(attempt * 2)
return row_number, page_url, "error", "", last_error
def main():
if len(sys.argv) != 2:
raise SystemExit("Usage: python capture_csv.py urls.csv")
if SIZE not in VALID_SIZES:
raise SystemExit(f"PAGEPEEKER_SIZE must be one of: {', '.join(sorted(VALID_SIZES))}")
if not re.fullmatch(r"[a-zA-Z0-9-]+", API_ENTRYPOINT):
raise SystemExit("PAGEPEEKER_ENTRYPOINT must be a simple entrypoint name")
csv_path = Path(sys.argv[1])
OUT_DIR.mkdir(parents=True, exist_ok=True)
jobs = []
results = []
with csv_path.open(newline="", encoding="utf-8-sig") as source:
reader = csv.DictReader(source)
if not reader.fieldnames or "url" not in reader.fieldnames:
raise SystemExit("CSV must have a header named 'url'")
for row_number, row in enumerate(reader, start=1):
page_url = (row.get("url") or "").strip()
if not valid_page_url(page_url):
results.append((row_number, page_url, "invalid", "", "expected an absolute http(s) URL"))
else:
jobs.append((row_number, page_url))
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
futures = [pool.submit(capture, row_number, page_url) for row_number, page_url in jobs]
for future in as_completed(futures):
results.append(future.result())
results.sort(key=lambda item: item[0])
with RESULTS_PATH.open("w", newline="", encoding="utf-8") as output:
writer = csv.writer(output)
writer.writerow(["row", "url", "status", "file", "error"])
writer.writerows(results)
saved = sum(result[2] == "saved" for result in results)
errors = len(results) - saved
print(f"Saved {saved}; rows needing attention: {errors}; log: {RESULTS_PATH}")
if __name__ == "__main__":
main()
The script uses a small concurrency limit as a conservative starting point; the reviewed PagePeeker API materials do not specify a request-rate limit. For a large run, check current account terms or ask PagePeeker what request pacing is appropriate. If you prefer sequential requests, set MAX_WORKERS = 1. The files use a neutral .img extension because the response format should be determined from the returned content type or image bytes rather than assumed to be PNG. Rename files after confirming the actual format if your downstream workflow requires extensions.
Run the capture
- Put the script and
urls.csvin a working directory. - Choose a documented entrypoint and size. The defaults above use
freeandx; use the entrypoint applicable to your account. - If your account uses a code, set it in the environment. For example, in a POSIX shell:
export PAGEPEEKER_CODE='your-code'. Keep it out of browser JavaScript, public repositories, and shared logs. - Run
python capture_csv.py urls.csv. - Review
screenshots/andcapture-results.csv. Retry only rows markederroror otherwise not ready, after determining the cause.
To avoid rerunning successful rows on a later attempt, preserve the results file and adapt the input selection to include only failures. For repeatable production jobs, store a stable source identifier alongside each URL and include it in the results log; row numbers alone can change when the CSV is reordered.
cURL, Python, and Node.js request examples
These examples make one request each. They use the documented V2 endpoint and encode the page URL as a query parameter. The CSV script above is the complete loop and operational example.
cURL
curl -G "http://free.pagepeeker.com/v2/thumbs.php" \
--data-urlencode "size=x" \
--data-urlencode "url=https://example.com/" \
--fail --output page.img
For an account code, add --data-urlencode "code=$PAGEPEEKER_CODE" in a shell where that environment variable is set. Do not put a real code in a command that will be stored in public logs or shared history.
Python
import requests
response = requests.get(
"http://free.pagepeeker.com/v2/thumbs.php",
params={"size": "x", "url": "https://example.com/"},
timeout=60,
)
response.raise_for_status()
if not response.headers.get("Content-Type", "").lower().startswith("image/"):
raise RuntimeError("The response was not an image; check readiness and account settings")
with open("page.img", "wb") as image_file:
image_file.write(response.content)
Node.js
const endpoint = new URL("http://free.pagepeeker.com/v2/thumbs.php");
endpoint.searchParams.set("size", "x");
endpoint.searchParams.set("url", "https://example.com/");
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 60_000);
try {
const response = await fetch(endpoint, { signal: controller.signal });
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const contentType = response.headers.get("content-type") ?? "";
if (!contentType.toLowerCase().startsWith("image/")) {
throw new Error(`Expected image response; received ${contentType || "no Content-Type"}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
if (bytes.length === 0) throw new Error("Empty image response");
const { writeFile } = await import("node:fs/promises");
await writeFile("page.img", bytes);
} finally {
clearTimeout(timer);
}
Readiness, refresh, and response handling
A successful HTTP response does not by itself mean that the desired final thumbnail is ready. PagePeeker documents a readiness API and a premium wait option. Consult the current API documentation for the exact readiness endpoint and supported parameters for your account; do not guess an endpoint or treat a placeholder as the finished capture. If a capture is not ready, check readiness or use the documented wait option where available, then retry according to your account’s guidance.
The API documentation describes response headers for paid and unbranded accounts, including capture method, capture duration, final redirected URL, capture timestamp, and error status. Log relevant headers with the source URL when diagnosing failures. Respect redirects: a site can land on a different final URL, which can explain why the resulting image differs from what the CSV value suggests.
The V2 pattern also lists refresh and wait. Their exact supported values and plan availability should be taken from the current documentation. Avoid setting refresh indiscriminately: it may cause a new capture when a cached thumbnail would suffice. A readiness check is also an API call for usage purposes.
Usage, cost, and planning
Estimate calls before processing a large file. PagePeeker’s FAQ says displaying a cached thumbnail, generating an uncached thumbnail, checking readiness, and any exposed API call each add one API call. Unused monthly quota does not roll over. A readiness workflow can therefore use more than one call per input URL. See the PagePeeker FAQ.
The PagePeeker pricing page reviewed for this article listed Basic at $5.99/month for 100,000 API calls and Advanced at $39.99/month for 1,000,000 API calls. Prices, quotas, account features, and API behavior can change; verify the current pricing page before committing to a large run. The reviewed docs do not state a maximum batch size or API request-rate limit, so do not assume unlimited throughput.
A rough planning model is input rows + readiness checks + intentional refreshes + any other API calls. Invalid rows should be filtered locally and do not need a request. Start with a small sample, inspect the result log and account usage, then process the remaining rows at a measured pace. This makes it easier to spot placeholder responses, bad inputs, or unexpectedly high call consumption early.
Performance and reliability
- Bound concurrency: parallel requests reduce wall-clock time but can increase load and may encounter account or service limits. Since the reviewed docs do not publish a rate limit, use low concurrency and confirm suitable pacing for high-volume work.
- Use timeouts and bounded retries: the sample has a 60-second timeout and three attempts with increasing delays. Retry transient network errors and server failures selectively; fix malformed URLs and authorization or account issues rather than retrying them repeatedly.
- Make retries resumable: retain per-row status and source URL. Keep successful outputs, and retry failed or not-ready rows instead of recapturing the entire file.
- Check the payload: validate HTTP status, content type, nonempty bytes, and, in production, decode the image with an image library. An image-like response can still be a placeholder or error image, so use documented readiness and error signals too.
- Choose the needed dimensions: larger thumbnails consume more storage and transfer bytes. Standard sizes top out at 480 × 360; use the smallest size that meets the display need.
- Plan for changing pages: thumbnails may reflect an earlier capture or cache state. Use documented refresh behavior when freshness matters and account for the extra calls.
Common errors and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| CSV header error | The file has no url column, or the heading differs in capitalization or spacing. |
Use a header exactly named url, or change the script to match the actual column name. |
| Rows marked invalid | Blank values, relative links, missing scheme, or malformed URLs. | Supply absolute http:// or https:// URLs and remove blank rows. |
| HTTP error or timeout | Network interruption, slow destination or capture, or an account/endpoint issue. | Check the URL and endpoint, lower concurrency, inspect the response, and retry transient failures with a bounded delay. |
| Response is not an image | The capture may not be ready, or the service returned an error response. | Check status and response headers; use PagePeeker’s documented readiness API or premium wait option as applicable. |
| Image is a placeholder or stale | The thumbnail may still be processing or may be cached. | Check readiness and error headers. Use the documented refresh behavior only when a fresh capture is required. |
| Account quota is consumed faster than expected | Readiness checks and other exposed API calls count, in addition to thumbnail usage. | Include these calls in the estimate, avoid unnecessary readiness polling and refreshes, and check the account dashboard and current plan details. |
| Code appears in a public page or repository | The account code was embedded in client-side or shared code. | Move it to a server-side environment variable and follow PagePeeker’s account guidance if it may have been exposed. |
| Capture omits below-the-fold content | The request produced a standard thumbnail rather than a full-page capture. | Review PagePeeker’s premium full-page options and verify the required dimensions and output settings before scaling the job. |
Site owner opt-out
If you operate a site and want to block PagePeeker’s robot, PagePeeker documents this robots.txt rule:
User-agent: PagePeeker
Disallow: /
PagePeeker says its robot attempts to fetch a site or page when a client requests it and that it tries to fetch a page only once every 5–7 days. This is PagePeeker’s stated behavior, not a general rule for screenshot services. See its robot information.
Or skip the browser setup
If you want a hosted API that processes a URL with one GET request, ScreenshotNeo returns a PNG, JPEG, WebP, or PDF. Its API accepts a URL directly, and its parameter names also work with those used by other screenshot APIs. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Does PagePeeker accept a CSV directly?
The reviewed V2 API documentation describes individual URL thumbnail requests, not a CSV upload or native bulk endpoint. Use a script to loop through the rows.
Can I get full-page screenshots from the standard thumbnail endpoint?
The standard V2 sizes describe thumbnails. PagePeeker presents full-page capture as a premium option; confirm the account-specific settings and output limits in its current documentation.
Can the script run in a browser?
A local or server-side script is preferable when using an account code. PagePeeker warns against exposing the code on public client-side pages because it can be used against the account quota.
How often should I check whether a capture is ready?
Follow the documented readiness API or premium wait behavior for the account. Each readiness check affects API usage, so avoid aggressive polling.
Can I use this workflow for thousands of URLs?
The same per-row pattern can process a large file, but the reviewed docs do not publish a native batch size or request-rate limit. Estimate API calls, process a sample, use bounded concurrency, and confirm suitable limits with PagePeeker for high-volume jobs.


