How Indian SEO Agencies Can Create Website Thumbnails in Bulk with CaptureKit
Build a repeatable CaptureKit workflow for capturing client and competitor URLs, saving consistent thumbnails, and handling failures in batches.
Indian SEO agencies can create website thumbnails in bulk by validating and deduplicating a URL list, then making one authenticated CaptureKit screenshot request per URL with the same capture settings. Save each image with a stable client and URL identifier, record a manifest, and retry failed requests deliberately. CaptureKit documents a one-credit cost per capture call; its cited capture reference describes individual calls, not a native bulk-job endpoint or guaranteed batch throughput.
This workflow suits agency directories, audit libraries, proposal examples, portfolio pages, and visual monitoring. CaptureKit lists thumbnail generation and SEO tracking among its use cases, but screenshots alone do not establish ranking improvements. See the CaptureKit capture reference and its API introduction for current behavior and account details.
1. Plan the batch
Start with a UTF-8 CSV containing one URL per row and a stable identifier. For example:
client_id,url
acme,https://example.com/
acme,https://example.com/services/
contoso,https://contoso.example/
Before capturing:
- Normalize URLs and reject rows without an absolute HTTP or HTTPS URL.
- Deduplicate exact normalized URLs, while retaining the original URL and client identifier for reporting.
- Choose one output format and viewport profile for pages that will appear side by side.
- Estimate calls: each CaptureKit capture call costs one credit. Check current quota and usage in the account before a large run.
- Choose an output directory or configure optional S3-compatible storage.
Keep the API key in an environment variable or secret manager. CaptureKit accepts the key in the x-api-key header; do not put it in a spreadsheet, browser code, or a public repository.
2. Choose consistent thumbnail settings
| Setting | When to use it | Practical note |
|---|---|---|
format |
PNG, JPEG/JPG, WebP, or PDF | PNG can preserve fine interface detail; JPEG or WebP can reduce image size. This is a format tradeoff, not a CaptureKit benchmark. The documented default is PNG. |
viewport_width, viewport_height |
Consistent desktop framing | Documented defaults are 1280 by 1024 pixels. Set explicit values when comparing pages. |
device |
Mobile or tablet preview | Use one device preset consistently. The reference lists iPhone, iPad, Galaxy, Pixel, Redmi, and Huawei presets; check its current list before use. |
scale_factor |
Higher-resolution output | Default is 1. Larger output can consume more storage and transfer bandwidth. |
image_quality |
JPEG or WebP quality control | The reference documents this for JPEG and WebP and gives a default of 80. |
full_page |
Long-page records rather than compact cards | Defaults to false. A full-page capture may not work well as a small thumbnail. |
full_page_scroll, full_page_scroll_duration |
Lazy-loaded content on full-page captures | Scrolling can help load lazy elements; duration is in milliseconds, with a documented default of 400. |
selector |
Capture one page element | Useful for a hero, product card, or report panel. Verify the selector exists on the target page. |
wait_until, wait_for_selector, delay |
Pages that render after navigation | Supported wait conditions include domcontentloaded, load, networkidle0, and networkidle2. Delay is in seconds and documented from 0 to 10. |
remove_selectors, remove_ads |
Hide known page elements or remove ads | Use CSS selectors for elements you identify. Review samples because site layouts differ. |
block_resources, block_urls |
Reduce loading work or exclude selected requests | Blocking scripts, styles, images, or fonts can change visual output. Test before applying across all clients. |
cache, cache_ttl |
Reuse captures when appropriate | Cache defaults to false. TTL is documented from 3,600 to 2,592,000 seconds. Decide whether a repeated URL should represent a fresh page or a cached result. |
| S3 parameters | Send results to AWS S3 or compatible storage | Optional parameters include bucket, region, object key, access key, secret key, endpoint, and s3_url. Protect storage credentials and separate client data with your bucket policy and naming scheme. |
For a directory grid, start with a viewport capture and a consistent desktop or mobile profile. Capture full pages separately when the purpose is audit documentation. CaptureKit documents device emulation and viewport dimensions; neither guarantees identical rendering across every site.
3. Make a single CaptureKit request
The API endpoint is GET https://api.capturekit.dev/v1/capture. The URL is required. The API key belongs in the x-api-key header. Successful synchronous calls return the capture; the reference documents one credit per call.
cURL
export CAPTUREKIT_API_KEY='YOUR_API_KEY'
curl --fail --silent --show-error \
-H "x-api-key: $CAPTUREKIT_API_KEY" \
--get 'https://api.capturekit.dev/v1/capture' \
--data-urlencode 'url=https://example.com/' \
--data-urlencode 'format=webp' \
--data-urlencode 'viewport_width=1280' \
--data-urlencode 'viewport_height=800' \
--output example.webp
To use another option, add its documented query parameter. URL-encode values rather than concatenating raw URLs into a query string.
Python: process a CSV with bounded concurrency
Install the dependency with python -m pip install requests. Save this as capture_batch.py and provide an input.csv with client_id and url columns. The script uses a small worker pool, a timeout, deterministic filenames, and a manifest. It retries transient network errors and HTTP 429/5xx responses with bounded backoff. Adjust concurrency to your account limits; the cited docs do not prescribe a safe concurrency level.
import csv
import hashlib
import os
import re
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from urllib.parse import urlparse
import requests
API_KEY = os.environ.get("CAPTUREKIT_API_KEY")
if not API_KEY:
raise SystemExit("Set CAPTUREKIT_API_KEY before running this script")
INPUT = Path("input.csv")
OUT = Path("thumbnails")
OUT.mkdir(exist_ok=True)
MANIFEST = OUT / "manifest.csv"
ENDPOINT = "https://api.capturekit.dev/v1/capture"
WORKERS = 3
MAX_ATTEMPTS = 4
def safe_name(value):
value = re.sub(r"[^A-Za-z0-9._-]+", "_", value.strip())
return value[:80] or "unknown"
def valid_url(value):
parsed = urlparse(value)
return parsed.scheme in ("http", "https") and bool(parsed.netloc)
def capture(row):
client_id = row["client_id"].strip()
url = row["url"].strip()
if not valid_url(url):
return {"client_id": client_id, "url": url, "status": "invalid_url", "path": ""}
digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:12]
path = OUT / f"{safe_name(client_id)}_{digest}.webp"
params = {
"url": url,
"format": "webp",
"viewport_width": 1280,
"viewport_height": 800,
}
for attempt in range(MAX_ATTEMPTS):
try:
response = requests.get(
ENDPOINT,
headers={"x-api-key": API_KEY},
params=params,
timeout=(10, 90),
)
if response.status_code == 200:
path.write_bytes(response.content)
return {"client_id": client_id, "url": url, "status": "ok", "path": str(path)}
if response.status_code not in (429, 500, 502, 503, 504):
return {
"client_id": client_id,
"url": url,
"status": f"http_{response.status_code}: {response.text[:300]}",
"path": "",
}
except requests.RequestException as exc:
if attempt == MAX_ATTEMPTS - 1:
return {"client_id": client_id, "url": url, "status": f"network_error: {exc}", "path": ""}
if attempt < MAX_ATTEMPTS - 1:
time.sleep(min(2 ** attempt, 16))
return {"client_id": client_id, "url": url, "status": "retry_exhausted", "path": ""}
with INPUT.open(newline="", encoding="utf-8-sig") as source:
reader = csv.DictReader(source)
if not {"client_id", "url"}.issubset(reader.fieldnames or []):
raise SystemExit("CSV must contain client_id and url columns")
rows = list(reader)
# Deduplicate exact URLs while preserving the first row's client mapping.
unique = []
seen = set()
for row in rows:
url = row.get("url", "").strip()
if url not in seen:
seen.add(url)
unique.append(row)
results = []
with ThreadPoolExecutor(max_workers=WORKERS) as pool:
futures = [pool.submit(capture, row) for row in unique]
for future in as_completed(futures):
result = future.result()
results.append(result)
print(result["status"], result["url"])
with MANIFEST.open("w", newline="", encoding="utf-8") as target:
fields = ["client_id", "url", "status", "path"]
writer = csv.DictWriter(target, fieldnames=fields)
writer.writeheader()
writer.writerows(sorted(results, key=lambda item: (item["client_id"], item["url"])))
print(f"Processed {len(results)} unique URLs; manifest: {MANIFEST}")
The example deliberately keeps concurrency low and records failures instead of silently discarding them. A production job can persist each result as it completes, resume from non-success rows, and write capture settings and timestamps into the manifest. Avoid automatic retries for permanent 4xx errors.
Node.js: single capture request
This example uses Node.js with the built-in fetch API. It writes the binary response to a file and rejects non-success HTTP responses.
import { writeFile } from 'node:fs/promises';
const key = process.env.CAPTUREKIT_API_KEY;
if (!key) throw new Error('Set CAPTUREKIT_API_KEY');
const params = new URLSearchParams({
url: 'https://example.com/',
format: 'webp',
viewport_width: '1280',
viewport_height: '800',
});
const response = await fetch(`https://api.capturekit.dev/v1/capture?${params}`, {
headers: { 'x-api-key': key },
signal: AbortSignal.timeout(90000),
});
if (!response.ok) {
throw new Error(`CaptureKit returned HTTP ${response.status}: ${await response.text()}`);
}
await writeFile('example.webp', Buffer.from(await response.arrayBuffer()));
4. Run and review the batch
- Check account usage and available quota, then run a small sample spanning different client site types.
- Inspect the images for blank captures, consent overlays, bot checks, missing fonts, mobile layout changes, and content that appears late.
- Adjust viewport, wait condition, selector, or delay for specific page groups when needed; do not assume one timing setting fits every site.
- Run the full list with bounded concurrency and write a manifest mapping client, input URL, settings, timestamp, status, and output path.
- Review all failures and retry only eligible transient failures. Confirm the account usage after the batch.
- Publish thumbnails from a controlled output directory or storage bucket, with client access separated according to your agency's data handling requirements.
CaptureKit's documented use cases include thumbnail generation for catalogs, portfolios, directories, and summaries, including a thumbnail_url pattern for directory grids. The cited materials do not establish India-specific latency, regional availability, special rupee pricing, or a rendering guarantee for Indian websites. Verify current account pricing and quota before quoting a client.
5. Handle failures and edge cases
| Symptom or status | Likely cause | What to do |
|---|---|---|
| 400 Bad Request | Missing or malformed URL, unsupported parameter value, or wrong parameter type. | Check the URL and parameter names against the live reference. Encode query values and send booleans/numbers in the documented form. |
| 401 Invalid API Key | Missing, incorrect, revoked, or misplaced key. | Send the active key in the x-api-key header. Rotate it in the dashboard if exposed. |
| 402 Payment required | Account billing or available credits prevent the request. | Check account billing and usage before resuming. Avoid repeatedly sending the same blocked request. |
| 429 Rate limit or quota exceeded | Request rate or key/account quota limit reached. | Slow the worker pool, honor any retry guidance, use exponential backoff, and check key limits and monthly quota. |
| 500 Internal Error or failed capture | Server-side or target-page failure. | Record the URL and response, retry a limited number of times, and leave unresolved rows in the manifest for review. |
| Timeout | Slow target site, heavy assets, stalled network, or wait condition that never settles promptly. | Set a finite client timeout, try a suitable wait condition or selector, and inspect the page manually. Avoid indiscriminately increasing delays for every URL. |
| Image is blank or incomplete | Page blocked automation, content renders late, a required resource was blocked, or target returned an error page. | Review the result and target behavior. Adjust wait settings or resource blocking and validate a sample. The cited docs do not promise universal CAPTCHA bypass or successful rendering of every page. |
| Consent banner, chat widget, or personalized page | Site-specific overlay, locale, cookies, or geolocation changes what the visitor sees. | Use documented removal selectors where appropriate, and check representative pages. Do not treat this as a guaranteed banner-removal or location-fidelity feature. |
| Invalid selector or missing element | Selector differs by page template or the element has not appeared. | Check the selector in the page DOM, wait for it if needed, or use a viewport capture for pages without that component. |
| Output file has the wrong extension or cannot be opened | Requested format and filename extension disagree, or an error response was saved as an image. | Only write the response body after a successful status; keep the extension aligned with format. |
CaptureKit's introduction documents 200 for billed synchronous success; 4xx and 5xx failures are generally free, with endpoint-specific billing exceptions for 404 noted in the introduction. The capture reference lists 1 credit per call. Check the current endpoint docs and request logs when reconciling usage.
6. Performance, reliability, and cost
- Cost planning: budget one credit per capture request according to the capture reference. A list with 500 unique URLs therefore represents 500 capture calls if every URL is requested once; this is a call count, not an account price estimate. Confirm current pricing and quota in the dashboard.
- Concurrency: use bounded workers and backoff. The reviewed references do not specify a safe concurrency level or throughput promise.
- Repeatability: keep format, viewport, device, and timing settings stable for comparisons. Record changes so future captures can be interpreted correctly.
- Cache: enable caching only when serving an existing capture is acceptable. Choose a TTL based on how often clients expect updated pages; cache behavior and billing should be verified in current docs and account logs.
- Payload size: choose an output format and quality suited to the grid. Full-page and high-scale captures can create larger artifacts; measure your own storage and transfer needs rather than assuming a particular size.
- Reliability: persist a manifest, record status and error details, and make the batch resumable. Treat bot checks, CAPTCHAs, consent overlays, language variants, and geographically personalized content as site-by-site validation concerns.
- Storage: optional S3 parameters can send results to AWS S3 or compatible storage. Keep credentials server-side and define bucket permissions and object naming with client separation in mind.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its one-call endpoint can return a screenshot or PDF; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
- Cookie banners and consent overlays, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
- Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are never billed; responses include
X-Page-VerdictandX-Billedheaders. - An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.
Frequently asked questions
Does CaptureKit provide a bulk endpoint?
The reviewed capture reference documents a request for one URL. This workflow processes a list through repeated requests; it does not rely on an undocumented bulk endpoint.
Can screenshots improve a client's search rankings?
Screenshots can serve as visual records for audits and monitoring. The cited materials do not show that creating screenshots by itself improves organic rankings.
Should an agency capture full pages for every directory card?
Usually a consistent viewport image is easier to scan in a grid. Use full-page captures when the deliverable is an audit or visual record of the full page.
Can I use the same settings for every Indian website?
Use a shared baseline for consistency, then validate representative pages. Language, consent behavior, bot checks, content timing, and location-personalized output can vary by site, and the cited documentation does not guarantee identical results.
How should I estimate the batch's CaptureKit cost?
Count the capture calls and use the documented one-credit-per-call unit, then verify the account's current pricing, available credits, and quota before running the job.


