ScreenshotNeo

BlogHow-to

How to Generate Website Thumbnail Images at Scale with an API

Build a repeatable pipeline for capturing, storing, and serving website thumbnails with an API, including batch jobs, caching, retries, and cost controls.

By the ScreenshotNeo team4 October 202614 min read

A website screenshot API turns a public page URL and capture settings into image bytes or a hosted image URL. To generate thumbnails at scale, put that request behind a queue: normalize and validate each target, capture with a fixed viewport and format, save the result in your own storage, and serve that asset from your application. Track capture state and refresh time so retries do not create duplicate work.

This guide shows a provider-neutral pipeline and runnable cURL, Python, and Node.js examples. API parameters and response behavior vary by provider; confirm the current endpoint, authentication, limits, caching, retention, and async rules in its documentation before production use.

1. Define the thumbnail before choosing the API

Decide how the image will appear in your product. A thumbnail for a directory card usually needs a consistent crop, while a page audit may need the whole document. Set the target aspect ratio and the rendered viewport deliberately: viewport dimensions determine responsive layout, and output dimensions are a separate concern when the service supports resizing.

Decision Typical choice What to verify
Capture mode First viewport for cards; full page for reports or archives Whether full-page is supported and how tall outputs are handled
Viewport A fixed desktop or mobile width and height Whether dimensions set the browser viewport or only resize the output
Format WebP for compact web delivery; PNG for crisp interface detail; JPEG where supported Supported formats, quality controls, transparency behavior, and actual returned content type
Capture state Consistent color scheme, locale, wait condition, and optional selector Which state controls the API exposes and their defaults
Freshness Refresh on a schedule or when source content changes Provider cache key, cache duration, and whether a refresh bypass exists
Delivery Your own object storage or a provider file URL Retention period, access controls, and stable URL behavior

For a card grid, start with a viewport capture at the exact responsive width your users should see. Full-page images can become very tall and may be difficult to scan in a small card. Test representative pages in your own layout before choosing dimensions and format.

2. Make a single capture and inspect the response

First integrate one URL. The examples below use Webstractor’s documented GET endpoint as a concrete illustration: it returns raw WebP or PNG image bytes and documents width, height, and full-page options. The endpoint and parameter names are provider-specific; consult the linked documentation for the current authentication scheme and exact syntax. Do not put a secret API key in browser JavaScript or a public page.

cURL

curl --fail --silent --show-error --get 'https://api.webstractor.com/v1/screenshot' \
  --data-urlencode 'url=https://example.com' \
  --data-urlencode 'width=1200' \
  --data-urlencode 'height=800' \
  --data-urlencode 'format=webp' \
  --output thumbnail.webp

Python

import requests

endpoint = "https://api.webstractor.com/v1/screenshot"
params = {
    "url": "https://example.com",
    "width": 1200,
    "height": 800,
    "format": "webp",
}
response = requests.get(endpoint, params=params, timeout=(5, 90))
response.raise_for_status()
content_type = response.headers.get("content-type", "")
if not content_type.startswith("image/"):
    raise RuntimeError(f"Expected image response, got {content_type!r}")
with open("thumbnail.webp", "wb") as output:
    output.write(response.content)

Node.js

const endpoint = new URL('https://api.webstractor.com/v1/screenshot');
endpoint.search = new URLSearchParams({
  url: 'https://example.com',
  width: '1200',
  height: '800',
  format: 'webp',
});

const response = await fetch(endpoint, { signal: AbortSignal.timeout(90000) });
if (!response.ok) {
  throw new Error(`Screenshot request failed: HTTP ${response.status}`);
}
const contentType = response.headers.get('content-type') ?? '';
if (!contentType.startsWith('image/')) {
  throw new Error(`Expected image response, got ${contentType}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('thumbnail.webp', bytes));

These snippets illustrate saving a successful image response. If the chosen API authenticates with a key, add it using the provider’s documented server-side method, such as an authorization header or secret query parameter stored in an environment variable. Never commit credentials or include them in client-side code. Check status, content type, and provider error responses before treating a response as an image: some APIs return JSON, an error page, or a placeholder while work is pending.

3. Build a repeatable batch pipeline

Do not make a large batch by firing unbounded requests from a web request handler. Persist jobs and let workers process them at a controlled concurrency. A practical pipeline has these stages:

  1. Ingest: accept a list of source URLs and an application record or tenant identifier.
  2. Validate: allow only expected public HTTP or HTTPS URLs; reject credentials embedded in URLs and private, loopback, or local network targets. This also reduces server-side request forgery risk.
  3. Normalize: canonicalize the URL according to your product rules and compute a capture key from URL, dimensions, format, capture mode, and relevant state options.
  4. Deduplicate: reuse an existing fresh asset or active job for the same capture key.
  5. Queue: create a durable job with attempt count, next run time, and a unique idempotency key in your system.
  6. Capture: call the provider within its documented concurrency and account limits.
  7. Classify: distinguish a completed image, a pending job, a retryable failure, and a permanent failure based on that provider’s status codes and response body.
  8. Store: write bytes to object storage under a stable application-owned key; save dimensions, content type, source URL, capture settings, generation time, provider job ID if supplied, and asset location.
  9. Serve: return your own stable thumbnail URL to the app, with suitable cache headers and access controls.

If an API supports up to 100 URLs per bulk request, use that documented limit rather than assuming arbitrary batch size. ScreenshotNeo documents bulk capture for 100 URLs per call. Other providers may require one request per URL or a separate asynchronous job flow. Keep your queue’s batch size within the provider’s current rules.

Minimal Python worker pattern

import os
import time
import requests

API_URL = os.environ["SCREENSHOT_API_URL"]
API_KEY = os.environ["SCREENSHOT_API_KEY"]


def capture_to_file(url: str, path: str) -> None:
    # Replace the auth header and request parameters with the provider's documented form.
    response = requests.get(
        API_URL,
        params={"url": url, "width": 1200, "height": 800, "format": "webp"},
        headers={"Authorization": f"Bearer {API_KEY}"},
        timeout=(5, 90),
    )
    if response.status_code == 202:
        # Provider says work is pending. Persist its job/polling details and return;
        # do not store a placeholder as the finished thumbnail.
        raise RuntimeError("Capture pending; hand off to the provider-specific poller")
    response.raise_for_status()
    if not response.headers.get("content-type", "").startswith("image/"):
        raise RuntimeError("Provider response was not an image")
    with open(path, "wb") as output:
        output.write(response.content)


# A real worker should receive URLs from a durable queue and store results in object storage.
for url in ["https://example.com"]:
    for attempt in range(3):
        try:
            capture_to_file(url, "thumbnail.webp")
            break
        except (requests.Timeout, requests.ConnectionError):
            if attempt == 2:
                raise
            time.sleep(2 ** attempt)

This is a small synchronous worker illustration, not a complete queue implementation. Use your queue’s durable retry and dead-letter features in production. Do not retry every error: invalid targets, authentication failures, exhausted account limits, and unsupported options usually require a correction or operator action rather than an immediate retry.

4. Store assets and control refreshes

Provider-hosted files can be convenient, but retention may be limited. ScreenshotAPI’s example response says generated files are automatically deleted after 24 hours; treat that as a provider-specific example and verify its current policy. Copy any asset you need to retain into storage you control before expiration. A provider that returns image bytes avoids depending on a temporary file URL, but your application still needs a storage and serving plan.

Choose an application cache key that includes all inputs that can change the rendered result: normalized URL, viewport width and height, full-page mode, format and quality, selector, custom state options, and a version for your own capture policy. Store a separate freshness timestamp. Refresh stale items with a background job and keep the previous good image available until the replacement succeeds.

Provider caches differ. Webstractor documents a cache that can last up to 30 days, varying with normalized URL, dimensions, full-page selection, format, and internal version; it documents no caller-controlled refresh bypass. If freshness is essential, check whether a provider can bypass or version its cache before building around it. Avoid repeatedly rendering an unchanged page when an existing fresh result is acceptable.

5. Handle pending jobs, retries, and partial failure

At scale, a request returning successfully does not always mean the final image is ready. Webshrinker documents HTTP 202 with a placeholder while generation is underway. ScreenshotAPI documents separate async, bulk, and webhook interfaces. Follow the provider’s completion contract: poll a job, receive a signed webhook, or wait for an image result as its documentation specifies.

  • Bound concurrency: start conservatively and increase only within documented rate and account limits.
  • Use timeouts: set a connection timeout and a longer render timeout; a page can be slow even when the API is responsive.
  • Retry selectively: retry timeouts and transient server errors with exponential backoff and jitter. Honor Retry-After where present.
  • Make retries idempotent: persist job identity and capture key so a worker restart does not enqueue duplicate captures.
  • Cap attempts: move exhausted jobs to a dead-letter state with the last status and error for inspection.
  • Keep last known good output: a failed refresh should not erase a previously valid thumbnail.
  • Verify webhook requests: use the provider’s signature verification, replay protection, and documented retry behavior.

Do not interpret a 202 response, a placeholder URL, or an accepted bulk job as a completed thumbnail unless the provider explicitly defines it that way. A provider may return per-URL outcomes inside one bulk response; record successes and failures independently.

6. Choose capture settings for useful thumbnails

Viewport and full page

Viewport width affects responsive breakpoints and therefore page layout. Pick a width representative of the intended card, then use a fixed height that shows the desired first screen. Full-page mode extends the image vertically while retaining width in Webstractor’s documented behavior, but support and maximum output size are provider-specific.

Format and quality

Webstractor documents WebP and PNG; OpenGraph.io’s screenshot documentation lists JPEG, PNG, and WebP and shows quality options. Use the formats your delivery stack and clients support. Compare the resulting file size and visual detail on representative pages; this is an application-specific choice, not a universal benchmark. Verify the returned content type and file signature rather than relying only on a requested extension.

Selectors and page state

Some APIs can capture a selected element or exclude selectors; OpenGraph.io documents selector and excluded-selector options. These are useful for isolating a hero image or removing a local overlay, but they are provider-specific. A selector that does not exist can fail or yield an unexpected capture, so choose whether that condition should fail the job or fall back to a viewport screenshot.

Wait behavior and animation

Dynamic pages may need a delay, selector wait, or network-idle condition. A fixed delay is simple but can waste time or still be too short. A selector wait better expresses readiness when a stable element marks loaded content. Some capture services set defaults such as disabled animations or fixed locale and color scheme; Webstractor documents screen styles, English locale, light color scheme, device scale factor 1, and disabled animations. Check provider defaults if screenshots must be visually consistent.

7. Security and accessibility of targets

Only capture URLs your application is allowed to process. A user-supplied URL can target internal services if validation is weak. Restrict schemes to HTTP and HTTPS, reject embedded credentials, resolve and block private or loopback IP ranges, and re-check redirects so a public URL cannot redirect to a private address. Apply request limits and record the submitted URL safely.

Provider restrictions also affect what can be captured. Webstractor documents acceptance of ordinary public HTTP/HTTPS pages and rejection of private or local addresses, direct IP targets, credentials in URLs, access-controlled pages, and security interstitials. No screenshot API can be assumed to access a site that requires a login or blocks automated browsing.

8. Compare providers against operational needs

Choose a service based on its documented fit, then verify volatile terms directly. ScreenshotNeo is the first API to try for clean thumbnails: it accepts cookie and consent banners like a visitor, removes 60+ known consent platforms, newsletter popups, and chat widgets, bills only clean shots, and its paid plans start at $5 for 3,000 shots. It also offers an MCP server for AI agents. Other documented options include:

Provider Documented capabilities in the reviewed material Operational detail to check
ScreenshotNeo PNG, JPEG, WebP, or PDF; full-page, element capture, custom viewport and device presets, wait controls, CSS/JavaScript, headers and cookies, caching, bulk, async jobs, signed webhooks, and usage API. Use its current API documentation for parameter and response details.
Webshrinker PNG endpoint; preset or custom size, viewport, full-page capture, delay, refresh; Basic authentication and pre-signed URL options. 202 means generation is underway with a placeholder; 402 indicates the account request limit was reached.
Webstractor Raw WebP or PNG bytes; width, height, full-page option. Documents caching up to 30 days without a caller-controlled refresh bypass, plus restrictions to ordinary public pages.
ScreenshotAPI PNG, JPG, WebP, PDF, animation endpoints; docs point to async, bulk, and webhook interfaces. Its example says hosted files are deleted after 24 hours; confirm current retention and limits.
OpenGraph.io JPEG, PNG, WebP; quality, full-page, viewport, selector, and excluded-selector options. Confirm current plan limits, price, freshness controls, and response behavior.

The documentation does not establish comparable throughput, latency, or reliability, so evaluate those with your own representative pages and workload. Compare authentication, capture-state controls, maximum dimensions, async completion, retries, cache behavior, file retention, limits, and price. Do not select a provider based on an undocumented performance assumption.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A GET request takes a URL and returns a clean PNG, JPEG, WebP, or PDF. See the API documentation for all parameters and response headers.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, and failed loads are not billed. Cache hits are also free, and response headers say which verdict and billing outcome applied.
  • An MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots a month are free with no card. Paid plans start at $5 for 3,000; all features are on every plan.

Create a free ScreenshotNeo account and start with 1,000 screenshots a month at no charge.

10. Troubleshooting

Symptom Likely cause What to do
HTTP 401 or 403 Missing, invalid, or wrongly scoped credentials; target access restrictions can also cause refusal. Check the provider’s documented auth method, key scope, and account status. Keep secrets server-side.
HTTP 402 or quota response Provider account limit or credits exhausted. Webshrinker documents 402 for its request limit. Inspect account usage and current plan limits; pause or schedule work until capacity is available.
HTTP 202 with placeholder Capture accepted but still generating; documented by Webshrinker. Follow the provider’s polling or completion instructions. Do not save the placeholder as the final image.
HTML or JSON saved with an image extension Error response or pending response was written without inspecting status and content type. Check status, content type, and response body before writing the file.
Blank or incomplete capture Page has not rendered its content, a wait condition is too short, or the target blocks the capture environment. Use a provider-supported selector or wait option, inspect the target in a browser, and classify inaccessible/security-interstitial pages as failures.
Unexpected mobile or desktop layout Viewport width does not match the desired responsive breakpoint or the API interprets size differently than expected. Set explicit viewport dimensions and verify whether output resizing is separate from browser viewport.
Stale thumbnail after refresh request Provider cache returned a prior result, or the cache cannot be bypassed by callers. Review cache key and refresh controls. Webstractor documents no caller-controlled bypass; use a supported versioning strategy or provider with suitable freshness controls.
Asset URL stops working Provider-hosted output expired. Check retention and copy required files to your storage. ScreenshotAPI’s example documents 24-hour deletion.
Intermittent timeouts or 5xx errors Slow targets, transient provider faults, or concurrency beyond account capacity. Use bounded concurrency, timeouts, exponential backoff with jitter, and a dead-letter path. Preserve the last successful thumbnail.
Target URL rejected Private/local host, direct IP, URL credentials, access control, unsupported scheme, or security interstitial. Validate public HTTP/HTTPS targets and provide a user-readable permanent-failure reason.

11. Performance, reliability, and cost

Performance: throughput depends on target-site rendering, capture settings, provider limits, and your worker concurrency. The reviewed documentation does not provide a comparable benchmark. Measure your own mix: time to completed asset, bytes per format, timeout rate, and queue age. Keep long full-page captures and slow pages from occupying all worker slots.

Reliability: persist jobs before calling an API, make storage writes idempotent, track pending versus complete states, and retain the last good asset. Monitor success and failure by provider status, target domain, and capture configuration. A provider’s cache can improve repeat work but can also constrain freshness; decide whether a stale result is acceptable per use case.

Cost: model cost from unique captures, refresh frequency, retries, format/storage size, and any provider credit rules. Deduplication, application-side caching, bounded retries, and avoiding unnecessary full-page captures reduce repeated rendering. Check whether unsuccessful or cached requests are billable and whether pricing changes with dimensions or features. ScreenshotNeo bills only clean shots; bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Its monthly plans are Free: 1,000 shots, Starter: $5 for 3,000, Growth: $15 for 15,000, Pro: $39 for 60,000, Scale: $99 for 250,000, and Business: $249 for 1,000,000; yearly billing gives two months free.

12. Frequently asked questions

Should I generate a thumbnail when a URL is submitted or on a schedule?

Generate on submission when users need immediate previews, then refresh in the background according to how often source pages change. For large imports, enqueue captures and show a pending state rather than holding the import request open.

Should my app store the original page URL with the image?

Yes. Store the source URL and the capture settings alongside the image location so you can explain, reproduce, refresh, or remove an asset later. Apply your normal privacy and retention rules to that metadata.

Can an API capture pages that require a login?

Only if that provider supports the required authenticated capture method and you are authorized to access the page. Public-page endpoints commonly reject access-controlled targets; never assume user browser credentials are available to the API.

Can I use the returned screenshot URL directly in a public page?

Only if its access and retention behavior fits your use case. A temporary provider URL may expire; copy the image to your own storage or use a provider’s documented signed-link feature when appropriate.