ScreenshotNeo

BlogHow-to

How to Run a Bulk Website Screenshot Job on an Indian Cloud VPS

Build a controlled Playwright worker for capturing many URLs on an Indian cloud VPS, with bounded retries, practical storage choices, and a one-call API alternative.

By the ScreenshotNeo team4 October 202612 min read

To run a bulk website screenshot job on an Indian cloud VPS, provision a VM in a suitable India region, install Python and Playwright with its browser dependencies, read URLs from a file, and capture them with a bounded number of workers. Save each result under a stable filename, log successes and failures, and retry only transient errors. Playwright provides the browser screenshot primitive; the input format, queue, concurrency, retries, and storage are choices your job must implement.

This guide uses Python and Playwright. It captures public pages; authentication, consent prompts, anti-bot checks, dynamic content, and very long pages can change what a screenshot contains or whether navigation succeeds. A cloud location changes the network path and may affect site presentation, but it does not bypass a website’s access rules.

1. Choose and verify an Indian cloud region

AWS lists Mumbai (ap-south-1) with opt-in not required and Hyderabad (ap-south-2) with opt-in required. AWS advises considering service and feature availability as well as proximity to the majority of users. Before creating the VM, confirm that the required instance family, storage, networking, and any other dependencies are available in the selected region. See the AWS Regions and Availability Zones table.

Azure lists Central India, South India, West India, and India South Central among its regions available or coming soon, and directs customers to check regional service availability. Do not assume that a specific VM family or dependent service is offered in every listed region. Check the current Azure geography and service availability information before following a provider-specific setup guide.

For either provider, choose a region based on the services you need and the capture location you want. The sources do not establish a recommended VPS size or guarantee that a given service is available in a region. Start with a small workload, observe memory and CPU use, then size the VM for your actual pages and concurrency.

2. Define the batch and its output

Use one URL per line in a plain text file. This keeps the example focused; a CSV reader or sitemap parser can replace the input function if your source data requires it. Decide how to handle duplicates before the run. This script removes exact duplicate URL strings while preserving their first-seen order.

Each URL gets a deterministic filename derived from its position and a short hash. The index avoids collisions when two distinct URLs have similar names, and the hash helps identify the URL behind a file. The script also writes a JSON Lines manifest with each URL, output path, attempt count, and outcome.

https://example.com/
https://example.com/pricing
https://www.example.org/docs

Create a virtual environment and install Playwright:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip playwright
python -m playwright install chromium
# On Linux, install browser operating-system dependencies if needed:
python -m playwright install-deps chromium

Playwright’s official Python screenshot documentation shows saving a page screenshot to a file, requesting the full page with full_page=True, capturing to a buffer, and taking locator or element screenshots. The screenshot API also supports format and quality parameters.

3. Run a bounded Playwright screenshot worker

Save the following as bulk_screenshots.py. It launches one Chromium browser, then uses a fixed number of asynchronous workers. Each worker creates its own browser context and page per URL, so pages do not share cookies or storage state. The default waits for the page load event, applies a per-navigation timeout, captures the full page, and retries a small number of times with backoff.

import asyncio
import hashlib
import json
import sys
from pathlib import Path
from urllib.parse import urlparse

from playwright.async_api import async_playwright

INPUT = Path(sys.argv[1] if len(sys.argv) > 1 else "urls.txt")
OUTPUT = Path(sys.argv[2] if len(sys.argv) > 2 else "screenshots")
WORKERS = 3
NAVIGATION_TIMEOUT_MS = 30_000
MAX_ATTEMPTS = 3


def read_urls(path: Path) -> list[str]:
    # Ignore blank lines and comment lines; preserve order and remove exact duplicates.
    seen = set()
    urls = []
    for raw in path.read_text(encoding="utf-8").splitlines():
        url = raw.strip()
        if not url or url.startswith("#") or url in seen:
            continue
        parsed = urlparse(url)
        if parsed.scheme not in {"http", "https"} or not parsed.netloc:
            print(f"Skipping invalid URL: {url}", file=sys.stderr)
            continue
        seen.add(url)
        urls.append(url)
    return urls


def output_path(index: int, url: str) -> Path:
    digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:12]
    return OUTPUT / f"{index:06d}-{digest}.png"


async def capture(browser, index: int, url: str, manifest) -> dict:
    destination = output_path(index, url)
    last_error = None
    for attempt in range(1, MAX_ATTEMPTS + 1):
        try:
            context = await browser.new_context(viewport={"width": 1440, "height": 1000})
            try:
                page = await context.new_page()
                response = await page.goto(
                    url,
                    wait_until="load",
                    timeout=NAVIGATION_TIMEOUT_MS,
                )
                # HTTP error statuses can still render a useful page. Record the status;
                # decide separately whether your workflow should keep such screenshots.
                status = response.status if response else None
                await page.screenshot(path=str(destination), full_page=True)
            finally:
                await context.close()
            result = {"url": url, "file": str(destination), "status": status,
                      "attempts": attempt, "outcome": "ok"}
            manifest.write(json.dumps(result, ensure_ascii=False) + "\n")
            manifest.flush()
            print(f"OK {url} -> {destination} (HTTP {status})")
            return result
        except Exception as exc:
            last_error = f"{type(exc).__name__}: {exc}"
            if attempt < MAX_ATTEMPTS:
                await asyncio.sleep(2 ** (attempt - 1))
    result = {"url": url, "file": None, "attempts": MAX_ATTEMPTS,
              "outcome": "error", "error": last_error}
    manifest.write(json.dumps(result, ensure_ascii=False) + "\n")
    manifest.flush()
    print(f"ERROR {url}: {last_error}", file=sys.stderr)
    return result


async def main() -> None:
    if not INPUT.is_file():
        raise SystemExit(f"Input file not found: {INPUT}")
    OUTPUT.mkdir(parents=True, exist_ok=True)
    urls = read_urls(INPUT)
    if not urls:
        raise SystemExit("No valid HTTP or HTTPS URLs found")

    queue = asyncio.Queue()
    for index, url in enumerate(urls, start=1):
        queue.put_nowait((index, url))

    results = []
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch(headless=True)
        async with open(OUTPUT / "manifest.jsonl", "a", encoding="utf-8") as manifest:
            async def worker():
                while True:
                    try:
                        index, url = queue.get_nowait()
                    except asyncio.QueueEmpty:
                        return
                    try:
                        result = await capture(browser, index, url, manifest)
                        results.append(result)
                    finally:
                        queue.task_done()

            await asyncio.gather(*(worker() for _ in range(WORKERS)))
        await browser.close()

    succeeded = sum(item["outcome"] == "ok" for item in results)
    failed = len(results) - succeeded
    print(f"Finished: {succeeded} succeeded, {failed} failed, {len(urls)} unique URLs")
    if failed:
        raise SystemExit(1)


if __name__ == "__main__":
    asyncio.run(main())

Run it with an input file and an output directory:

python bulk_screenshots.py urls.txt run-2026-10-04

The code appends to manifest.jsonl if that file already exists. Use a new output directory for each run to keep artifacts and records isolated. If you want to resume jobs, add a manifest reader that skips URLs with a successful prior result; do not infer success just because a screenshot file exists.

4. Tune capture behavior for your pages

Need Implementation choice Trade-off
Capture the visible viewport only Use full_page=False or omit the argument. Smaller, faster images; content below the fold is absent.
Capture the full scrollable page Keep full_page=True. Very long pages can consume more time and memory; lazy content may need scrolling or additional waits.
Wait for a client-rendered section After navigation, call await page.locator("SELECTOR").wait_for(state="visible", timeout=...). A missing selector fails the job; choose a selector that represents the content you need.
Wait for a fixed delay Call await page.wait_for_timeout(milliseconds). Simple but wastes time on fast pages and may still be too short on slow pages.
Wait for network quiet Use wait_until="networkidle" in navigation where appropriate. Pages with analytics, polling, or persistent connections may not become idle. It is not a universal readiness signal.
Capture one element Wait for a locator and call await page.locator("SELECTOR").screenshot(path="element.png"). The element must exist and be visible; its screenshot does not represent the whole page.
Change image format Use screenshot options such as type="jpeg" and a supported quality value, or choose a file extension consistent with the output. JPEG is lossy; PNG is larger but preserves sharp edges. Check the installed Playwright API for supported options.

To add a selector wait to the example, place this before the screenshot call:

await page.locator("main").wait_for(state="visible", timeout=10_000)

To capture only an element:

await page.locator("main article").screenshot(path=str(destination))

Do not add arbitrary sleeps to every URL without measuring their effect. Prefer a readiness condition tied to the page content, and make the condition configurable when sites differ.

5. Handle retries, duplicates, and partial runs

  • Retry transient failures only. Navigation timeouts and intermittent network errors may succeed on another attempt. An authentication wall, invalid URL, or persistent anti-bot block usually needs an explicit policy rather than repeated retries.
  • Bound retries. The sample makes at most three attempts and backs off between them. For a large job, cap total run time and record the final error so one difficult page cannot stall the batch.
  • Keep per-job output isolated. Give every invocation its own directory. This avoids confusing old captures with current results and makes cleanup straightforward.
  • Persist the manifest as the job runs. JSON Lines allows completed outcomes to survive process interruption. For a resumable job, treat the manifest as the source of truth and make the skip/retry rule explicit.
  • Control concurrency. Start low and increase only after checking VM memory, CPU, network use, and target-site behavior. More pages at once can increase resource pressure and may trigger site defenses.
  • Store artifacts durably. Local VM disk is useful for a simple job, but can be lost when the instance or disk is removed. Copy completed output to persistent disk or object storage according to your retention needs.

These are application design recommendations. The Playwright screenshot documentation describes capture options; it does not prescribe a bulk queue, retry strategy, concurrency limit, or VPS size.

6. Operate the job on a VPS

Run the process inside a persistent session manager or an appropriate service supervisor so an SSH disconnect does not terminate your batch. Redirect standard output and error to a job log if your environment does not collect them. Keep the input, output directory, manifest, and log associated with a run identifier.

For a small one-off job, local disk plus a copied archive may be enough. For recurring or larger jobs, plan disk capacity around the number and dimensions of images, retain only what you need, and move completed artifacts to storage designed for persistence. The sources reviewed do not provide comparable provider prices, throughput benchmarks, or service-level guarantees, so estimate cost using the VM, disk, network transfer, and storage prices shown for your chosen region and actual workload.

7. Troubleshooting

Symptom Likely cause What to do
Executable doesn't exist or browser launch fails Chromium was not installed for the active Playwright environment, or system libraries are missing. Activate the same virtual environment used for installation, run python -m playwright install chromium, then install Linux dependencies with python -m playwright install-deps chromium where supported.
Navigation timeout The site is slow, blocked, waiting on resources, or the timeout is too short. Inspect the URL and error; raise the timeout selectively or use a more appropriate readiness condition. Do not retry indefinitely.
Screenshot is blank or incomplete Content may render after load, require scrolling, depend on scripts, or be blocked. Wait for a meaningful selector, test a bounded delay, inspect the page response and logs, and check whether the site requires permitted authentication or interaction.
Some URLs fail while others succeed Per-site network behavior, invalid input, access controls, or temporary server errors. Use the manifest to isolate failures, classify transient versus persistent errors, and rerun only eligible URLs.
VM runs out of memory or becomes very slow Too many simultaneous browser pages, unusually large pages, or full-page captures. Reduce WORKERS, close contexts reliably, and consider a larger VM only after observing resource use.
Files overwrite or cannot be identified Names are based only on a URL slug or output is reused across runs. Use an index plus URL hash as in the example and write each run to a fresh directory.
Hyderabad region is unavailable in the account AWS marks ap-south-2 as opt-in required. Enable the region for the account if appropriate, or choose an available region after checking the required services.
Cloud console does not offer the needed VM or service Regional product availability differs. Check the provider’s current region/service matrix and select a region where all required components are available.

8. Managed alternatives

ScreenshotNeo is a website screenshot API and MCP server for developers. It is the first alternative to consider when you want to avoid maintaining browser installation and batch orchestration yourself: it offers one-call screenshots, bulk capture of up to 100 URLs per call, and bills only clean shots. See ScreenshotNeo and its API documentation.

Other managed options have different boundaries. AddScreenshots describes asynchronous bulk capture for URL lists, sitemaps, and domains, with output images stored in the customer’s cloud repository; verify its live endpoint and current terms before adopting it. Microsoft Playwright Workspaces is described as managed browser infrastructure, but the cited region overview lists Australia East, East Asia, East US, Japan East, Switzerland North, West Europe, and West US 3; India is absent from that displayed list. That page is a snapshot, so check current regional availability before making a decision.

For managed browser infrastructure, compare control over runtime and dependencies, region availability, job orchestration and artifact storage, and current terms. The available sources do not establish comparable prices, throughput benchmarks, or service-level guarantees for these alternatives.

Or skip the browser setup

Make one GET request to capture a URL. Replace the target URL and API key for your job. The API returns an image response; the example saves it as a WebP file.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For batch work, call the API for each URL with a concurrency limit that fits your job, and inspect the response headers to distinguish clean shots from other outcomes. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See the docs for request options and sign up for 1,000 free screenshots a month with no card.

Frequently asked questions

Can I use a sitemap instead of a URL file?

Yes. Parse the sitemap into the same ordered URL list before queuing work. Decide whether to include only canonical page URLs and how to handle sitemap index files.

Does choosing an Indian VPS make every site appear as it does to Indian visitors?

No guarantee follows from region alone. The capture origin changes the network path and may affect site presentation, but sites can also vary content by account, cookies, headers, or other signals.

Should I run one browser per URL?

The example reuses one browser and creates isolated contexts per capture. This avoids launching a separate browser process for every URL while keeping page state separate.

Can the job capture pages behind a login?

Only if you configure an allowed authentication flow or storage state for the site. The sample deliberately uses fresh contexts and does not include credentials; handle secrets carefully and follow the site’s access rules.

What should I retain when a run finishes?

Keep the manifest and screenshots needed for your workflow, then apply a retention policy to temporary files and logs. Persist important artifacts outside an ephemeral VM disk.