ScreenshotNeo

BlogHow-to

How to Save Batch Website Screenshots to Google Cloud Storage in India

Capture a list of web pages with Playwright and upload each screenshot to a Mumbai or Delhi Cloud Storage bucket, with runnable Node.js, Python and cURL examples.

By the ScreenshotNeo team4 October 202610 min read

Direct answer: To save batch website screenshots to Google Cloud Storage in India, create a bucket in Mumbai (ASIA-SOUTH1) or Delhi (ASIA-SOUTH2), capture each URL with a browser such as Playwright, then upload each screenshot file or byte buffer as an object. The example below uses Node.js, Playwright, and the official Google Cloud Storage client. It is an implementation pattern assembled from those APIs, not a built-in batch feature.

The bucket’s location controls where its object data resides; it does not by itself determine where the browser runs. Google Cloud states: “A bucket’s location defines the physical place where object data in the bucket resides.” See Google Cloud bucket locations.

1. Choose and create an India bucket

Choose the region that fits your deployment, data-location requirements, and resilience design. Google lists Mumbai as ASIA-SOUTH1 and Delhi as ASIA-SOUTH2. The research does not establish one as best for every workload.

Create the bucket in the selected location using the Google Cloud bucket creation guide or your preferred infrastructure tooling. Set access controls deliberately; screenshots are not automatically safe to expose publicly. The bucket location concerns stored object data and does not establish the location of browser execution, logs, or other processing.

2. Set up credentials and permissions

Run the script with Google Cloud application credentials configured for the environment and authenticated to the project. The Google client library uses that authentication context when it creates a storage client. Follow the official Cloud Storage upload guide for the supported credential setup in your runtime.

Grant the workload only the permissions it needs. Google identifies Storage Object User as a role with upload permissions. Replacing an existing object also requires storage.objects.delete. This example generates unique object names to avoid routine overwrites, so it does not need replacement behavior. For a more restrictive setup, verify the exact permissions against your chosen IAM configuration.

3. Capture and upload a batch with Node.js

This version keeps screenshot bytes in memory and uploads them directly. It captures the visible viewport by default; set FULL_PAGE=true to capture the whole scrollable page. Playwright supports both screenshot buffers and file paths, as well as full-page capture, in its screenshot guide and Page API.

npm install playwright @google-cloud/storage
npx playwright install chromium
// save-batch.js
const { chromium } = require('playwright');
const { Storage } = require('@google-cloud/storage');
const { createHash } = require('node:crypto');

const bucketName = process.env.GCS_BUCKET;
if (!bucketName) throw new Error('Set GCS_BUCKET to an existing bucket name');

const urls = [
  'https://example.com/',
  'https://example.com/about',
];
const fullPage = process.env.FULL_PAGE === 'true';
const storage = new Storage();
const bucket = storage.bucket(bucketName);

function objectKey(url, index) {
  const digest = createHash('sha256').update(url).digest('hex').slice(0, 16);
  return `screenshots/${new Date().toISOString().slice(0, 10)}/${String(index).padStart(4, '0')}-${digest}.png`;
}

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  const results = [];

  try {
    for (const [index, url] of urls.entries()) {
      const key = objectKey(url, index);
      try {
        const response = await page.goto(url, {
          waitUntil: 'networkidle',
          timeout: 45000,
        });
        if (!response || !response.ok()) {
          throw new Error(`Navigation returned ${response ? response.status() : 'no HTTP response'}`);
        }
        const bytes = await page.screenshot({ fullPage, type: 'png' });
        await bucket.file(key).save(bytes, {
          resumable: false,
          contentType: 'image/png',
          metadata: { cacheControl: 'private, max-age=0' },
        });
        results.push({ url, key, ok: true, bytes: bytes.length });
        console.log(`Uploaded ${url} -> gs://${bucketName}/${key} (${bytes.length} bytes)`);
      } catch (error) {
        results.push({ url, key, ok: false, error: error.message });
        console.error(`Failed ${url}: ${error.message}`);
      }
    }
  } finally {
    await browser.close();
  }

  const failures = results.filter(result => !result.ok);
  console.log(JSON.stringify({ total: results.length, failed: failures.length, results }, null, 2));
  if (failures.length) process.exitCode = 1;
})();

Run it after setting the bucket and credentials:

export GCS_BUCKET='your-existing-india-bucket'
node save-batch.js

Use a globally unique bucket name when creating the bucket; the script expects it to exist already. Object keys include the date, list index, and a short URL hash. If you run the same list more than once on the same date, those keys will repeat, so either add a run identifier to the key or deliberately configure overwrite permissions and behavior.

4. Adapt capture and storage options

Viewport or full page

Viewport captures are smaller and represent the visible browser area. Full-page captures include the full scrollable document and can be much taller or larger. For pages with lazy-loaded content, scrolling may be necessary before capture so images and other deferred content have loaded; account for that in page-specific logic rather than assuming navigation alone loads every asset.

Buffer or file

The example captures a buffer, avoiding temporary image files. This is convenient for modest batches and image sizes. To persist locally as well, pass a path to page.screenshot({ path: ... }), then upload it using the Cloud Storage file-upload method. The upload guide covers file uploads; Google also documents uploading from memory in its in-memory upload sample.

Image format and metadata

Playwright can capture PNG, JPEG, or WebP where supported by its screenshot API. Set the matching content type when saving the object, such as image/jpeg or image/webp. PNG is used above for broad compatibility. Consider adding metadata such as source URL, capture timestamp, and job identifier if your downstream workflow needs to trace objects back to inputs; avoid placing secrets or sensitive query parameters in object names or metadata.

5. Python alternative: Playwright and Cloud Storage

Install the libraries and browser, then authenticate the runtime as described in the Google Cloud documentation. This example uploads each screenshot buffer directly and continues after individual URL failures.

python -m pip install playwright google-cloud-storage
python -m playwright install chromium
# save_batch.py
import hashlib
import os
from datetime import datetime, timezone
from urllib.parse import urlsplit

from google.cloud import storage
from playwright.sync_api import sync_playwright

bucket_name = os.environ['GCS_BUCKET']
urls = [
    'https://example.com/',
    'https://example.com/about',
]
full_page = os.environ.get('FULL_PAGE', '').lower() == 'true'
client = storage.Client()
bucket = client.bucket(bucket_name)
day = datetime.now(timezone.utc).strftime('%Y-%m-%d')

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={'width': 1440, 'height': 900})
    failures = 0
    try:
        for index, url in enumerate(urls):
            digest = hashlib.sha256(url.encode()).hexdigest()[:16]
            key = f'screenshots/{day}/{index:04d}-{digest}.png'
            try:
                response = page.goto(url, wait_until='networkidle', timeout=45000)
                if response is None or not response.ok:
                    status = response.status if response else 'no HTTP response'
                    raise RuntimeError(f'Navigation returned {status}')
                image_bytes = page.screenshot(full_page=full_page, type='png')
                blob = bucket.blob(key)
                blob.upload_from_string(image_bytes, content_type='image/png')
                print(f'Uploaded {url} -> gs://{bucket_name}/{key} ({len(image_bytes)} bytes)')
            except Exception as exc:
                failures += 1
                print(f'Failed {url}: {exc}')
    finally:
        browser.close()

if failures:
    raise SystemExit(f'{failures} URL(s) failed')

The Python client also supports uploading a local file with blob.upload_from_filename(path). Use that route when you need a local archive or want to separate capture and upload stages.

6. cURL upload for already captured files

cURL does not render websites or create screenshots. Use it to upload an image file that another capture step has already produced. The following uses a Google Cloud access token available in the environment and the JSON API object upload endpoint. Obtain credentials through an approved Google authentication flow; do not hard-code or log access tokens.

export BUCKET='your-existing-india-bucket'
export OBJECT='screenshots/manual/example.png'
export ACCESS_TOKEN='your-short-lived-google-access-token'

curl --fail-with-body -X POST \
  -H "Authorization: Bearer ${ACCESS_TOKEN}" \
  -H 'Content-Type: image/png' \
  --data-binary @shot.png \
  "https://storage.googleapis.com/upload/storage/v1/b/${BUCKET}/o?uploadType=media&name=${OBJECT}"

URL-encode bucket object names when they contain spaces or reserved characters. For production, prefer the Cloud Storage client library’s credential handling and retry support over manually managing bearer tokens.

7. Make batch runs safe and repeatable

  1. Validate inputs. Reject malformed URLs and decide whether redirects are acceptable. Avoid logging full URLs if they can contain tokens or personal data.
  2. Use stable but collision-safe object names. Include a run ID or timestamp when retaining multiple versions. If intentionally replacing objects, grant the required delete permission and consider generation preconditions to prevent races.
  3. Record per-URL outcomes. Preserve the URL identifier, object key, status, duration, and error category in job results. Do not let one failed page silently make the whole batch appear successful.
  4. Retry transient errors selectively. Retry temporary navigation or upload failures with a bounded attempt count and backoff. Do not blindly retry deterministic errors such as invalid URLs, permission denial, or unsupported content.
  5. Limit concurrency. The examples run sequentially, which is easier on target sites and keeps memory use predictable. For larger batches, use a bounded worker pool and separate browser contexts or pages as appropriate; tune against available CPU, memory, target-site limits, and upload capacity.
  6. Keep capture environments fixed when comparing images. Browser version, operating system, settings, hardware, power, and headless mode can affect rendering. Playwright discusses these sources of variation in its visual comparisons guide; do not assume pixel-identical results across different environments.

8. Performance, reliability, and cost considerations

For a large batch, the main time costs are page loading, screenshot rendering, and upload. Reusing a browser process, as in the examples, avoids launching a browser for every URL. Sequential processing reduces resource spikes but takes longer; bounded concurrency can improve throughput, while increasing browser resource use and the load placed on destination websites.

Buffers avoid disk I/O but occupy memory until uploaded; very large screenshots or high concurrency can increase memory pressure. File-based capture can reduce the amount held in application memory, at the cost of local disk use and cleanup. Use explicit navigation timeouts and ensure a failed capture does not create a misleading success record.

Cloud Storage charges and broader cost details depend on current pricing, storage amount, operations, network transfer, and retention choices; check the current Google Cloud Storage pricing for your workload. This workflow also consumes compute to run the browser. The cited implementation sources do not establish a universal cost per screenshot or a performance benchmark.

9. Troubleshooting

Symptom Likely cause What to do
Bucket not found Typo, wrong project, or bucket not created. Check GCS_BUCKET, the authenticated project, and that the bucket exists.
403 or permission denied on upload Credentials are missing or the identity lacks object creation permission. Configure application credentials for the runtime and grant an appropriate upload role such as Storage Object User.
Overwrite fails with permission error The object key already exists and replacement requires delete permission. Use a unique run-specific key, or grant storage.objects.delete if replacement is intended.
Playwright says browser executable is missing The browser binary was not installed for the installed Playwright package. Run npx playwright install chromium or python -m playwright install chromium in the same environment.
Navigation times out or returns no response The site is slow, unreachable, blocked, or did not produce a normal HTTP response. Check connectivity and URL, increase the timeout only when justified, and record the URL as failed. A timeout should not be uploaded as a valid screenshot.
Capture is blank or missing images The page may still be rendering, defer images until scroll, require interaction, or render differently in headless mode. Wait for a relevant selector or page-specific readiness condition, scroll to trigger lazy loading, and inspect the page in the same browser environment.
Images differ across runs Browser or host environment, fonts, timing, animations, or dynamic content changed. Pin browser and runtime versions, use a consistent viewport and settings, and account for dynamic page content. Pixel identity is not guaranteed across environments.
High memory use Large screenshots and too many concurrent pages or buffers. Lower concurrency, upload promptly, or capture to temporary files and remove them after successful upload.

Or skip the browser setup

ScreenshotNeo can return a screenshot from one GET request, and its API documentation describes the available options. You still choose the India bucket and upload the returned bytes to it; the API call does not itself store the image in your bucket.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp
import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://example.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
// Upload bytes to your Google Cloud Storage bucket with the Storage client.

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media, with clean shots and billing only for clean shots. Visit ScreenshotNeo and sign up free for 1,000 screenshots a month with no card.

FAQ

Does choosing an India bucket mean the browser runs in India?

No. The bucket location describes where its object data resides. Browser execution location is a separate deployment choice.

Should I choose Mumbai or Delhi?

Choose based on your deployment arrangement, proximity needs, and data-location requirements. Both are documented India regional locations; neither is universally best.

Can I keep the screenshots private?

Yes. Uploading with authenticated credentials does not require making objects public. Set access according to your application’s needs.

Can this handle a very large URL list?

Yes, if you process it in bounded batches, track per-URL results, and tune concurrency and memory use for the browser host and target sites.