ScreenshotNeo

BlogHow-to

How to Capture Website Screenshots in Bulk with Google Cloud Functions

Capture a small batch in Cloud Run functions, or fan out larger URL lists with Cloud Run jobs, Puppeteer, and Cloud Storage.

By the ScreenshotNeo team4 October 202612 min read

Direct answer: For a small, bounded batch, a Google Cloud function can start headless Chrome, capture each page with Puppeteer, and save the resulting images to Cloud Storage. For a genuinely large batch, use Cloud Run jobs to distribute URLs across tasks and store each task’s output in Cloud Storage. Google’s documented bulk screenshot architecture uses Cloud Run jobs, Workflows, and Eventarc; jobs are a separate run-to-completion product, not a function feature. Google’s screenshot example shows that pattern.

Google Cloud Functions is now called Cloud Run functions. The title’s older name is still widely recognizable, but use the current product name and check quotas for the generation and trigger you deploy. Google’s naming announcement.

1. Choose the right execution shape

Approach Good fit What to plan for
One HTTP Cloud Run function A few URLs, on-demand capture, or a bounded request Browser startup, navigation, capture, and uploads all count against one invocation deadline. Validate the request and cap the number of URLs.
One event-driven function per URL Independent work items triggered by events Each event should contain a URL or work-item reference, not screenshot bytes. Add idempotency and retry handling.
Cloud Run job tasks Large independent lists and repeatable batches Split the manifest across tasks, bound task parallelism and timeout, and write results to object storage.

For large batches, the job pattern avoids holding a single HTTP request open for the whole workload. Google’s example uploads a text file of URLs to Cloud Storage, uses Eventarc to start a Workflow, and has the Workflow run a Cloud Run job with a task per URL. Each task uses its task index to select a URL and uploads the screenshot. The example retains job data after failures for investigation. Architecture details.

2. Small-batch implementation with a Cloud Run function

This Node.js HTTP function accepts a JSON body containing a bounded urls array, captures each URL sequentially, and uploads PNG files to a configured Cloud Storage bucket. It returns a small per-URL result list rather than image bytes. Sequential processing keeps this simple example’s memory and browser load bounded; for more throughput, move work into independently retryable job tasks rather than simply increasing concurrency in one invocation.

Files

{
  "name": "bulk-screenshot-function",
  "version": "1.0.0",
  "type": "module",
  "main": "index.js",
  "engines": { "node": ">=20" },
  "dependencies": {
    "@google-cloud/functions-framework": "^3.4.0",
    "@google-cloud/storage": "^7.0.0",
    "puppeteer": "^24.0.0"
  }
}

Save as package.json. Pin versions to the versions you qualify for your runtime and rebuild and validate the browser after dependency upgrades.

import { http } from '@google-cloud/functions-framework';
import { Storage } from '@google-cloud/storage';
import puppeteer from 'puppeteer';
import { createHash } from 'node:crypto';

const storage = new Storage();
const bucketName = process.env.OUTPUT_BUCKET;
const maxUrls = Number(process.env.MAX_URLS ?? 10);
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);

function validPublicHttpUrl(value) {
  try {
    const u = new URL(value);
    return (u.protocol === 'http:' || u.protocol === 'https:') &&
      !u.username && !u.password;
  } catch {
    return false;
  }
}

http('captureBatch', async (req, res) => {
  if (req.method !== 'POST') {
    res.set('Allow', 'POST');
    return res.status(405).json({ error: 'Use POST.' });
  }
  if (!bucketName) return res.status(500).json({ error: 'OUTPUT_BUCKET is not configured.' });
  const urls = req.body?.urls;
  if (!Array.isArray(urls) || urls.length === 0 || urls.length > maxUrls ||
      !urls.every(u => typeof u === 'string' && validPublicHttpUrl(u))) {
    return res.status(400).json({ error: `Provide 1-${maxUrls} valid HTTP(S) URLs in {"urls": [...]} .` });
  }

  let browser;
  const results = [];
  try {
    browser = await puppeteer.launch({
      headless: true,
      args: ['--no-sandbox', '--disable-setuid-sandbox']
    });
    const bucket = storage.bucket(bucketName);
    for (const url of urls) {
      const id = createHash('sha256').update(url).digest('hex').slice(0, 20);
      let page;
      try {
        page = await browser.newPage();
        await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
        const response = await page.goto(url, {
          waitUntil: 'networkidle2',
          timeout: navigationTimeoutMs
        });
        if (!response) throw new Error('Navigation returned no main-document response.');
        const bytes = await page.screenshot({ type: 'png', fullPage: true });
        const objectName = `screenshots/${id}.png`;
        await bucket.file(objectName).save(bytes, {
          resumable: false,
          contentType: 'image/png',
          metadata: { cacheControl: 'private, max-age=0' }
        });
        results.push({ url, ok: true, httpStatus: response.status(), object: objectName });
      } catch (error) {
        results.push({ url, ok: false, error: String(error?.message ?? error).slice(0, 500) });
      } finally {
        await page?.close().catch(() => {});
      }
    }
  } catch (error) {
    console.error('Batch setup failed', error);
    return res.status(500).json({ error: 'Browser or batch setup failed.' });
  } finally {
    await browser?.close().catch(() => {});
  }
  const failed = results.filter(r => !r.ok).length;
  return res.status(failed ? 207 : 200).json({ total: results.length, failed, results });
});

Save as index.js. This validates URL syntax but does not prevent access to private network addresses or metadata endpoints. If callers are not fully trusted, enforce a destination allowlist and block private, loopback, link-local, and metadata IP ranges at the network layer; DNS resolution and redirects must also be considered. Do not expose an unrestricted screenshot endpoint to the public internet.

Deploy and call

gcloud services enable cloudfunctions.googleapis.com run.googleapis.com cloudbuild.googleapis.com artifactregistry.googleapis.com

gcloud storage buckets create gs://YOUR_UNIQUE_OUTPUT_BUCKET --location=YOUR_REGION

gcloud functions deploy captureBatch \
  --gen2 \
  --runtime=nodejs20 \
  --region=YOUR_REGION \
  --source=. \
  --entry-point=captureBatch \
  --trigger-http \
  --memory=1GiB \
  --timeout=300s \
  --set-env-vars=OUTPUT_BUCKET=YOUR_UNIQUE_OUTPUT_BUCKET,MAX_URLS=10,NAVIGATION_TIMEOUT_MS=30000 \
  --no-allow-unauthenticated

Grant the function’s runtime service account permission to create objects in the output bucket using the narrowest applicable Cloud Storage role. The deployer identity and runtime identity are different concerns: the deployer needs deployment permissions, while the running function needs bucket access. For a private function, obtain an identity token and pass it in the Authorization header:

FUNCTION_URL="$(gcloud functions describe captureBatch --gen2 --region=YOUR_REGION --format='value(serviceConfig.uri)')"
TOKEN="$(gcloud auth print-identity-token)"
curl -X POST "$FUNCTION_URL" \
  -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"urls":["https://example.com","https://www.google.com"]}'

The example uses networkidle2, which is a practical default, not a universal definition of visual readiness. Some sites keep analytics or application connections open. For those pages, use waitUntil: 'domcontentloaded' and an explicit selector or bounded delay appropriate to the target, or apply a per-site readiness policy. For image-heavy pages, full-page capture can require substantially more memory than a viewport screenshot.

3. Bulk architecture with Cloud Run jobs

For many URLs, make each task handle one URL or a small chunk and keep the manifest and outcomes in Cloud Storage. Add Workflows and Eventarc when uploads or another event should start the batch automatically. A job is a containerized run-to-completion workload; it does not receive an HTTP request like the function above.

Task worker pattern

The following Node.js worker illustrates the core per-task work pattern from Google’s documented design: task index selects a URL, Puppeteer captures it, and Cloud Storage receives the image. The manifest is a newline-delimited text object referenced by MANIFEST_URI. This code assumes the manifest is available locally at /workspace/urls.txt; adapt the startup step to download the configured Cloud Storage object before running the worker. The task-index approach requires one assigned URL per task; for chunking, partition the manifest deterministically instead.

import { Storage } from '@google-cloud/storage';
import puppeteer from 'puppeteer';
import { readFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';

const lines = (await readFile('/workspace/urls.txt', 'utf8'))
  .split(/\r?\n/).map(x => x.trim()).filter(Boolean);
const index = Number(process.env.CLOUD_RUN_TASK_INDEX ?? 0);
const url = lines[index];
if (!url || !/^https?:\/\//i.test(url)) throw new Error(`No valid URL for task ${index}`);
const bucketName = process.env.OUTPUT_BUCKET;
if (!bucketName) throw new Error('Set OUTPUT_BUCKET');

let browser;
try {
  browser = await puppeteer.launch({ headless: true, args: ['--no-sandbox', '--disable-setuid-sandbox'] });
  const page = await browser.newPage();
  await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
  const response = await page.goto(url, { waitUntil: 'networkidle2', timeout: 30000 });
  if (!response) throw new Error('No main-document response');
  const image = await page.screenshot({ type: 'png', fullPage: true });
  const id = createHash('sha256').update(url).digest('hex').slice(0, 20);
  const object = `screenshots/${index}-${id}.png`;
  await new Storage().bucket(bucketName).file(object).save(image, {
    resumable: false, contentType: 'image/png'
  });
  console.log(JSON.stringify({ index, url, status: response.status(), object }));
} finally {
  await browser?.close();
}

Package this as a container with a pinned Node.js base image and compatible Puppeteer/Chrome installation. Google’s Cloud Run browser automation guide lists headless Chrome with Puppeteer, Playwright, or the Chrome DevTools Protocol as approaches. Validate the exact image and library versions together when upgrading. Browser and OS automation in Cloud Run.

At orchestration time, ensure task count matches the number of work items if using one URL per task. Configure parallelism to protect both your Cloud project and target sites. A failed task should leave enough metadata to identify its URL and error. Retry transient navigation or infrastructure failures selectively; do not retry permanent errors indefinitely. Use stable object names or a completion manifest so a retried task is safe to run again.

4. Capture behavior and configuration choices

Choice Options and guidance
Capture scope fullPage: true captures the full document; omit it for the current viewport. Tall pages increase image size and memory requirements.
Image format Puppeteer screenshots support PNG, JPEG, and WebP in applicable versions. PNG preserves sharp text and transparency behavior; JPEG or WebP can reduce storage size. Set matching content type and file extension.
Viewport Set width, height, and device scale factor before navigation or capture. A larger scale factor increases output pixel dimensions and memory use.
Readiness Use a finite navigation timeout. Choose a load condition per workload; for app-specific readiness, wait for a selector or known state. Avoid an unbounded wait.
Output Upload bytes to Cloud Storage. Use deterministic names plus a run identifier if preserving multiple captures. Keep the HTTP response to statuses and object references.
Failure policy Record results per URL. Separate permanent failures from timeouts and temporary infrastructure errors, and cap retries with backoff.
Security Restrict who can invoke the function and who can read outputs. Validate destinations to prevent server-side request forgery when input is not trusted.

5. Limits, performance, reliability, and cost

Invocation limits

Google’s Cloud Run functions quota page lists maximum function durations by generation and trigger: 1st generation is 540 seconds; 2nd generation is up to 60 minutes for HTTP, 1,800 seconds for scheduled/task-queue functions, and 540 seconds for event-driven functions. The same documentation lists request and response size limits, including 32 MB for an uncompressed HTTP request and 32 MB for a non-streaming 2nd-generation HTTP response. These are maximum quotas, not recommended batch sizes; confirm the live quota and deployment configuration before relying on them. Cloud Run functions quotas.

The Google quota reference also lists first-generation maximum memory as 8 GiB. Available settings and behavior depend on generation and configuration; consult the current documentation for the deployed function and Cloud Run job. Generation-specific limits.

Estimate capacity with measurements

  1. Capture representative pages in the intended region and runtime, including browser startup, readiness wait, screenshot generation, and upload.
  2. Measure memory and elapsed time for both ordinary and unusually long pages.
  3. Set task timeout below the platform maximum with room for shutdown and reporting.
  4. Choose task count and parallelism so each task stays within its deadline and memory budget.
  5. Repeat the measurement after browser, library, runtime, or page mix changes.

There is no universal URLs-per-invocation figure: page complexity, network, readiness policy, viewport, image dimensions, and concurrency all affect throughput. Google’s reviewed examples do not provide a general performance benchmark.

Reliability and data handling

  • Use per-URL results so one inaccessible site does not erase the status of every other URL.
  • Use idempotent object naming or a run ID to make retries predictable.
  • Retry only transient failures, with a small cap and backoff. Preserve terminal errors for diagnosis.
  • Close pages and the browser in cleanup paths. Treat browser crashes as task failures that can be retried.
  • Store the source URL, capture time, status, and error metadata alongside the image, while limiting sensitive URL data in logs.
  • Define retention and bucket access to match the sensitivity of captured pages.

Cost considerations

Cloud cost depends on function or job compute time and memory, browser execution, network egress, Cloud Storage operations and stored bytes, and any orchestration services used. Larger full-page images and excessive parallelism can increase resource use. Check current Google Cloud pricing for the region and services you choose; the research sources provide no fixed cost or benchmark for this workload. Reduce unnecessary recaptures with a cache or schedule policy when freshness requirements allow.

6. Troubleshooting

Symptom Likely cause Fix
Browser fails to launch Browser binary missing, incompatible library/browser versions, insufficient memory, or missing container dependencies Pin and package compatible versions, inspect startup logs, and validate the built runtime image. Increase memory only after checking the actual failure.
Navigation timeout Slow target, blocked request, or a page that never becomes idle Set a finite timeout and use a readiness condition suited to the site, such as DOM content plus a specific selector. Record the timeout as a per-URL result.
Screenshot is blank or incomplete Capture occurred before client-side rendering or lazy content finished Wait for a meaningful selector or known page state. For lazy-loaded content, scroll in controlled increments and wait for images before capture.
Function returns 413 or request rejected Request body exceeds the function’s applicable size limit Send a manifest reference or object name instead of a huge URL array. Keep invocation payloads small.
Function times out on a batch Combined browser startup, page rendering, and uploads exceed the configured deadline Reduce the batch size or move the work to Cloud Run job tasks. Measure representative pages before sizing.
Permission denied on upload Runtime service account lacks bucket write permission, or bucket/project configuration is wrong Verify the runtime identity and grant the least privilege needed to create objects in the destination bucket.
Some URLs return errors while others succeed Redirect, TLS issue, bot defense, invalid URL, or site-specific failure Keep per-URL outcomes, inspect response status and logs, and retry only likely transient failures. Respect site access rules.
Repeated objects overwrite one another Output names are not unique for the chosen batch/run semantics Use a run identifier and stable per-item key, or explicitly define that each URL’s latest capture replaces the prior one.
High memory use or browser crash on long pages Full-page rasterization or concurrent pages consume too much memory Capture viewport-only when sufficient, reduce scale or concurrency, or assign one page per task.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF, and it can handle capture options such as full-page screenshots, device presets, waiting for a selector, and custom headers. Its API parameter names used by other screenshot APIs also work. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers say the page verdict and whether the request was billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, no card required.

8. FAQ

Can I return all screenshot files directly from the function?

It is usually a poor fit for a batch: binary responses grow quickly and are subject to HTTP response limits. Upload images to Cloud Storage and return object names and statuses.

Does networkidle2 mean the screenshot is visually complete?

No. It describes a network activity condition, not a guarantee that every application has finished rendering or that lazy content has loaded. Define readiness for the sites you capture.

Should every URL be a separate task?

One URL per task makes retries and failure attribution straightforward, as in Google’s example. Small chunks can reduce task overhead, but then the worker needs per-URL error handling and a bounded chunk size.

Can I use Playwright instead of Puppeteer?

Yes. Google lists both Puppeteer and Playwright as ways to control headless Chrome in Cloud Run. Choose one browser automation library and qualify its runtime dependencies with the browser image you deploy.

Sources