How to Schedule Website Screenshots with Google Cloud Scheduler
Use Cloud Scheduler to trigger a Chromium screenshot service on Cloud Run. Includes runnable Playwright code, secure scheduling, Cloud Storage, retries, and troubleshooting.
Short answer: Cloud Scheduler does not render pages or take screenshots itself. Configure it to send an authenticated, cron-scheduled HTTP request to a Cloud Run service. That service runs Chromium with Playwright or Puppeteer, captures the page, and saves the image to Cloud Storage or another destination.
This guide builds a small Node.js service with Playwright and Cloud Run, uploads PNGs to a private Cloud Storage bucket, and invokes it using an authenticated Cloud Scheduler job. It also covers duplicate deliveries, deadlines, alternatives for batch work, and common failures.
1. Choose the architecture
Use Cloud Scheduler as the trigger and Cloud Run as the browser worker. The Scheduler sends a request at the configured time; the service launches or connects to Chromium, navigates to the requested URL, takes the screenshot, and stores it. Google documents browser automation, including headless Chromium, as a Cloud Run use case. Cloud Run browser automation
- Trigger: Cloud Scheduler makes an HTTP request on a cron schedule.
- Worker: An authenticated Cloud Run service validates the request and captures the page using Playwright.
- Output: The service writes an image to Cloud Storage and returns a small success response.
For a recurring single-site capture or a small number of captures per invocation, an HTTP service is a straightforward fit. For finite batch work, compare Cloud Run jobs: a job runs a container to completion and exits, and can be a better match for batch-oriented tasks. The choice depends on whether you need request-response handling or batch execution, how you want to split URLs into tasks, where output goes, and how you handle failures. There is no universal workload-size threshold that determines which to use. Cloud Run jobs codelab
2. Create a Playwright screenshot service
The example accepts a URL in a JSON request, captures a full-page PNG, and stores it in a private bucket. It uses the scheduled execution time as the output key, so a retry for the same scheduled run overwrites the same object instead of creating another one. The service account attached to Cloud Run needs permission to create and overwrite objects in that bucket.
Application code
Create package.json:
{
"name": "scheduled-screenshot",
"version": "1.0.0",
"type": "module",
"scripts": { "start": "node server.js" },
"dependencies": {
"@google-cloud/storage": "^7.16.0",
"express": "^4.21.2",
"playwright": "^1.51.0"
}
}
Create server.js:
import express from 'express';
import { chromium } from 'playwright';
import { Storage } from '@google-cloud/storage';
import { createHash } from 'node:crypto';
const app = express();
app.use(express.json({ limit: '32kb' }));
const storage = new Storage();
const bucketName = process.env.SCREENSHOT_BUCKET;
const allowedHosts = (process.env.ALLOWED_HOSTS || '')
.split(',').map(s => s.trim().toLowerCase()).filter(Boolean);
function safeTarget(raw) {
const u = new URL(raw);
if (u.protocol !== 'https:' && u.protocol !== 'http:') {
throw new Error('Only http and https URLs are allowed');
}
if (allowedHosts.length && !allowedHosts.includes(u.hostname.toLowerCase())) {
throw new Error('Host is not on the allowed list');
}
return u;
}
app.post('/capture', async (req, res) => {
let browser;
try {
if (!bucketName) throw new Error('SCREENSHOT_BUCKET is not configured');
const target = safeTarget(req.body?.url || 'https://example.com');
const scheduleTime = req.get('X-CloudScheduler-ScheduleTime') || new Date().toISOString();
const runId = createHash('sha256').update(`${target.href}\n${scheduleTime}`).digest('hex').slice(0, 32);
const objectName = `scheduled/${runId}.png`;
browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(target.href, { waitUntil: 'networkidle', timeout: 60000 });
const bytes = await page.screenshot({ fullPage: true, type: 'png' });
await storage.bucket(bucketName).file(objectName).save(bytes, {
resumable: false,
contentType: 'image/png',
metadata: { metadata: { sourceUrl: target.href, scheduleTime } }
});
res.status(200).json({ ok: true, bucket: bucketName, object: objectName, bytes: bytes.length });
} catch (err) {
console.error(err);
res.status(500).json({ error: 'Capture failed', detail: err.message });
} finally {
if (browser) await browser.close().catch(() => {});
}
});
const port = Number(process.env.PORT || 8080);
app.listen(port, '0.0.0.0', () => console.log(`Listening on ${port}`));
The allowlist is optional, but recommended when the service only needs to capture known sites. Avoid accepting arbitrary public URLs in a publicly reachable endpoint: browser navigation can expose the worker to requests for internal addresses and other unwanted targets. Cloud Run authentication protects invocation; the allowlist constrains what an authorized caller can ask the browser to visit.
Container image
Create a Dockerfile. The Playwright image includes browser dependencies and Chromium; keep the Playwright package version aligned with the image tag.
FROM mcr.microsoft.com/playwright:v1.51.0-noble
WORKDIR /app
COPY package.json package-lock.json* ./
RUN npm install --omit=dev
COPY server.js ./
ENV NODE_ENV=production
CMD ["npm", "start"]
Generate and commit a lockfile with npm install before building for repeatable dependency resolution. Build and deploy from a Google Cloud project with Cloud Run, Cloud Build, Artifact Registry, Cloud Scheduler, and Cloud Storage APIs enabled.
3. Deploy Cloud Run and grant storage access
Set shell variables for your project and region. Replace the example values with your own names.
PROJECT_ID="your-project-id"
REGION="us-central1"
BUCKET="your-private-screenshot-bucket"
SERVICE="scheduled-screenshot"
RUNTIME_SA="screenshot-runtime@${PROJECT_ID}.iam.gserviceaccount.com"
SCHEDULER_SA="screenshot-scheduler@${PROJECT_ID}.iam.gserviceaccount.com"
gcloud config set project "$PROJECT_ID"
gcloud services enable run.googleapis.com cloudbuild.googleapis.com \
cloudscheduler.googleapis.com storage.googleapis.com
gcloud storage buckets create "gs://${BUCKET}" --location="$REGION"
gcloud iam service-accounts create screenshot-runtime
gcloud iam service-accounts create screenshot-scheduler
gcloud storage buckets add-iam-policy-binding "gs://${BUCKET}" \
--member="serviceAccount:${RUNTIME_SA}" \
--role="roles/storage.objectUser"
gcloud run deploy "$SERVICE" \
--source . \
--region "$REGION" \
--service-account "$RUNTIME_SA" \
--no-allow-unauthenticated \
--timeout 180 \
--memory 1Gi \
--set-env-vars "SCREENSHOT_BUCKET=${BUCKET},ALLOWED_HOSTS=example.com,www.example.com"
Use the runtime service account for Cloud Storage access. Do not put a service-account key in the container. The bucket remains private; grant readers access separately or create an application flow for sharing objects. Cloud Storage object names are not public URLs by themselves.
4. Create the authenticated Cloud Scheduler job
Give the Scheduler service account permission to invoke the service, then create the job with an OIDC token. Google recommends OIDC for authenticated HTTP targets such as Cloud Run; OAuth is generally for Google API targets. Cloud Scheduler HTTP target authentication
gcloud run services add-iam-policy-binding "$SERVICE" \
--region "$REGION" \
--member="serviceAccount:${SCHEDULER_SA}" \
--role="roles/run.invoker"
SERVICE_URL="$(gcloud run services describe "$SERVICE" \
--region "$REGION" --format='value(status.url)')"
gcloud scheduler jobs create http daily-screenshot \
--location "$REGION" \
--schedule "0 8 * * *" \
--time-zone "Etc/UTC" \
--uri "${SERVICE_URL}/capture" \
--http-method POST \
--update-headers "Content-Type=application/json" \
--message-body '{"url":"https://example.com"}' \
--oidc-service-account "$SCHEDULER_SA" \
--oidc-token-audience "$SERVICE_URL" \
--attempt-deadline 180s
This example runs daily at 08:00 UTC. Change --schedule and --time-zone for your desired cadence. Cloud Scheduler uses cron-style schedules. Specify a time zone deliberately; daylight-saving changes can shift local wall-clock behavior. Review Cloud Scheduler cron job configuration for schedule syntax and time-zone details.
The token audience should match the Cloud Run service URL. The HTTP request URI includes /capture; the token audience is the service URL. Ensure the service account named in the OIDC option has the Cloud Run Invoker role on the target.
5. Test the endpoint and schedule
Cloud Scheduler can run a job on demand for a smoke test:
gcloud scheduler jobs run daily-screenshot --location "$REGION"
gcloud scheduler jobs describe daily-screenshot --location "$REGION"
Inspect service logs and confirm that the expected object exists in the bucket:
gcloud run services logs read "$SERVICE" --region "$REGION" --limit 50
gcloud storage ls "gs://${BUCKET}/scheduled/"
For local development, run the container and send a request:
docker build -t scheduled-screenshot .
docker run --rm -p 8080:8080 \
-e SCREENSHOT_BUCKET=your-private-screenshot-bucket \
-e ALLOWED_HOSTS=example.com \
scheduled-screenshot
curl -X POST http://localhost:8080/capture \
-H 'Content-Type: application/json' \
-d '{"url":"https://example.com"}'
The local container needs Google credentials with permission to write to the bucket, or you can adapt the handler to write to a local path during development.
6. Configure capture behavior
The example uses Playwright’s page.goto and page.screenshot. Playwright supports viewport and full-page captures, file output or screenshot bytes, and PNG, JPEG, or WebP. See the Playwright Page API.
| Need | Playwright approach | Consideration |
|---|---|---|
| Viewport image | page.screenshot({ type: 'png' }) |
Captures the current viewport size. |
| Entire page | page.screenshot({ fullPage: true, type: 'png' }) |
Very long pages produce large images and may take longer to render. |
| JPEG or WebP | Set type: 'jpeg' or type: 'webp' where supported by the installed Playwright version. |
JPEG can use a quality value; confirm format support and downstream compatibility. |
| Specific element | await page.locator('.report').screenshot() |
Wait for the locator to be visible and ensure it is not clipped by layout. |
| Wait for app content | await page.locator('[data-ready="true"]').waitFor() |
Prefer a page-specific readiness signal when network idle is unreliable. |
| Output location | Save the screenshot bytes with the Cloud Storage client, as above, or pass a path to the screenshot API. |
Cloud Run’s local filesystem is ephemeral; persist results to Cloud Storage or another durable store. |
Other useful capture adjustments include setting a deliberate viewport, device scale factor, locale, timezone, color scheme, and authentication state in the browser context. Keep credentials out of the request body and source code; load secrets through an appropriate secret-management setup. If a site renders content after navigation, wait for a meaningful selector or a bounded delay rather than assuming the initial document load means the screenshot is ready.
7. Make retries safe and deadlines realistic
Cloud Scheduler delivers at least once, so rare duplicate invocations can occur. Google documents the job name and X-CloudScheduler-ScheduleTime header as useful identifiers for a scheduled execution. Cloud Scheduler job behavior
The sample derives a deterministic object key from the URL and scheduled time. Repeated attempts for that same URL and scheduled time write the same object name. If each attempt must preserve a separate artifact, use a deliberate versioning or attempt identifier strategy instead. If capture triggers additional side effects, such as notifications or database writes, make those operations idempotent too.
Align browser navigation timeout, Cloud Run request timeout, and Scheduler attempt deadline. The Cloud Scheduler HTTP target deadline defaults to three minutes and can be configured from 15 seconds to 30 minutes. Cloud Scheduler RPC reference Cloud Run services on a schedule
- Set navigation and readiness waits below the service timeout, leaving time for upload and response.
- Set the Scheduler deadline to cover the expected service duration, within the documented bounds.
- Do not acknowledge success before the image is safely stored if successful storage is the job’s purpose.
- Retries can repeat work. Use idempotent output names and inspect logs when Scheduler reports a failure after the service may have completed.
A mismatch can make Scheduler report a failed attempt even if the target completed. Use service logs and the object timestamp/name to distinguish a timeout from a missing capture.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Scheduler returns 401 or 403 | OIDC token audience mismatch, missing invoker role, or wrong service account. | Use the Cloud Run service URL as the token audience and grant the configured Scheduler service account roles/run.invoker. |
| Scheduler reports deadline exceeded | Page navigation, readiness waiting, or upload exceeded the attempt deadline. | Set bounded waits, increase the service and Scheduler deadlines as appropriate, or move longer batch work to a Cloud Run job. |
| Screenshot is blank or incomplete | The page needs client-side rendering, fonts or images have not loaded, or a readiness condition was missed. | Wait for a meaningful element or app-ready marker; inspect browser logs and use a bounded additional wait only when needed. |
| Navigation times out on a site that eventually loads | networkidle may not occur for pages with persistent connections or ongoing requests. |
Use waitUntil: 'domcontentloaded' or 'load', then wait for the target content selector. |
| Cloud Run exits or Chromium fails to launch | Browser dependencies are missing, package and container browser versions differ, or memory is insufficient. | Use a Playwright image aligned with the package version, retain required system dependencies, and adjust memory based on observed failures. |
| Upload returns permission denied | The Cloud Run runtime identity lacks bucket write permissions, or the bucket name is wrong. | Grant the runtime service account object-write access on the intended bucket and verify SCREENSHOT_BUCKET. |
| Duplicate screenshots or repeated side effects | At-least-once delivery retried a run. | Use a stable idempotency key based on the scheduled time and target, and make downstream writes safe to repeat. |
| Only part of a long page appears | Full-page capture occurred before lazy content loaded, or the page has unusual scrolling behavior. | Scroll through the page to trigger lazy loading, wait for content, and consider element captures or splitting the page. |
| Images are unexpectedly large | Full-page PNG at a large viewport preserves substantial detail. | Choose viewport-only capture, JPEG/WebP where suitable, or resize/compress after capture. |
9. Performance, reliability, and cost
Browser startup and page rendering dominate a small screenshot worker’s latency. Keep captures bounded, avoid loading resources the page does not need when appropriate, and choose viewport dimensions and output format based on downstream use. A full-page image can consume more memory and storage than a viewport image. Measure your own pages and workload; this guide makes no fixed performance claim.
Cloud Run instances may serve requests with different startup and concurrency behavior depending on configuration. For predictable browser isolation, keep the handler’s browser lifecycle explicit and close pages/browser processes after each capture. If optimizing startup by reusing a browser, isolate contexts per request and verify that state such as cookies and cache does not leak between captures.
Account for Cloud Scheduler job charges, Cloud Run compute and networking, Cloud Storage operations and stored bytes, and any egress. Exact costs depend on region, configuration, run duration, image size, and current pricing. Check the current Google Cloud pricing pages before estimating a production schedule. Retention policies or lifecycle rules can control how long captures remain stored.
10. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API returns an image or PDF from one GET request, so you can schedule an HTTP request without building and maintaining a Chromium container. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Store the API key as a secret in your scheduled caller. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000.
Sign up for 1,000 free screenshots a month, no card required.
FAQ
Can Cloud Scheduler take a screenshot by itself?
No. It schedules and invokes the worker. Chromium and browser automation code perform the capture.
Where does the screenshot get saved?
In this example, it is saved as a private object in the configured Cloud Storage bucket. You can change the handler to store it elsewhere, but Cloud Run’s local filesystem should not be treated as durable storage.
Can I capture several URLs in one scheduled run?
Yes. Extend the request schema to accept a list and define per-URL limits, failure handling, and idempotent object names. For finite batch workloads, evaluate a Cloud Run job.
What time zone does the cron schedule use?
The job specifies a time zone explicitly. Choose the zone that matches your intended schedule and account for daylight-saving transitions if using a local zone.
Should I choose Playwright or Puppeteer?
Either can automate Chromium for this workflow. This implementation uses Playwright; Google’s Cloud Run browser automation material and screenshot codelab also describe browser automation approaches. Choose the library that fits your existing code and deployment dependencies.


