How to schedule recurring website screenshots with Google Cloud Run
Build a Cloud Run job that captures website screenshots on a schedule, saves them to Cloud Storage, and reports execution results.
Use a Cloud Run job for a finite screenshot task: package Chromium and browser automation code in a container, save each capture to Cloud Storage, then trigger the job with Cloud Scheduler. A Cloud Run job runs tasks to completion; a Cloud Run service handles HTTP requests. Google documents both the job trigger and the service-based alternative. Google’s job scheduling guide covers the direct job pattern.
This guide uses Node.js and Playwright for the capture container. The code is a practical implementation of Google’s documented headless browser support, not a Google-provided complete screenshot application. Google’s browser automation guide describes headless Chrome, Puppeteer, and Playwright for tasks such as taking screenshots.
1. Choose the scheduling architecture
For a recurring finite batch, use Cloud Scheduler to invoke a Cloud Run job. The job starts a container, captures one or more pages, writes the output, and exits. You can also use Scheduler to send an authenticated HTTP request to a Cloud Run service if your existing implementation is a request handler. Keep that service authenticated; Google’s guide advises against allowing public access for this pattern.
| Choice | Use it when | Trigger |
|---|---|---|
| Cloud Run job (recommended here) | A scheduled run performs a finite capture batch and exits. | Scheduler calls the Cloud Run Jobs :run API. |
| Cloud Run service | You already have a request-driven screenshot handler. | Scheduler sends an authenticated HTTP request to the service URL. |
The job and its scheduler caller have separate identities and permissions: the caller needs permission to run the job; the job identity needs permission to write captures to storage.
2. Create a runnable screenshot container
The following minimal application captures a viewport screenshot of the URL in TARGET_URL. It waits for DOM content and then for a short configurable settling period. Pages with delayed or interactive content may need a more specific readiness condition, discussed below.
Files
package.json
{
"name": "scheduled-site-capture",
"version": "1.0.0",
"private": true,
"scripts": { "start": "node capture.js" },
"dependencies": { "playwright": "1.52.0" }
}
capture.js
const { chromium } = require('playwright');
const fs = require('node:fs/promises');
const path = require('node:path');
async function main() {
const target = process.env.TARGET_URL;
if (!target) throw new Error('TARGET_URL is required');
const outputDir = process.env.OUTPUT_DIR || '/captures';
const settleMs = Number(process.env.SETTLE_MS || '1000');
const fullPage = process.env.FULL_PAGE === 'true';
await fs.mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 }, deviceScaleFactor: 1 });
page.setDefaultNavigationTimeout(Number(process.env.NAVIGATION_TIMEOUT_MS || '60000'));
const response = await page.goto(target, { waitUntil: 'domcontentloaded' });
if (!response) throw new Error('Navigation returned no main document response');
if (!response.ok()) throw new Error(`Navigation failed with HTTP ${response.status()}`);
if (process.env.READY_SELECTOR) await page.locator(process.env.READY_SELECTOR).waitFor({ state: 'visible', timeout: 30000 });
if (settleMs > 0) await page.waitForTimeout(settleMs);
// Use a unique name so retries and overlapping runs do not replace a prior capture.
const stamp = new Date().toISOString().replaceAll(':', '-');
const file = path.join(outputDir, `${stamp}-${process.pid}.png`);
await page.screenshot({ path: file, fullPage });
console.log(JSON.stringify({ event: 'capture_complete', url: target, file, status: response.status(), fullPage }));
} finally {
await browser.close();
}
}
main().catch((error) => {
console.error(JSON.stringify({ event: 'capture_failed', message: error.message, stack: error.stack }));
process.exitCode = 1;
});
The pinned Playwright version makes builds repeatable; choose and maintain a version compatible with your runtime. Playwright’s browser binaries must be present in the container. The Docker image below installs the Chromium browser dependencies and browser through Playwright’s installer.
Dockerfile
FROM node:22-bookworm
WORKDIR /app
COPY package.json ./
RUN npm install --omit=dev
RUN npx playwright install --with-deps chromium
COPY capture.js ./
ENV OUTPUT_DIR=/captures
CMD ["npm", "start"]
Build and push the image to Artifact Registry, then deploy it as a Cloud Run job. Replace project, region, repository, and image values with yours.
gcloud services enable run.googleapis.com cloudscheduler.googleapis.com artifactregistry.googleapis.com
gcloud artifacts repositories create screenshots \
--repository-format=docker --location=us-central1
gcloud builds submit --tag us-central1-docker.pkg.dev/PROJECT_ID/screenshots/site-capture:latest
gcloud run jobs create site-capture \
--image=us-central1-docker.pkg.dev/PROJECT_ID/screenshots/site-capture:latest \
--region=us-central1 \
--tasks=1 \
--max-retries=1 \
--task-timeout=10m \
--memory=2Gi \
--set-env-vars=TARGET_URL=https://example.com,SETTLE_MS=1000,FULL_PAGE=false
Set TARGET_URL to the site you are authorized to capture. For a page requiring an element before capture, set READY_SELECTOR, for example main. A selector that never appears makes the task fail at its timeout rather than saving a misleading partial capture.
3. Persist captures in Cloud Storage
A job’s local filesystem disappears when its task ends. Mount a bucket as a Cloud Run job volume when you want ordinary filesystem writes to persist. Grant the job’s service identity the Storage Object User role on the bucket (or an appropriately scoped resource) for object write access. See Google’s Cloud Storage volume mount documentation.
gcloud run jobs update site-capture \
--region=us-central1 \
--add-volume=name=captures,type=cloud-storage,bucket=YOUR_BUCKET \
--add-volume-mount=volume=captures,mount-path=/captures
Configure the job to run as a dedicated service account and grant that identity storage access. For example, after creating capture-writer@PROJECT_ID.iam.gserviceaccount.com and granting it the bucket role:
gcloud run jobs update site-capture \
--region=us-central1 \
--service-account=capture-writer@PROJECT_ID.iam.gserviceaccount.com
Cloud Storage FUSE volume writes consume container memory and do not provide full POSIX filesystem behavior. Google documents limitations including lack of file locking and last-writer-wins behavior when concurrent writes replace one file. Use unique object paths per run and avoid having parallel tasks write the same path. The mount can also affect startup time.
4. Schedule the job with Cloud Scheduler
Create a scheduler service account as the caller, then grant it only the permission needed to invoke the job. Google’s documented CLI pattern uses an authenticated OAuth POST to the Cloud Run Jobs :run endpoint. The Scheduler region does not have to match the job region.
gcloud iam service-accounts create screenshot-scheduler
gcloud run jobs add-iam-policy-binding site-capture \
--region=us-central1 \
--member=serviceAccount:screenshot-scheduler@PROJECT_ID.iam.gserviceaccount.com \
--role=roles/run.invoker
gcloud scheduler jobs create http site-capture-daily \
--location=us-central1 \
--schedule="0 12 * * *" \
--time-zone="Etc/UTC" \
--uri="https://run.googleapis.com/v2/projects/PROJECT_ID/locations/us-central1/jobs/site-capture:run" \
--http-method=POST \
--oauth-service-account=screenshot-scheduler@PROJECT_ID.iam.gserviceaccount.com
Replace the example cron expression and time zone with the desired recurrence. 0 12 * * * means daily at noon in the selected zone. Choose an explicit zone, especially when the schedule should follow local civil time across daylight-saving changes. The scheduler caller identity is distinct from the job identity that writes to the bucket.
5. Configure capture behavior and reliability
- Viewport or full page:
FULL_PAGE=truecaptures the page’s full vertical extent; the default captures the configured viewport. Very long pages can produce large images and use more memory. - Readiness:
domcontentloadedis a useful starting point, not proof that a single-page app has finished rendering. PreferREADY_SELECTORwhen a known element indicates useful content is present. Use a bounded settle delay for animations or deferred layout; avoid unbounded waits. - Timeouts: The sample uses a 60-second navigation timeout and a 30-second selector wait. Tune these for the pages you capture and keep the total within the Cloud Run task timeout.
- Retries: A retry can be useful for transient network failures, but it can also create duplicate captures. Unique filenames make retries safe for stored output. Set retry policy based on whether repeating the capture is acceptable.
- Parallelism: Start with one task. Increase task count or parallelism only when the URL list and target sites can handle the load. Consider target-site access rules and rate limits. The cited platform guidance does not establish a suitable rate for any specific website.
- Exit status: Throw on navigation, selector, or storage-write failures so Cloud Run records a failed task. Log structured details without including credentials, cookies, or other secrets.
Cloud Run job documentation lists a 10-minute default task timeout, a 168-hour (7-day) maximum for ordinary tasks, and three retries by default. These are platform defaults and limits; verify the current values in Google’s Create jobs documentation before deployment. Browser memory demand depends on page complexity, viewport and full-page height, and concurrency. Observe actual executions and adjust memory and timeout accordingly; no workload-specific performance benchmark is implied here.
6. Inspect executions and logs
Start a manual execution while bringing up the workflow, then inspect the execution and logs:
gcloud run jobs execute site-capture --region=us-central1 --wait
gcloud run jobs executions list --job=site-capture --region=us-central1
gcloud run jobs executions describe EXECUTION_NAME --region=us-central1
Job execution logs go to Cloud Logging and monitoring data goes to Cloud Monitoring. The console execution details view includes the most recent 1,000 executions and those from the previous seven days; older logs and metrics follow their respective retention policies. See Google’s job execution guide.
7. Cost and operational notes
The dossier does not establish a current cost estimate for a particular capture workload. Cloud Run compute and Cloud Storage use depend on task resources, run duration, image size, and retention; Scheduler and other Google Cloud charges also depend on current pricing and usage. Check the current pricing pages and estimate using your expected frequency, runtime, memory, and retained storage before setting a budget. Avoid retaining every run forever unless historical comparison requires it. A failed attempt still consumes execution resources even if it yields no usable image, so bounded retries and sensible timeouts help control waste.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Container exits with missing Chromium or browser executable | The browser binary or OS dependencies are absent from the image. | Install Chromium with Playwright during the image build, rebuild, and deploy the updated image. Keep the Playwright package and installed browser compatible. |
| Navigation timeout | The page is slow, network idle never occurs, or the timeout is too short. | Use a bounded navigation timeout appropriate to the page, wait for DOM content rather than network idle when suitable, and confirm the Cloud Run task timeout allows the whole run. |
| Screenshot is blank or incomplete | The page rendered after the capture, a required element was absent, or client-side content had not loaded. | Wait for a meaningful selector or a short bounded delay. Inspect the response status and execution logs; do not treat a successful navigation alone as proof of a complete page. |
| Scheduler returns permission denied | The Scheduler caller identity lacks permission to invoke the job, or the OAuth identity is misconfigured. | Grant the caller the documented invocation role on the job and ensure the Scheduler job uses that service account. |
| Scheduler job succeeds but no image appears | The trigger accepted the request but the Cloud Run execution failed, or output was written to ephemeral local storage. | Inspect Cloud Run execution status and Cloud Logging. Mount the bucket and use its mount path for persistent output. |
| Storage permission denied | The job’s runtime service identity lacks bucket write access. | Grant the job identity Storage Object User on the target bucket and verify the job is configured to use that identity. |
| Earlier screenshot is overwritten | Runs used the same destination filename. | Include a timestamp or execution identifier in each object name, especially when retries or overlapping executions are possible. |
| Task is killed or runs out of memory | Chromium, large pages, full-page capture, or concurrent work exceeds allocated memory. | Increase memory, reduce viewport/page size or concurrency, and avoid unnecessary simultaneous browser pages. Observe resource use on representative pages. |
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A single GET request returns a PNG, JPEG, WebP, or PDF. Its API and options are documented at ScreenshotNeo docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. For recurring captures, call the API from your scheduler or application and store the returned image where your workflow needs it. Sign up for 1,000 free screenshots a month, with no card.
FAQ
Does Cloud Scheduler need to run in the same region as the job?
No. The Scheduler region can differ from the Cloud Run job region.
Can one job capture multiple websites?
Yes. Extend the entry point to read a controlled URL list and capture each URL with a distinct output name. Bound concurrency and account for each page in the task timeout.
Should I use full-page screenshots for every run?
Only if the whole document is needed. Viewport captures use a predictable image size; full-page captures include content below the fold but can use more memory and produce larger files.
Can Scheduler call a Cloud Run service instead?
Yes. That is appropriate for an existing HTTP handler. Configure authenticated invocation rather than public access, as described in Google’s scheduling documentation.


