How to Capture Bulk Website Screenshots with the Google PageSpeed Insights API
PageSpeed Insights analyzes URLs one request at a time and does not document screenshot output in its current v5 API. Here’s how to automate audits and capture actual screenshots in bulk.
Short answer: The current Google PageSpeed Insights API v5 does not document screenshot capture. It analyzes one URL per request, so a bulk workflow means looping over your URL list and making a separate request for each page. If you need image files, run a separate screenshot capture workflow alongside the audits.
This guide shows how to run PageSpeed audits for a list of URLs, retain useful result metadata, and capture actual screenshots without confusing Lighthouse results with screenshots.
1. What the PageSpeed Insights API can and cannot do
The current method is GET https://pagespeedonline.googleapis.com/pagespeedonline/v5/runPagespeed. It requires a url parameter and returns JSON containing Lighthouse analysis. Optional parameters include category, locale, and strategy. The documented strategy choices are desktop and mobile; desktop is the default. The method reference does not document a screenshot parameter or screenshot image field. See Google’s current v5 method reference.
Google’s older v4 reference described screenshot and snapshot fields. Those historical fields are not evidence that current v5 supports screenshots. Do not add screenshot=true to v5 expecting an image. Use PageSpeed for performance analysis, and a browser capture tool or service for PNG, JPEG, WebP, or PDF output.
| Need | PageSpeed Insights API v5 |
|---|---|
| Audit a URL | Yes, one URL per request |
| Process a list | Yes, by orchestrating separate requests |
| Return a screenshot image | Not documented by the current v5 method |
| Choose mobile or desktop strategy | Yes |
| Choose a test geography | Not described by the method reference |
2. Prepare your URL list and request options
- Use fully qualified page URLs, including
https://. - Decide whether each audit should use
desktopormobile. - Specify categories when you need Lighthouse results beyond the default Performance category.
- Use a URL builder or query-parameter encoder. A target URL may itself contain query parameters, which must be encoded as a value inside the API request.
- For frequent automated queries, follow Google’s getting-started guide and use an API key. Check the current limits for your Google Cloud project; the sources reviewed here do not establish a numeric quota or batch size.
The API accepts repeated category parameters. Common documented category values include performance, accessibility, best-practices, and seo. The API response includes the resolved document ID, analysis timestamp, Lighthouse result, and version information. Preserve these with each result so you can identify what was analyzed and which Lighthouse version produced it.
3. Run audits for multiple URLs with cURL
This Bash example reads one URL per line from urls.txt, requests an audit for each, and saves each JSON response. Set PAGESPEED_API_KEY in your environment if using an API key. The target URL is encoded by cURL’s --data-urlencode.
#!/usr/bin/env bash
set -euo pipefail
API_KEY="${PAGESPEED_API_KEY:-}"
mkdir -p results
index=0
while IFS= read -r page_url || [[ -n "$page_url" ]]; do
[[ -z "$page_url" || "$page_url" == \#* ]] && continue
index=$((index + 1))
args=(--get "https://pagespeedonline.googleapis.com/pagespeedonline/v5/runPagespeed"
--data-urlencode "url=$page_url"
--data-urlencode "strategy=mobile"
--data-urlencode "category=performance")
if [[ -n "$API_KEY" ]]; then
args+=(--data-urlencode "key=$API_KEY")
fi
curl --fail-with-body --silent --show-error "${args[@]}" \
-o "results/audit-$index.json"
done < urls.txt
Change strategy to desktop when that is the run you want. Add another --data-urlencode "category=..." argument to request an additional category. This simple loop is sequential; see the reliability section before increasing concurrency.
4. Run the same workflow with Python
Install the HTTP client with python -m pip install requests. Put fully qualified URLs in urls.txt, one per line. This version writes each API response to a JSON file and records request failures separately rather than treating them as successful audits.
import json
import os
from pathlib import Path
import requests
ENDPOINT = "https://pagespeedonline.googleapis.com/pagespeedonline/v5/runPagespeed"
API_KEY = os.environ.get("PAGESPEED_API_KEY")
OUT = Path("results")
OUT.mkdir(exist_ok=True)
urls = [line.strip() for line in Path("urls.txt").read_text().splitlines()
if line.strip() and not line.lstrip().startswith("#")]
failures = []
for index, page_url in enumerate(urls, start=1):
params = {
"url": page_url,
"strategy": "mobile",
"category": ["performance", "accessibility"],
}
if API_KEY:
params["key"] = API_KEY
try:
response = requests.get(ENDPOINT, params=params, timeout=120)
response.raise_for_status()
payload = response.json()
(OUT / f"audit-{index}.json").write_text(
json.dumps({"requestedUrl": page_url, "response": payload}, indent=2)
)
except (requests.RequestException, ValueError) as exc:
failures.append({"url": page_url, "error": str(exc)})
(OUT / "failures.json").write_text(json.dumps(failures, indent=2))
Requests encodes the URL and repeated category values in the query string. The timeout is a client-side bound for waiting on a response, not a promise about Google’s processing time. Adjust it for your job runner and handle timeouts as failed attempts, not as audit results.
5. Run the workflow with Node.js
This example uses built-in fetch and runs on a modern Node.js runtime with global fetch support. It builds query parameters with URLSearchParams, including repeated categories, and saves each response.
import { readFile, mkdir, writeFile } from 'node:fs/promises';
const endpoint = 'https://pagespeedonline.googleapis.com/pagespeedonline/v5/runPagespeed';
const apiKey = process.env.PAGESPEED_API_KEY;
const outDir = 'results';
await mkdir(outDir, { recursive: true });
const urls = (await readFile('urls.txt', 'utf8'))
.split(/\r?\n/)
.map(line => line.trim())
.filter(line => line && !line.startsWith('#'));
const failures = [];
for (let i = 0; i < urls.length; i++) {
const pageUrl = urls[i];
const query = new URLSearchParams({ url: pageUrl, strategy: 'mobile' });
query.append('category', 'performance');
query.append('category', 'accessibility');
if (apiKey) query.set('key', apiKey);
try {
const response = await fetch(`${endpoint}?${query}`, {
signal: AbortSignal.timeout(120_000),
});
const body = await response.text();
if (!response.ok) throw new Error(`HTTP ${response.status}: ${body}`);
const payload = JSON.parse(body);
await writeFile(`${outDir}/audit-${i + 1}.json`, JSON.stringify({
requestedUrl: pageUrl,
response: payload,
}, null, 2));
} catch (error) {
failures.push({ url: pageUrl, error: String(error) });
}
}
await writeFile(`${outDir}/failures.json`, JSON.stringify(failures, null, 2));
6. Capture actual screenshots for a URL list
Keep screenshot capture as a separate step from PageSpeed analysis. The capture workflow should accept the same URL list, save image files, and report failures independently. Decide whether you need a viewport screenshot or a full-page image, which viewport or device dimensions matter, and which image format is useful downstream. PageSpeed’s strategy selects Lighthouse’s desktop or mobile analysis; it does not configure an image capture viewport.
If you build your own capture worker, use a browser automation library with a headless browser, navigate to each URL, wait for the page condition your workflow needs, and write the screenshot to a file. Be explicit about whether to wait for document load, a selector, a delay, or network idle: pages with analytics, long polling, or streaming requests may never become network-idle. Handle redirects, authentication, cookie banners, lazy-loaded content, and cross-origin restrictions according to your own capture requirements.
For a reliable batch, pair each requested URL with a stable output name, save the final URL after redirects if available, and capture a failure record when navigation or image writing fails. Avoid naming files directly from arbitrary URL text; sanitize filenames or use an index plus a URL hash to prevent collisions and invalid path characters.
7. Save traceable results and handle API changes
For every successful audit, retain at least:
- The requested URL and your selected strategy.
- The response’s resolved document
idandanalysisUTCTimestamp. - The returned version information and the Lighthouse result.
- Which categories you requested and the time the job completed in your system.
Do not assume every nested result field will always be present or stay identical across Lighthouse releases. Google’s release notes distinguish PageSpeed API v5 from Lighthouse versioning and report that PSI and its API were updated to Lighthouse 13.0 on October 20, 2025. Preserve returned version metadata and make parsers tolerant of missing or changing fields. See Google’s PageSpeed Insights release notes.
8. Reliability, performance, and cost considerations
Concurrency and request pacing
A list is not a single batch API call: every URL triggers its own request. Sequential processing is easiest to reason about and reduces pressure on your project limits. If you parallelize, use a small configurable worker pool, capture errors per URL, and tune against the actual Google Cloud limits and service guidance for your project. The reviewed getting-started material does not establish a numeric quota or safe concurrency value, so do not bake a guessed rate into the workflow.
Retries and partial completion
Keep successful responses even if other URLs fail. Record HTTP status, response body, URL, attempt count, and error category for failed items. Retry only transient failures according to the applicable Google Cloud quota and service guidance; do not retry indefinitely or treat an error payload as a valid Lighthouse result. Make jobs resumable by skipping already completed URLs or by writing each result atomically.
Run consistency
Performance results can vary between runs as the page and Lighthouse environment change. Store timestamps, strategy, requested categories, and returned version metadata. The method reference does not describe selecting a test region, so do not claim that a run was performed from a user-selected geography.
Cost and field data
Check current Google Cloud project limits and billing details before running large recurring jobs; the reviewed sources do not establish a numerical API quota or a per-request price. Google’s getting-started material says real-world CrUX data in the PSI API is planned for discontinuation and recommends the CrUX API or CrUX History API for that data. It does not provide an end date in the retrieved text, so treat this as Google’s stated plan, not as a completed removal. See Google’s getting-started guide.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The response is JSON, not an image | That is the documented v5 analysis response. | Use a separate screenshot capture workflow for image files. |
A screenshot query parameter has no effect |
The current v5 method does not document screenshot capture; screenshot fields appeared in historical v4 documentation. | Remove the assumption that v5 returns screenshots and use a capture tool or service. |
| Invalid or missing URL error | The required url parameter is absent, malformed, or not fully qualified. |
Supply a complete page URL and let a query encoder encode it as one parameter value. |
Target URLs with & or query strings fail |
The target URL was concatenated into the API URL without encoding. | Use --data-urlencode, URLSearchParams, or a library’s parameter dictionary. |
| Unauthorized or key-related error | The API key is missing, incorrect, or not configured for the request. | Follow Google’s API key setup guidance and verify the key and project configuration. |
| Some URLs fail while others succeed | Per-request failures are independent; a batch loop does not make the requests atomic. | Persist each success and failure separately, then resume or retry eligible failures. |
| Parser crashes on a missing nested field | Response content can change with Lighthouse versions or page conditions. | Check for field presence and preserve response version metadata. |
| Capture worker waits forever for network idle | The page may keep connections open for analytics, streaming, or long polling. | Use a bounded timeout or wait for a specific selector or a deliberate delay instead. |
10. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its bulk capture accepts up to 100 URLs per call, and its one-call screenshot endpoint can return an image or PDF. The PageSpeed API and ScreenshotNeo serve different jobs: use PSI for Lighthouse analysis and ScreenshotNeo when you need screenshot files.
For one URL, the basic request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the request options and bulk capture. Cookie banners are accepted before capture and more than 60 known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card required.
11. FAQ
Can I get a full-page screenshot from PageSpeed Insights v5?
The current method reference does not document screenshot output or a full-page capture option. Use a separate capture workflow for full-page images.
Does one PageSpeed request accept many URLs?
No. The documented method takes one url. Your script should issue one request per page.
Does mobile strategy create a mobile screenshot?
No. It selects the Lighthouse analysis strategy. Configure device dimensions in a separate screenshot capture workflow.
Where do CrUX field metrics fit?
Google’s getting-started material states that CrUX data in the PSI API is planned for discontinuation and points readers to the CrUX API or CrUX History API. Check the current guidance before designing a long-lived field-data pipeline.


