How to Use Playwright to Capture SERP Screenshots for a Keyword List
Capture SERP screenshots for a keyword list with Playwright. Build a repeatable batch script with controlled browser settings, clear filenames, and useful error handling.
Use Playwright to open each prepared search-results URL, wait for the page to reach a useful state, and save a screenshot with a deterministic filename. The script below processes a keyword list sequentially, records capture metadata, and lets you choose a viewport or full-page image. It does not construct search URLs or determine whether a search provider permits automated queries; prepare URLs using the provider’s current guidance and review its terms before collecting results.
1. Install Playwright and prepare your inputs
This example uses Node.js and the Playwright library. It expects a JSON file containing a list of keywords and their already-prepared target URLs. Keeping URL construction separate makes locale, query encoding, and provider-specific behavior explicit rather than assuming a universal SERP URL format.
mkdir serp-captures
cd serp-captures
npm init -y
npm install playwright
npx playwright install chromium
Create keywords.json:
[
{
"keyword": "best running shoes",
"url": "https://www.google.com/search?q=best%20running%20shoes"
},
{
"keyword": "weather in London",
"url": "https://www.google.com/search?q=weather%20in%20London"
}
]
The sample URLs illustrate the input shape only. Search providers can change URL behavior and may apply different rules to automated requests. Use URLs and collection methods permitted for your situation.
2. Capture the list with a repeatable script
Save this as capture-serps.mjs. It creates an output directory, handles duplicate or filesystem-hostile keywords, visits each target in order, saves a viewport screenshot by default, and writes a JSONL record for each successful capture or failure.
import { chromium } from 'playwright';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import path from 'node:path';
const inputPath = process.argv[2] ?? 'keywords.json';
const outputDir = process.argv[3] ?? 'captures';
const fullPage = process.env.FULL_PAGE === '1';
const locale = process.env.LOCALE ?? 'en-US';
const width = Number(process.env.WIDTH ?? 1365);
const height = Number(process.env.HEIGHT ?? 900);
const navigationTimeoutMs = Number(process.env.NAV_TIMEOUT_MS ?? 45000);
const settleMs = Number(process.env.SETTLE_MS ?? 1000);
if (!Number.isInteger(width) || width < 1 || !Number.isInteger(height) || height < 1) {
throw new Error('WIDTH and HEIGHT must be positive integers');
}
const entries = JSON.parse(await readFile(inputPath, 'utf8'));
if (!Array.isArray(entries)) throw new Error('Input JSON must be an array');
for (const [index, item] of entries.entries()) {
if (typeof item?.keyword !== 'string' || typeof item?.url !== 'string') {
throw new Error(`Item ${index + 1} must have string keyword and url fields`);
}
const parsed = new URL(item.url);
if (!['http:', 'https:'].includes(parsed.protocol)) {
throw new Error(`Item ${index + 1} URL must use HTTP or HTTPS`);
}
}
await mkdir(outputDir, { recursive: true });
const metadataPath = path.join(outputDir, 'metadata.jsonl');
const usedNames = new Map();
function safeName(value) {
const base = value.normalize('NFKD')
.replace(/[\\/:*?"<>|\u0000-\u001f]/g, '-')
.replace(/\s+/g, '-')
.replace(/[^\p{L}\p{N}._-]/gu, '')
.replace(/^[.-]+|[.-]+$/g, '')
.slice(0, 90);
return base || 'keyword';
}
function uniqueName(keyword, index) {
const base = safeName(keyword);
const seen = (usedNames.get(base) ?? 0) + 1;
usedNames.set(base, seen);
return `${String(index + 1).padStart(3, '0')}-${base}${seen > 1 ? `-${seen}` : ''}`;
}
const browser = await chromium.launch({ headless: true });
try {
const context = await browser.newContext({
viewport: { width, height },
deviceScaleFactor: 1,
locale,
colorScheme: 'light'
});
for (const [index, item] of entries.entries()) {
const page = await context.newPage();
const filename = `${uniqueName(item.keyword, index)}.png`;
const screenshotPath = path.join(outputDir, filename);
const startedAt = new Date().toISOString();
try {
page.setDefaultNavigationTimeout(navigationTimeoutMs);
const response = await page.goto(item.url, { waitUntil: 'domcontentloaded' });
// A short settle delay can allow client-rendered content to appear. Adjust
// it for the target page; fixed delays do not guarantee identical results.
if (settleMs > 0) await page.waitForTimeout(settleMs);
await page.screenshot({ path: screenshotPath, fullPage });
await appendFile(metadataPath, JSON.stringify({
keyword: item.keyword,
url: item.url,
capturedAt: new Date().toISOString(),
startedAt,
status: response?.status() ?? null,
finalUrl: page.url(),
browser: 'Chromium',
viewport: { width, height },
deviceScaleFactor: 1,
locale,
fullPage,
screenshot: filename
}) + '\n');
console.log(`Saved ${screenshotPath} (HTTP ${response?.status() ?? 'unknown'})`);
} catch (error) {
await appendFile(metadataPath, JSON.stringify({
keyword: item.keyword,
url: item.url,
startedAt,
failedAt: new Date().toISOString(),
error: String(error)
}) + '\n');
console.error(`Failed ${item.keyword}: ${error.message}`);
} finally {
await page.close();
}
}
await context.close();
} finally {
await browser.close();
}
Run it with:
node capture-serps.mjs keywords.json captures
Use environment variables to change the capture dimensions, locale, navigation timeout, settling delay, and full-page mode:
WIDTH=1440 HEIGHT=1000 LOCALE=en-GB SETTLE_MS=1500 node capture-serps.mjs keywords.json captures-uk
FULL_PAGE=1 node capture-serps.mjs keywords.json captures-full
The script records each failure and continues to later keywords. Review metadata.jsonl alongside the images: a screenshot is a snapshot of the page under the recorded conditions, not a guarantee that every searcher sees the same results.
3. Choose the right screenshot and browser state
Viewport, full page, element, or buffer
| Capture | Playwright call | Use it when | Trade-off |
|---|---|---|---|
| Viewport | page.screenshot({ path }) |
You need the initial visible results and a compact image. | Content below the current viewport is omitted. |
| Full page | page.screenshot({ path, fullPage: true }) |
You need below-the-fold results in one image. | The image can become very tall and larger to store or review. |
| Element | page.locator(selector).screenshot({ path }) |
A stable selector isolates the results region you want. | Search page markup can change; a brittle selector can fail or capture the wrong region. |
| Buffer | const bytes = await page.screenshot() |
You want to upload, hash, or process image bytes without writing first. | Your code must handle the bytes and storage destination. |
For an element capture, replace the screenshot call after navigation with a locator that you have verified for the target page:
const results = page.locator('main');
await results.waitFor({ state: 'visible', timeout: 10000 });
await results.screenshot({ path: screenshotPath });
The selector above is only an example; inspect the page and choose a selector that matches the result area for the specific provider and page version.
Make visual comparisons reproducible
Set the browser engine, viewport, device scale factor, locale, color scheme, and other relevant context settings explicitly. Use the same Playwright/browser version and execution environment for captures you intend to compare. Rendering can vary with operating system, browser version and settings, hardware, power source, and headless mode. Playwright’s visual comparison guide documents these sources of variation and its screenshot baseline workflow: Visual comparisons.
For baseline-based visual regression tests, Playwright Test provides toHaveScreenshot(). That is useful for comparing an expected page rendering in a controlled test environment; it is distinct from collecting a batch of changing search results.
Each browser context has isolated browser storage. Reuse one context when the batch should share its cookies and local storage. Create a fresh context per capture when each URL needs a clean state or separately configured locale, viewport, or storage. A context can also contain several pages; pages can navigate independently, though a sequential loop is the simpler starting point.
4. Adapt the workflow for a keyword list
- Prepare targets: store the original keyword and full URL for every capture. Encode query parameters with a URL builder when generating URLs in your own system.
- Choose a capture condition: fix viewport, locale, device scale, color scheme, and browser version. Record them in the metadata.
- Choose readiness behavior: this example uses
domcontentloadedplus a short configurable delay. If a known results selector exists, waiting for it to become visible is often more meaningful than choosing a long fixed delay. - Run sequentially first: it is easier to diagnose failures and uses fewer concurrent pages. Increase throughput only after checking resource use and the provider’s current rules.
- Review artifacts: match filenames to metadata, inspect failed records, and recapture only the affected inputs where practical.
For a site-specific readiness check, add a selector wait after navigation and before the screenshot:
await page.goto(item.url, { waitUntil: 'domcontentloaded' });
await page.locator('YOUR_STABLE_RESULTS_SELECTOR').waitFor({
state: 'visible',
timeout: 15000
});
await page.screenshot({ path: screenshotPath });
Use this only if you control or have a reliable selector for the page. Provider markup may vary across regions, sessions, experiments, or time.
5. Optional command-line capture
The Playwright CLI is convenient for a one-off or interactive capture. Install the CLI through the Playwright package, then use its screenshot command for a viewport, full page, element, output filename, image format, or high-resolution capture. Check the current options in the Playwright CLI documentation. For a keyword list, the library loop above is the more direct building block because it associates each URL with a deterministic output name and metadata.
6. cURL, Python, and Node.js examples with ScreenshotNeo
If you want to avoid managing a browser installation and screenshot loop, ScreenshotNeo provides a website screenshot API. One GET request returns an image or PDF. These examples capture a prepared SERP URL; use the target URL and query parameters appropriate to your workflow. The ScreenshotNeo API documentation describes its request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url="https://www.google.com/search?q=best%20running%20shoes" \
-o serp.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://www.google.com/search?q=best%20running%20shoes",
},
timeout=90,
)
r.raise_for_status()
with open("serp.webp", "wb") as image:
image.write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://www.google.com/search?q=best%20running%20shoes'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
writeFile('serp.webp', Buffer.from(await res.arrayBuffer()))
);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. See ScreenshotNeo for the service and the API documentation for options.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
Executable doesn't exist or browser launch fails |
The browser binary for the installed Playwright version is missing. | Run npx playwright install chromium in the project and ensure deployment includes the browser dependencies. |
| Navigation timeout | The page is slow, the network is unavailable, or a page keeps loading resources. | Check the target URL and network first. Set an appropriate NAV_TIMEOUT_MS; use domcontentloaded rather than waiting for every network request, and record the timeout as a failed capture. |
| Screenshot is blank or incomplete | Client-side content had not rendered, a consent overlay obscures results, or the target returned a bot check/error page. | Inspect the final URL and response status. Wait for a known visible results element where appropriate. Do not treat an error or challenge page as an ordinary SERP. |
| Images differ between runs | Locale, viewport, browser version, time, session state, experiments, or rendering environment changed. | Hold capture settings and environment steady, record metadata, and remember search results themselves can change. |
| Duplicate filename or invalid path | Terms normalize to the same filename or contain reserved characters. | The sample sanitizes names, prefixes an item number, and adds a duplicate suffix. Keep the original keyword in metadata. |
| Very large full-page image | The document is long or continues loading content while capture occurs. | Prefer viewport capture if only initial results matter. For full-page work, review image dimensions and storage needs and wait for the relevant content to stabilize. |
| Provider blocks requests or shows a challenge | Automated access may be restricted or the request pattern triggered provider controls. | Stop and review the provider’s current terms and guidance. Do not attempt to bypass access controls; use an authorized data source or permitted workflow. |
8. Performance, reliability, and cost considerations
- Start sequentially: each page has browser and network costs. A single page per keyword keeps memory use and failure handling straightforward.
- Bound concurrency if you add it: multiple pages can improve throughput, but unbounded navigation increases memory, CPU, network load, and the chance of throttling. There is no universal concurrency setting established by Playwright for SERP collection.
- Reuse the browser process: the sample launches once for the batch and closes pages after capture. Reusing a context is efficient when shared state is intended; use separate contexts when storage isolation matters.
- Choose capture size deliberately: viewport images are smaller and easier to compare; full-page images preserve below-the-fold content but can use more time and storage.
- Make failures visible: retain per-item errors and metadata. Retry transient network failures selectively with a small bounded retry policy; avoid retry loops that hammer a provider.
- Budget the work: local Playwright avoids a per-screenshot API charge but uses compute, storage, and maintenance time. With ScreenshotNeo, the free tier is 1,000 shots monthly and paid plans range from Starter at $5 for 3,000 to Business at $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Check current account usage and plan details before a recurring batch.
9. FAQ
Does a screenshot show the SERP everyone sees?
No. It documents the result rendered for that URL, browser state, locale, viewport, and capture time. Search results can differ across conditions and later change.
Can I capture several keywords at once with the Playwright CLI?
The CLI is suited to individual commands; the Playwright library API is the straightforward way to loop over a list and map each item to a filename.
Should I use a fresh browser context for each keyword?
Use a fresh context if each capture needs isolated cookies and local storage. Reuse one context if shared browser state is intended and documented in your metadata.
Does Playwright say whether automated searches are allowed?
No. The cited Playwright documentation explains browser automation and screenshots, not a search provider’s collection rules. Check the provider’s current terms and seek advice for your circumstances where needed.
Official Playwright references
- Screenshots: viewport, full-page, buffer, and element capture.
- Pages: creating pages and navigating.
- Browser contexts: independent contexts and browser state.
- Emulation: viewport, locale, and device settings.
- Visual comparisons: screenshot comparisons and rendering variability.
- CLI: screenshot command options.


