How to Make an AI Agent Take Website Screenshots from a Google Sheets URL List
Connect a Google Sheets URL list to Playwright, capture each site, and write clear success or failure results back to the right rows.
Direct answer: Read the URL column from Google Sheets, validate and filter the rows, then have an agent or workflow controller send each URL to a browser automation tool such as Playwright. The browser opens each page and saves a screenshot. Record the result against its original sheet row so failures are visible and duplicate URLs stay distinguishable.
The agent coordinates the workflow; Playwright performs the page navigation and screenshot capture. This guide uses a local Node.js script with the Google Sheets API and Playwright. You can adapt the same steps to another runtime or workflow platform.
1. Prepare the spreadsheet and choose the workflow
Create a sheet with a header row and one URL per row. Decide which tab and column contain the input URLs, and use an explicit A1 range such as Queue!A2:A. Google’s Sheets API needs the spreadsheet ID and A1 range to read values; if the range omits the tab name, it applies to the first sheet. See the Google Sheets API values guide.
| Approach | Best for | Trade-offs |
|---|---|---|
| Sheets API + Playwright script | Control over browser behavior, output names, retries, and storage | You manage credentials, runtime, browser installation, storage, and error handling. |
| Workflow platform + screenshot service | Low-code orchestration from spreadsheet rows | Check the provider’s current limits, pricing, retention, and terms yourself. The n8n integration listing shows a Google Sheets and GetScreenshot integration path; it does not establish those provider details. |
In either design, keep the row number or another stable row identifier with every job. A URL alone is not a sufficient identifier when the sheet contains duplicates.
2. Set up Google Sheets API access and Playwright
- In Google Cloud, create or select a project and enable the Google Sheets API.
- Create a service account, download its JSON credentials, and share the spreadsheet with the service account email. Grant only the access the workflow needs. For a read-only capture flow, use read access; add write access only if the script will update the sheet.
- Keep the credentials file outside source control. Set its path in an environment variable rather than embedding secrets in the script.
- Install Node.js, then create a project and install the Google API client and Playwright.
- Install the Playwright browser binaries for the browser engine you plan to use.
npm init -y
npm install googleapis playwright
npx playwright install chromium
Playwright supports browser engines and page controls; its Page API documents navigation and screenshot output. See the Playwright Page API and browser installation documentation.
3. Run a complete capture script
Set the spreadsheet ID, range, output directory, and credentials path. This script validates HTTP(S) URLs, skips the header and blank cells, records a result for each valid URL, continues after per-site errors, and writes the output file path and status to columns B and C of the source rows. It uses a viewport screenshot by default; set FULL_PAGE=true for full-page captures.
export GOOGLE_APPLICATION_CREDENTIALS="/secure/path/sheets-service-account.json"
export SPREADSHEET_ID="your_spreadsheet_id"
export INPUT_RANGE="Queue!A2:A"
export OUTPUT_DIR="./screenshots"
node capture.mjs
import { google } from 'googleapis';
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const spreadsheetId = process.env.SPREADSHEET_ID;
const inputRange = process.env.INPUT_RANGE ?? 'Queue!A2:A';
const outputDir = process.env.OUTPUT_DIR ?? './screenshots';
const fullPage = process.env.FULL_PAGE === 'true';
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);
const spreadsheetWriteRange = process.env.WRITE_RANGE ?? 'Queue!B2:C';
if (!spreadsheetId) throw new Error('Set SPREADSHEET_ID');
if (!process.env.GOOGLE_APPLICATION_CREDENTIALS) {
throw new Error('Set GOOGLE_APPLICATION_CREDENTIALS to the service-account JSON path');
}
if (!Number.isFinite(navigationTimeoutMs) || navigationTimeoutMs <= 0) {
throw new Error('NAVIGATION_TIMEOUT_MS must be a positive number');
}
const auth = new google.auth.GoogleAuth({
scopes: ['https://www.googleapis.com/auth/spreadsheets'],
});
const sheets = google.sheets({ version: 'v4', auth });
const { data } = await sheets.spreadsheets.values.get({
spreadsheetId,
range: inputRange,
});
const rows = data.values ?? [];
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
for (let i = 0; i < rows.length; i++) {
const sheetRow = i + 2; // INPUT_RANGE starts at row 2 in this example.
const raw = String(rows[i]?.[0] ?? '').trim();
if (!raw) continue;
let url;
try {
url = new URL(raw);
if (!['http:', 'https:'].includes(url.protocol) || !url.hostname) {
throw new Error('Only http and https URLs are accepted');
}
} catch (error) {
results.push([sheetRow, 'invalid_url', String(error.message)]);
continue;
}
const safeHost = url.hostname.replace(/[^a-zA-Z0-9.-]/g, '_');
const filename = `row-${sheetRow}-${safeHost}.png`;
const outputPath = path.resolve(outputDir, filename);
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
page.setDefaultNavigationTimeout(navigationTimeoutMs);
try {
const response = await page.goto(url.href, { waitUntil: 'domcontentloaded' });
if (response && response.status() >= 400) {
throw new Error(`HTTP ${response.status()}`);
}
await page.screenshot({ path: outputPath, fullPage });
results.push([sheetRow, 'success', outputPath]);
} catch (error) {
results.push([sheetRow, 'failed', String(error.message ?? error)]);
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
// Write row outcomes back to Queue columns B and C, beginning at row 2.
// Each item is [source row number, status, detail]. Writing a full contiguous
// range keeps statuses aligned with source rows. Empty rows are represented blank.
if (results.length) {
const lastRow = Math.max(...results.map(([row]) => row));
const byRow = new Map(results.map(([row, status, detail]) => [row, [status, detail]]));
const values = [];
for (let row = 2; row <= lastRow; row++) values.push(byRow.get(row) ?? ['', '']);
await sheets.spreadsheets.values.update({
spreadsheetId,
range: spreadsheetWriteRange,
valueInputOption: 'RAW',
requestBody: { values },
});
}
console.log(`Processed ${results.length} non-empty rows; outputs are in ${path.resolve(outputDir)}`);
Save the code as capture.mjs. If your input range starts somewhere other than row 2, adjust the sheetRow calculation and the write-back range to match. The example writes status and a local file path; that path is useful only to processes that can access the same machine or mounted storage. For shared access, upload files to storage you control and write back an access-controlled location.
What the script waits for
domcontentloaded waits for the document to be parsed without waiting for every image, ad, analytics request, or other resource to finish. Use load if the page needs its load event before capture, or wait for a specific selector or fixed delay when the page renders important content later. Network idle can be a poor fit for sites with long-running requests. Match the wait to the content you need; no one wait condition gives identical results across sites.
4. Configure the capture and the agent loop
- Viewport or full page: viewport captures show the initial visible area;
fullPage: truecaptures the full document and may create very tall images. Lazy-loaded content might require scrolling before the screenshot. - Browser engine: use Chromium, Firefox, or WebKit according to your compatibility needs. Install the chosen browser and keep it consistent between runs if comparable output matters.
- Viewport and scale: set viewport width and height for the intended device size. Configure device scale factor when pixel density matters.
- Format: Playwright screenshots can be PNG or JPEG; choose based on transparency and file size needs. The sample uses PNG.
- Authentication and state: provide cookies or storage state only when authorized. Do not put passwords or tokens in URLs, filenames, logs, or sheet cells.
- Per-row isolation: use a fresh page for each URL, close it after capture, and record the source row and original URL with the outcome. If sites must not share cookies or local storage, use a fresh browser context per site.
- Agent responsibility: let the agent decide which rows need work, invoke the capture worker with a validated URL and row ID, then inspect the status result. Keep the worker’s URL validation and network restrictions in place even if the agent has already checked inputs.
For larger queues, process rows in bounded batches and use modest concurrency until you understand the target sites and runtime limits. Retry transient navigation or network failures selectively with a small retry limit; do not repeatedly retry permanent HTTP errors or malformed URLs. No universal safe concurrency or batch size is established here, so measure the actual workload.
5. Handle URL safety, storage, and write-back
A sheet can contain arbitrary input. If a sheet is shared with others or the workflow runs on a server, consider an allowed-domain list and network-level restrictions. Otherwise, a browser with network access could be directed to unintended destinations. Block loopback, private-network, link-local, and internal service addresses where appropriate for your deployment; account for redirects too. Validate after URL parsing, and do not treat a valid URL syntax as proof that a destination is safe.
For screenshots containing confidential information, choose storage with access controls appropriate to that content. Avoid public links by default. Decide whether to retain local files, upload them, or delete them after a defined period. The sources do not prescribe a universal storage service or retention policy.
Write outcomes to the same source rows or to a separate results tab keyed by row ID. Include a clear status such as success, invalid_url, or failed, plus a useful detail. If multiple workers can process the sheet, prevent them from overwriting each other’s results; use a queue or claim/status mechanism. For a read-only workflow, omit the write scope and write-back code.
6. Reliability, performance, and cost
Each website is an independent job. Apply a bounded navigation timeout, close pages even after errors, continue after per-URL failure, and preserve enough context to diagnose the row. Store the original URL and timestamp in your run records if you need auditability. A screenshot is a point-in-time result: page appearance can differ with time, location, authentication, cookies, consent state, and dynamic content.
Browser startup and page rendering consume runtime and memory. Reusing one browser while creating isolated pages or contexts avoids launching a browser for each row; bounded parallelism can improve throughput but increases resource use and load on target sites. Large full-page captures can use more memory and storage than viewport images. The reviewed sources establish no benchmark, universal throughput, or quantitative cost comparison, so estimate from your own deployment and target list.
Self-hosting has costs in compute, storage, maintenance, and credential handling even when the software itself has no per-capture fee. A hosted capture step may reduce browser operations you manage, but check its price, limits, retention, and failure reporting before adopting it. The n8n integration listing confirms an integration path, not those commercial details.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Sheets API returns 403 | API is disabled, or the spreadsheet is not shared with the service account. | Enable the Sheets API in the project and share the sheet with the service account email. Confirm the account has the needed permission. |
| Sheets API returns 400 or range not found | Malformed A1 notation, incorrect tab name, or wrong spreadsheet ID. | Check the ID and exact tab name; quote tab names containing spaces, for example 'URL Queue'!A2:A. |
| Playwright says browser executable is missing | The browser binary was not installed for the selected engine or environment. | Run npx playwright install chromium in the deployment environment and ensure dependencies are available. |
| Navigation timeout | The site is slow, unreachable, or keeps requests open; the selected wait condition may be too strict. | Check the URL and network access, set a suitable bounded timeout, and use a targeted selector or domcontentloaded if appropriate. |
| Screenshot is blank or incomplete | The page may have failed to load, content renders later, or the wrong viewport or wait condition was used. | Record the navigation response, wait for the content selector, and verify viewport versus full-page behavior. Treat an empty or failed result as an error to investigate. |
| Output paths are inaccessible from the sheet | A local filesystem path is not a shared URL. | Upload to storage accessible to intended readers, set access controls, and write that location back. |
| Statuses are written to the wrong rows | The input range start row does not match the row-number calculation or write-back range. | Derive the source row from the actual A1 start row and test the mapping on a small sheet copy before processing a live queue. |
| Some URLs fail while others work | Individual sites may reject automated access, require login, or have network and content differences. | Respect access controls and site terms. Record the per-site failure and handle it manually or with an authorized authentication setup. |
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. A single GET request returns an image or PDF. For Google Sheets, read and validate the rows as above, then make one request per URL and save the returned bytes under a row-based filename.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
FAQ
Can an AI agent take the screenshot without a browser?
The agent still needs a capture mechanism. It can invoke a browser automation tool such as Playwright, or call a screenshot API such as ScreenshotNeo.
Can I use this with a sheet that has multiple URL columns?
Yes. Read the range covering those columns and define which column is the URL and which values identify the row before dispatching capture jobs.
Does a successful screenshot prove a page is safe or accurate?
No. It records what the capture system received at that time. Validate destinations, protect stored images, and interpret site-specific failures using their context.


