How to Screenshot a List of Authenticated Web App Pages in Bulk
Reuse a Playwright login session to capture authenticated pages in bulk, with reliable readiness checks, safe credential handling, and per-page error records.
To screenshot a list of authenticated pages in bulk, sign in once with Playwright, save the browser’s authenticated state to a private file, then load that state into a browser context and visit each URL. Wait for an application-specific readiness signal on each page, save a viewport or full-page screenshot, and record failures per URL so one expired session or inaccessible page does not stop the rest of the batch.
This guide uses Playwright with Node.js. It also includes the equivalent capture pattern in Python. Keep the state file secret: it may contain cookies and other credentials that can let someone impersonate the account. [Playwright’s authentication guide](https://playwright.dev/docs/auth) recommends saving and reusing state and strongly discourages checking it into repositories.
1. Prepare a dedicated login state
Use a dedicated automation account with only the access the pages require. Install Playwright, create a project directory, and add the state file to your local ignore rules before signing in.
npm init -y
npm install playwright
npx playwright install chromium
mkdir -p screenshots
printf '\nplaywright/.auth/\n' >> .gitignore
mkdir -p playwright/.auth
Create login.mjs. Replace the URL and selectors with those used by your application. This script waits for an explicit signal that login succeeded before writing the storage state.
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: false });
const context = await browser.newContext();
const page = await context.newPage();
try {
await page.goto('https://app.example.com/login', { waitUntil: 'domcontentloaded' });
// Replace these selectors and credentials with your app's login flow.
await page.getByLabel('Email').fill(process.env.APP_EMAIL);
await page.getByLabel('Password').fill(process.env.APP_PASSWORD);
await page.getByRole('button', { name: 'Sign in' }).click();
await page.getByRole('heading', { name: 'Dashboard' }).waitFor({ state: 'visible' });
await context.storageState({ path: 'playwright/.auth/state.json', indexedDB: true });
console.log('Saved authenticated state.');
} finally {
await context.close();
await browser.close();
}
Provide the credentials through your environment or secret manager; do not put them in source code or logs. Run the script in an environment where you can complete any required multi-factor or interactive login step. Playwright storage state covers cookies, local storage, IndexedDB when requested, and passkey-related state supported by the framework. It does not automatically preserve sessionStorage; see the edge cases below if the application relies on it. See [Playwright authentication](https://playwright.dev/docs/auth) for state setup details.
2. Capture the URL list with the saved state
Save the input URLs in urls.txt, one URL per line. The script below uses one browser process and one authenticated context, navigates sequentially, waits for an application-specific selector, captures each page, and writes a JSON-lines result record for every URL. Sequential capture is a conservative starting point for shared accounts and services with rate limits.
https://app.example.com/reports/monthly
https://app.example.com/settings/profile
https://app.example.com/projects/42
Create capture.mjs. Change READY_SELECTOR to a stable element present when the requested content is ready, and set FULL_PAGE=true if you need the full document instead of the viewport.
import { chromium } from 'playwright';
import { readFile, appendFile, mkdir } from 'node:fs/promises';
import path from 'node:path';
const STATE = 'playwright/.auth/state.json';
const URL_FILE = 'urls.txt';
const OUTPUT_DIR = 'screenshots';
const READY_SELECTOR = '[data-testid="page-content"]'; // Replace for your app.
const FULL_PAGE = process.env.FULL_PAGE === 'true';
const TIMEOUT_MS = Number(process.env.PAGE_TIMEOUT_MS || 30000);
function safeName(url, index) {
const parsed = new URL(url);
const slug = (parsed.pathname.split('/').filter(Boolean).join('-') || 'home')
.replace(/[^a-zA-Z0-9_-]/g, '-')
.slice(0, 70);
return `${String(index + 1).padStart(3, '0')}-${slug}.png`;
}
await mkdir(OUTPUT_DIR, { recursive: true });
await appendFile('capture-results.jsonl', '');
const urls = (await readFile(URL_FILE, 'utf8'))
.split(/\r?\n/).map(line => line.trim()).filter(Boolean);
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
storageState: STATE,
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
});
try {
for (const [index, url] of urls.entries()) {
const page = await context.newPage();
const startedAt = new Date().toISOString();
const output = path.join(OUTPUT_DIR, safeName(url, index));
let record;
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: TIMEOUT_MS });
if (response && response.status() >= 400) {
throw new Error(`Navigation returned HTTP ${response.status()}`);
}
await page.locator(READY_SELECTOR).waitFor({ state: 'visible', timeout: TIMEOUT_MS });
// Optional: wait for app-specific loading indicators to disappear.
await page.screenshot({ path: output, fullPage: FULL_PAGE, animations: 'disabled' });
record = { url, output, startedAt, finishedAt: new Date().toISOString(), ok: true };
console.log(`Saved ${output}`);
} catch (error) {
record = {
url, startedAt, finishedAt: new Date().toISOString(), ok: false,
error: error instanceof Error ? error.message : String(error),
finalUrl: page.url(),
};
console.error(`Failed ${url}: ${record.error}`);
} finally {
await page.close();
await appendFile('capture-results.jsonl', `${JSON.stringify(record)}\n`);
}
}
} finally {
await context.close();
await browser.close();
}
Run the login flow, then the capture job:
APP_EMAIL='you@example.com' APP_PASSWORD='read-from-a-secret-manager' node login.mjs
node capture.mjs
FULL_PAGE=true node capture.mjs
The example records a failure and proceeds to the next URL. Treat a final URL that points to a login page, an access-denied page, or a challenge as a failed capture even if navigation itself returned successfully. Add an application-specific check for that state rather than silently accepting an image of the wrong page.
3. Choose capture and waiting options
| Choice | Use it when | Trade-off |
|---|---|---|
| Viewport screenshot | You need a consistent at-a-glance view or fixed-size comparison. | Content below the viewport is omitted. |
| Full-page screenshot | You need the entire document in one image. | Very tall images can be cumbersome; lazy content may need to be triggered or awaited first. |
domcontentloaded plus a readiness selector |
The app renders after document parsing, especially in single-page apps. | You must choose a signal that represents the content you actually need. |
| Network-idle style waiting | The app becomes quiet after its requests complete. | Long polling, analytics, or streaming can keep a page busy; an app-specific signal is often more reliable. |
| Sequential iteration | You need predictable load on a shared account. | Batch duration grows with the number of pages. |
| Parallel pages or workers | You have confirmed the app and account allow concurrent reads. | Can trigger rate limits, session conflicts, or resource pressure. |
Playwright’s [screenshot guide](https://playwright.dev/docs/screenshots) documents viewport and full-page capture. The [Page API](https://playwright.dev/docs/api/class-page) documents navigation, waiting, and screenshot options. Disable animations when captures need a stable frame. For visual regression, keep browser version, operating system, viewport, device scale, and runtime conditions consistent; Playwright notes that rendering can vary across environments in its [visual comparisons guide](https://playwright.dev/docs/test-snapshots).
4. Python capture variant
Playwright’s Python package can reuse the same kind of saved browser state. Install the package and browser, then run this script after creating playwright/.auth/state.json through an appropriate login flow. It expects urls.txt and uses the same replaceable readiness selector.
python -m pip install playwright
python -m playwright install chromium
import asyncio
import json
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright
STATE = Path('playwright/.auth/state.json')
OUT = Path('screenshots')
READY_SELECTOR = '[data-testid="page-content"]' # Replace for your app.
FULL_PAGE = False
TIMEOUT_MS = 30_000
def filename(url: str, index: int) -> str:
path = urlparse(url).path.strip('/') or 'home'
slug = re.sub(r'[^a-zA-Z0-9_-]', '-', path.replace('/', '-'))[:70]
return f'{index + 1:03d}-{slug}.png'
async def main():
OUT.mkdir(parents=True, exist_ok=True)
urls = [line.strip() for line in Path('urls.txt').read_text().splitlines() if line.strip()]
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
storage_state=str(STATE),
viewport={'width': 1440, 'height': 1000},
device_scale_factor=1,
)
try:
with Path('capture-results.jsonl').open('a', encoding='utf-8') as log:
for index, url in enumerate(urls):
page = await context.new_page()
started = datetime.now(timezone.utc).isoformat()
output = OUT / filename(url, index)
try:
response = await page.goto(url, wait_until='domcontentloaded', timeout=TIMEOUT_MS)
if response is not None and response.status >= 400:
raise RuntimeError(f'Navigation returned HTTP {response.status}')
await page.locator(READY_SELECTOR).wait_for(state='visible', timeout=TIMEOUT_MS)
await page.screenshot(path=str(output), full_page=FULL_PAGE, animations='disabled')
record = {'url': url, 'output': str(output), 'startedAt': started,
'finishedAt': datetime.now(timezone.utc).isoformat(), 'ok': True}
except Exception as exc:
record = {'url': url, 'startedAt': started,
'finishedAt': datetime.now(timezone.utc).isoformat(), 'ok': False,
'error': str(exc), 'finalUrl': page.url}
print(f"Failed {url}: {exc}")
finally:
await page.close()
log.write(json.dumps(record) + '\n')
log.flush()
finally:
await context.close()
await browser.close()
asyncio.run(main())
5. cURL, Node.js fetch, and Python requests for public pages
A raw HTTP client is not a replacement for a browser session when the target requires a JavaScript-driven login, but these runnable examples show the simple unauthenticated request pattern for a public page. For authenticated browser pages, use the Playwright workflow above so application login and rendering happen in a browser.
cURL
curl -L --fail --output page.html https://example.com/
Python
import requests
response = requests.get('https://example.com/', timeout=30)
response.raise_for_status()
with open('page.html', 'wb') as f:
f.write(response.content)
Node.js
const response = await fetch('https://example.com/');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
await Bun.write('page.html', await response.text());
These examples fetch HTML; they do not render a screenshot. If your task is a screenshot, use the browser automation code or a screenshot API.
6. Authentication edge cases and safeguards
- Session storage: Playwright does not provide built-in persistence for
sessionStorage. If the app stores its login there, use the documented save-and-restore technique only after reviewing how it handles origin scope and secrets, or use an app-supported authentication flow. Do not assume cookies and local storage cover every app. - Expiry, device binding, and network rules: A saved state may expire or be rejected when reused from a different host, IP, browser, or device context. Refresh it through the normal login flow; do not repeatedly retry a rejected session as if it were a page-load problem.
- Multi-factor authentication and passkeys: Complete supported interactive login during state creation. Confirm the resulting state works in the same environment as the capture job.
- Redirects to login: Authentication can fail while navigation succeeds. Check the final URL and a page-specific authenticated marker before saving the image.
- State mutation: If pages alter server-side data, use an account and workflow designed for those changes. Parallel runs may interfere even if screenshots are the goal.
- Secrets in logs and artifacts: Never print cookie, authorization, or storage-state contents. Limit access to state files and remove temporary copies according to your credential policy.
- Filename collisions: Distinct URLs can share a path. Include a stable index or a hash of the complete URL if you need unique names across reordered or repeated input lists; avoid putting sensitive query values into filenames.
7. Reliability, performance, and cost
A single browser and sequential pages keep the workflow simple and reduce pressure on the target service, though elapsed time increases with batch size. For a large list, group work into manageable batches, preserve per-URL results, and retry only transient failures with a small bounded retry count. Do not retry authentication failures indefinitely. Before increasing concurrency, confirm account and service rules and monitor for rate limiting, session revocation, and browser resource limits.
Use the narrowest readiness condition that accurately signals usable content. Fixed delays add time when a page is fast and can still be too short when it is slow. A full-page capture can consume more memory and storage than a viewport image, especially for unusually long pages. Choose the smallest capture that answers the review question, and archive outputs with access controls appropriate to the page contents.
Playwright is browser automation software; this workflow’s direct costs depend on the compute and storage environment where you run it and the target application’s policies. The source documentation provides no universal runtime or price benchmark, so measure the job in your own environment instead of assuming a fixed throughput or cost. For visual comparisons, consistent browser and host conditions make diffs more meaningful.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Capture shows the login page | State expired, wrong origin, missing session storage, or login flow did not finish. | Refresh state through the normal login flow; verify the final URL and authenticated marker before capture. |
| Timeout waiting for readiness selector | Selector is wrong, content is delayed, or the page failed to load its data. | Inspect the page and choose a stable app-specific selector; check loading and error states, then tune the timeout to observed needs. |
| Navigation returns 401 or 403 | Account lacks access, state is invalid, or policy blocks the request. | Check account permissions and session validity. Do not treat access denial as a successful screenshot. |
| Page loads but screenshot is blank or incomplete | Capture began before content rendered, lazy images were not triggered, or the app uses a loading overlay. | Wait for the content marker and relevant loading indicator. Scroll or use an app-specific method to trigger lazy content before full-page capture. |
| Screenshot differs between runs | Animations, dynamic data, viewport, browser build, or host environment changed. | Fix viewport and device scale, disable animations, wait for stable content, and use a consistent runtime for comparisons. |
| Too many requests or session interruptions | Parallelism exceeds service or account limits. | Return to sequential capture or reduce worker count; consult the application’s concurrency and rate-limit rules. |
| Image output overwrites another page | Names were derived from paths that are not unique. | Include an index or stable URL hash in the filename and keep query secrets out of names. |
Or skip the browser setup
For pages the ScreenshotNeo API can reach, one GET request returns a screenshot. See the [ScreenshotNeo documentation](https://screenshotneo.com/docs/) for request options. For example:
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say which page verdict and billing status applied. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server made by Yorker Media. The API examples above illustrate a public URL request; do not assume a browser state file can be supplied for an authenticated application. Review the docs for supported authentication and options before moving a protected-page workflow.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Can I reuse one state file for every URL?
Yes, when the pages share the same origin and session and the application accepts that state. Confirm access for each page and refresh state when the session expires.
Should I capture pages in parallel?
Start sequentially. Raise concurrency only after checking the target application’s rate limits, account rules, and session behavior.
Does Playwright save sessionStorage?
No built-in storage-state persistence is provided for session storage. It needs separate, carefully reviewed handling if your app depends on it.
Why not just wait a fixed number of seconds?
A delay cannot tell whether the page is ready: it wastes time on fast pages and may still finish too early on slow ones. Wait for an app-specific signal instead.


