How to Take Screenshots of Many URLs with Custom Headers in Playwright
Capture a URL batch with shared or per-page headers in Playwright. Learn safe credential scoping, readiness checks, concurrency, error handling, and a one-call API option.
Set shared headers on a Playwright BrowserContext, then open a page for each URL, navigate, wait for the content you need, and save its screenshot. Context headers are sent with every request initiated by every page in that context, so use separate contexts when URLs require different credentials or trust boundaries. For a small sequential batch, one page at a time is simple and resource-conscious; use bounded concurrency only after considering target-site limits and local resources.
This guide uses JavaScript and Playwright’s Chromium browser. The same approach works with other Playwright browser types. See the official BrowserContext API, Page API, and Pages guide.
1. Install Playwright and prepare the output directory
In a new Node.js project, install Playwright and its Chromium browser. Create the output directory before running the script because screenshot paths do not create parent directories.
npm init -y
npm install playwright
npx playwright install chromium
mkdir -p screenshots
Save the runnable batch script below as capture.js, then run node capture.js. Provide headers using environment variables or another secret store in real deployments; do not commit tokens or cookies to source control.
2. Complete sequential batch example
const { chromium } = require('playwright');
const path = require('node:path');
const urls = [
'https://example.com',
'https://playwright.dev',
];
// Header values must be strings. Replace or load these safely for your use case.
const headers = {
'X-Custom-Header': 'example-value',
// Authorization: `Bearer ${process.env.API_TOKEN}`,
};
async function captureUrls(urlList, extraHeaders) {
const browser = await chromium.launch();
const context = await browser.newContext({
viewport: { width: 1440, height: 900 },
});
const results = [];
try {
await context.setExtraHTTPHeaders(extraHeaders);
for (const [index, url] of urlList.entries()) {
const page = await context.newPage();
try {
const response = await page.goto(url, {
waitUntil: 'load',
timeout: 30_000,
});
// HTTP 404/500 responses may still produce a response object.
const status = response ? response.status() : null;
if (status !== null && status >= 400) {
throw new Error(`HTTP ${status} for ${url}`);
}
// For dynamic sites, replace this with a meaningful page-specific signal:
// await page.locator('[data-report-ready="true"]').waitFor();
const outputPath = path.join(
'screenshots',
`page-${String(index + 1).padStart(3, '0')}.png`
);
await page.screenshot({
path: outputPath,
fullPage: true,
animations: 'disabled',
});
results.push({ url, outputPath, status, ok: true });
} catch (error) {
results.push({ url, ok: false, error: error.message });
console.error(`Capture failed for ${url}: ${error.message}`);
} finally {
await page.close();
}
}
} finally {
await context.close();
await browser.close();
}
return results;
}
captureUrls(urls, headers).then((results) => {
console.log(JSON.stringify(results, null, 2));
if (results.some((result) => !result.ok)) process.exitCode = 1;
}).catch((error) => {
console.error(error);
process.exitCode = 1;
});
The filenames here use the input order. If URLs are generated dynamically or batches can contain duplicate entries, consider adding a stable record ID or a sanitized URL hash to filenames. Avoid using raw URLs as filenames: they can contain characters unsupported by filesystems and may expose query-string secrets.
3. Choose where headers apply
Context-level headers for a shared batch
context.setExtraHTTPHeaders(headers) applies the supplied headers to requests initiated by every page in the context. This is convenient when the whole batch has the same header policy. It also means headers can accompany subresource requests, such as scripts and images, rather than only the top-level navigation.
Header values must be strings. The order of headers in outgoing requests is not guaranteed. If a page sets its own extra HTTP headers, its matching values take precedence over the corresponding context-level values. See the BrowserContext API.
Page-level headers for one-off overrides
const page = await context.newPage();
await page.setExtraHTTPHeaders({
'X-Custom-Header': 'page-specific-value',
});
await page.goto('https://example.com');
Use this when a page needs a different value for a matching header. Be deliberate: page-level headers are not a substitute for separating unrelated identities or credentials if pages share other context state.
Separate contexts for different credentials or policies
Group URLs that share an identity and header policy into one context. For URLs that must not receive the same authorization header, create separate contexts and set each group’s headers separately. A context is also where browser state and settings are scoped, so this boundary is easier to reason about than changing shared credentials repeatedly during a batch.
async function captureGroup(browser, urls, headers, label) {
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
try {
await context.setExtraHTTPHeaders(headers);
for (const [index, url] of urls.entries()) {
const page = await context.newPage();
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.screenshot({ path: `screenshots/${label}-${index + 1}.png` });
} finally {
await page.close();
}
}
} finally {
await context.close();
}
}
const browser = await chromium.launch();
try {
await captureGroup(browser, ['https://example.com/account'],
{ Authorization: `Bearer ${process.env.ACCOUNT_TOKEN}` }, 'account');
await captureGroup(browser, ['https://example.org/public'],
{ 'X-Custom-Header': 'public-batch' }, 'public');
} finally {
await browser.close();
}
This snippet uses top-level await; save it as an ES module (for example, capture.mjs) and define the chromium import and output directory as in the first example. The example is a usage pattern, not a claim of executed testing.
4. Decide when each screenshot is ready
page.goto() supports load, domcontentloaded, networkidle, and commit as navigation completion conditions; its documented default is load. A page can continue rendering content after navigation, so select a readiness signal that matches what the screenshot must contain.
load: wait for the load event. A reasonable starting point for ordinary static pages.domcontentloaded: proceed after HTML parsing and deferred scripts; may be faster where images and other resources are not required for readiness.commit: proceed once the response is received and document navigation begins. Use only when earlier capture is intentional.networkidle: Playwright discourages this as a testing readiness strategy. Persistent connections and background polling can make it unsuitable. Prefer a locator or application-specific signal.
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('main h1').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: outputPath, fullPage: true });
A locator is only useful if that element really indicates the required content is ready. For a report page, wait for the report container or a known completion marker; for a page with delayed images, wait for the specific image or application signal. Fixed delays can help with a known animation or delayed widget, but add latency and do not prove readiness. The official Page API describes navigation and screenshot behavior.
5. Screenshot options and batch choices
| Choice | Use it when | Trade-off |
|---|---|---|
fullPage: true |
You need the full scrollable document. | Tall pages produce large files and can expose more content than a viewport capture. |
Viewport capture (omit fullPage) |
You need a consistent above-the-fold image. | Content below the viewport is excluded. |
| One page at a time | The batch is modest or stability is the priority. | Total elapsed time grows with the number of URLs. |
| Bounded parallel pages | Throughput matters and resource limits are understood. | More browser memory, CPU, and simultaneous requests; target sites may throttle or block traffic. |
Set a consistent viewport and browser configuration if the batch is for visual comparison. Playwright notes that rendering can differ with operating system, browser version, hardware, headless mode, and other environmental factors; keep the capture environment consistent with the baseline. See Playwright’s visual snapshot guidance.
Do not start one page per URL without a concurrency limit for a large input set. A simple worker pool can cap active tasks, but there is no universally safe number: choose based on available memory, the sites’ rate limits, and observed failures. A low cap is a safer initial setting than unbounded concurrency.
6. Errors, partial failures, and retries
The sequential example records a failure and continues with the rest of the list. This is useful for batch jobs where one unavailable site should not discard successful captures. A production batch should also persist its input URL, outcome, status code, error message, and output path so failed items can be retried selectively.
- Navigation timeout: raise the timeout only when slow responses are expected; otherwise record the URL as failed. Check network access and whether the site is stalled.
- HTTP error page: a valid HTTP 404 or 500 response does not necessarily make
page.goto()throw. Inspect the returned response status if such captures should be rejected. - Challenge or access denied: the site may require a permitted authentication flow or may reject automation. Do not treat a screenshot of the challenge page as a successful content capture.
- Missing screenshot directory: create the parent folder before capture, or create it from Node.js with
fs.mkdir(path.dirname(outputPath), { recursive: true }). - Overwritten files: use unique names derived from a stable input identifier and batch ID.
- Bad header values: ensure every value is a string and names/values are valid HTTP header data. Avoid line breaks in values.
- Secret leakage: do not log authorization headers or full URLs if their query strings contain sensitive data. Keep credentials out of filenames and error telemetry.
Retries should be selective. Retry transient navigation failures with a small attempt limit and backoff, but do not repeatedly retry deterministic 4xx responses or an intentional access denial. Make output writes idempotent or use attempt-specific temporary files so a failed retry does not leave a partial artifact that looks complete.
7. Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| The server does not receive the header | Header was set after navigation, misspelled, or supplied at the wrong scope. | Set context headers before creating/navigating pages; verify the exact header name and value at a safe test endpoint. |
| A header appears on unrelated requests | Context headers apply to every request initiated by pages in that context. | Split URL groups into contexts with separate policies, or use a page-specific override where appropriate. |
| The screenshot is blank or incomplete | Capture happened before application content rendered, or the site returned a challenge/error page. | Check response status and wait for a meaningful visible locator or app-ready marker. |
| The script hangs on navigation | Slow resources, long-lived connections, or an unsuitable readiness condition. | Set a finite timeout; use a suitable navigation condition and a locator for actual content. Avoid relying on network idle for persistent applications. |
| A 404/500 screenshot was saved as success | Navigation completed with an HTTP error status. | Inspect response.status() and mark disallowed statuses as failures. |
| Captures differ between runs | Environment, viewport, browser version, dynamic content, fonts, or animations vary. | Keep the runtime, browser, viewport, and readiness rule consistent; disable animations where appropriate and account for changing content. |
| Some images or widgets are absent | They are lazy-loaded or appear after the readiness signal. | Wait for the particular image/widget or use an app signal; full-page capture alone does not guarantee every lazy resource has loaded. |
| Browser crashes or the machine runs out of memory | Too many pages or large full-page captures are active at once. | Reduce concurrency, close each page promptly, and capture only the needed viewport when suitable. |
8. Performance, reliability, and cost
Playwright’s software itself does not charge per screenshot; the practical costs are compute, storage, network traffic, and time spent maintaining browser infrastructure. Full-page images and high-resolution pages can increase memory and file size. Batch size, page weight, and concurrency determine the load placed on both your machine and destination sites.
For reliability, use finite navigation and locator timeouts, close every page/context/browser in finally blocks, record individual outcomes, and make retries limited and selective. Use separate contexts for distinct credential groups. Keep the browser version and operating environment stable when screenshots feed visual comparisons. No particular concurrency level or completion time is guaranteed by the cited documentation.
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Send a GET request with the target URL and receive a PNG, JPEG, WebP, or PDF. The example below captures a page as WebP; see the ScreenshotNeo API documentation for request options and formats.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, and cache hits are never billed, and response headers say which verdict and billing outcome applied. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.
10. FAQ
Can one BrowserContext contain several pages?
Yes. Pages in a context share its context-level settings, including extra headers. Close each page when its capture is finished.
Will context headers be sent only to the main page URL?
No. They apply to requests initiated by any page in that context, so consider subresources and keep credential scope intentional.
Does a successful navigation mean the page returned HTTP 200?
No. Check the navigation response status when HTTP success is a requirement; error statuses can still produce a response.
Does full-page mode guarantee lazy content is loaded?
No. Wait for the content your use case needs before capturing, especially when the page loads content as it scrolls or after application events.


