ScreenshotNeo

BlogHow-to

How to Capture Bulk Screenshots of Websites with Basic Authentication

Use Playwright to capture authorized Basic-auth pages in bulk, with scoped credentials, repeatable filenames, error handling, and screenshot options.

By the ScreenshotNeo team4 October 20269 min read

To capture bulk screenshots of websites with Basic authentication, configure credentials on a Playwright browser context, visit each authorized URL, and save a screenshot for each page. Reuse a context for URLs that share the same credential and authentication origin; set the credential origin narrowly, use HTTPS, and supply secrets at runtime rather than committing them to source code.

This guide covers HTTP Basic authentication only. A form-based sign-in, SSO flow, or application session needs its own authentication workflow. HTTP Basic encodes the user name and password with Base64; Base64 does not encrypt them. RFC 7617 says Basic authentication is not secure without an external secure system such as TLS. RFC 7617

1. Set up a URL list and the capture environment

Use a newline-delimited file so the input is easy to review and edit. These example URLs are placeholders: use only pages you are authorized to access.

https://example.com/reports/monthly
https://example.com/reports/weekly
https://example.com/dashboard

Install Playwright and its Chromium browser for the JavaScript example below:

npm init -y
npm install playwright
npx playwright install chromium

Set credentials in the process environment using your shell, a secrets manager, or your deployment platform. Do not add real credentials to the script, URL list, source control, logs, or command history.

export SITE_USER='your-user'
export SITE_PASSWORD='your-password'

The authentication origin includes the scheme, hostname, and port. For example, https://example.com is a different origin from http://example.com and https://example.com:8443. Use the exact origin that serves the protected pages.

2. Capture the URL list with Playwright in Node.js

Save this as capture.mjs. It reads urls.txt, creates the output directory, reuses one authenticated browser context, writes a numbered PNG for each successful URL, reports failures per URL, and closes the browser even if something goes wrong.

import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';

const username = process.env.SITE_USER;
const password = process.env.SITE_PASSWORD;
const origin = process.env.SITE_ORIGIN ?? 'https://example.com';

if (!username || !password) {
  throw new Error('Set SITE_USER and SITE_PASSWORD in the environment.');
}

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line.length > 0 && !line.startsWith('#'));

await mkdir('screenshots', { recursive: true });
const browser = await chromium.launch();
const context = await browser.newContext({
  httpCredentials: { username, password, origin },
  viewport: { width: 1440, height: 1000 },
});

try {
  for (const [index, url] of urls.entries()) {
    const filename = `screenshots/${String(index + 1).padStart(4, '0')}.png`;
    const page = await context.newPage();
    try {
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: 30_000,
      });
      if (!response) {
        throw new Error('Navigation returned no main-resource response.');
      }
      if (response.status() >= 400) {
        throw new Error(`Main document returned HTTP ${response.status()}.`);
      }

      // Replace or supplement this with an application-specific readiness signal.
      await page.screenshot({ path: filename, fullPage: true });
      console.log(`OK ${url} -> ${filename}`);
    } catch (error) {
      console.error(`FAILED ${url}: ${error.message}`);
    } finally {
      await page.close();
    }
  }
} finally {
  await context.close();
  await browser.close();
}

Run it with SITE_ORIGIN set if the pages use a different origin:

SITE_ORIGIN='https://example.com' node capture.mjs

The script intentionally handles URLs sequentially. That limits simultaneous load and makes failures easy to associate with their input. For larger jobs, add a small bounded worker pool only after confirming the site can handle parallel requests; concurrency is a workload choice, not a Playwright throughput guarantee.

3. Capture pages with Python Playwright

If the surrounding workflow is Python, the same context-level credential configuration works with Playwright’s Python API. Install the package and browser first:

python -m pip install playwright
python -m playwright install chromium

Save the following as capture.py and set SITE_USER, SITE_PASSWORD, and optionally SITE_ORIGIN in the environment. It uses the same newline-delimited urls.txt input and sequential handling.

import os
from pathlib import Path
from playwright.sync_api import sync_playwright

username = os.environ.get("SITE_USER")
password = os.environ.get("SITE_PASSWORD")
origin = os.environ.get("SITE_ORIGIN", "https://example.com")
if not username or not password:
    raise RuntimeError("Set SITE_USER and SITE_PASSWORD in the environment.")

urls = [
    line.strip()
    for line in Path("urls.txt").read_text(encoding="utf-8").splitlines()
    if line.strip() and not line.lstrip().startswith("#")
]
Path("screenshots").mkdir(parents=True, exist_ok=True)

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(
        http_credentials={"username": username, "password": password, "origin": origin},
        viewport={"width": 1440, "height": 1000},
    )
    try:
        for index, url in enumerate(urls, start=1):
            filename = Path("screenshots") / f"{index:04d}.png"
            page = context.new_page()
            try:
                response = page.goto(url, wait_until="domcontentloaded", timeout=30_000)
                if response is None:
                    raise RuntimeError("Navigation returned no main-resource response.")
                if response.status >= 400:
                    raise RuntimeError(f"Main document returned HTTP {response.status}.")
                page.screenshot(path=str(filename), full_page=True)
                print(f"OK {url} -> {filename}")
            except Exception as exc:
                print(f"FAILED {url}: {exc}")
            finally:
                page.close()
    finally:
        context.close()
        browser.close()

4. Choose the capture extent and page readiness

Viewport or full page

Mode Playwright option Use when
Viewport Omit fullPage, or set it to false You need the visible area at a fixed viewport size.
Full page fullPage: true in Node.js; full_page=True in Python You need the full scrollable document in one image.

Set the viewport in the browser context to keep captures consistent. A full-page capture can become very tall for long documents. If the deliverable needs consistent page-sized panels, capture the viewport or use a PDF workflow instead.

Wait for the content you actually need

domcontentloaded means the document has been parsed; it does not guarantee that client-rendered data, images, fonts, or other asynchronous content is ready. Avoid assuming one fixed delay fits every site. Prefer an application-specific signal, such as waiting for a report heading or a known content container:

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('[data-report-ready="true"]').waitFor({ state: 'visible', timeout: 15_000 });
await page.screenshot({ path: filename, fullPage: true });

The selector above is an example; replace it with a real readiness marker from the authorized site. If a page has no reliable marker, choose a suitable navigation state and a measured delay for that application, then record that choice so repeat runs use the same rule.

5. Handle credential scope, URL groups, and secrets

  • One origin and one credential pair: use one context with a specific origin for pages under that origin.
  • Several origins or accounts: create separate contexts per credential/origin group. This keeps session state and credential use easier to reason about.
  • Do not rely on URL paths for credential scope: the origin is scheme, host, and port; it does not narrow credentials to a particular path.
  • Redirects: if a protected page redirects to another host, do not assume the original origin-scoped credentials will authenticate that destination. Configure only credentials for an origin you are authorized to access.
  • Secrets: protect environment values and CI logs. Close the context and browser when finished, and protect the screenshots too: they may contain confidential information.

Playwright documents httpCredentials on browser contexts, including an optional origin restriction. The browser context option is for browser requests; do not confuse it with the separate API request context option. Playwright Browser API

6. cURL for a single Basic-authenticated page

cURL is useful for checking whether an endpoint accepts Basic credentials or downloading a response body. It does not render a webpage or produce a browser screenshot, so use Playwright for the actual screenshot capture.

curl --fail --show-error --silent \
  --user "$SITE_USER:$SITE_PASSWORD" \
  'https://example.com/reports/monthly' \
  --output response.html

Use HTTPS. Avoid placing literal secrets in the command; shell history and process inspection can expose command-line arguments depending on the environment. If the password includes characters that complicate the user:password form, use a protected cURL configuration or another secret injection mechanism rather than constructing an unsafe command string.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its screenshot endpoint is a single GET request; see the ScreenshotNeo API documentation. The one-call path is useful when you do not want to install and maintain a browser locally. This API example does not establish compatibility with a Basic-auth-only origin, so verify that your page is reachable through the API’s supported authentication configuration before using it for protected pages.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/reports/monthly \
  -o shot.webp

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the capture. Bot checks, blank pages, and failed loads are never billed; responses include page-verdict and billing headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.

8. Reliability, performance, and cost

  • Browser resources: a single browser process and one page at a time is a conservative starting point. Reusing the context avoids recreating the browser session for every URL. Increase concurrency only with a bounded limit and after considering site load and local memory.
  • Timeouts: navigation and readiness waits need finite timeouts. A timeout should mark that URL as failed; it should not silently create a screenshot that looks like a successful capture.
  • Repeatability: keep browser version, viewport, capture mode, readiness rule, and output naming consistent. Rendering can vary with browser and environment, so identical URLs do not guarantee pixel-identical results across different setups.
  • Partial success: log a status for each input URL and preserve the mapping from URL to filename. Retry transient failures separately with a limit rather than rerunning successful pages blindly.
  • Cost: a local Playwright workflow has no per-screenshot API charge, but it uses compute, storage, and maintenance time. For a hosted API, check current plan limits and whether unsuccessful captures are billed. ScreenshotNeo states that only clean shots are billed and lists its plans on its site.

9. Troubleshooting

Symptom Likely cause Fix
Browser keeps showing an authentication prompt or returns 401 Wrong username/password, wrong origin, or the endpoint is not using HTTP Basic. Confirm the credential pair and exact scheme, host, and port. Verify the site’s authentication type. Form login and SSO need a different flow.
Some pages authenticate while others fail The URLs span more than one origin, or a redirect reaches another host. Group URLs by origin and configure an authorized credential pair per context. Inspect redirect destinations before changing scope.
Screenshot is an error page The page returned an HTTP error, the credentials were rejected, or the server rendered an application error. Check the recorded response status and page content before capture. Keep the per-URL failure record; do not treat an error page as a successful screenshot.
Screenshot is blank or missing dynamic content Capture happened before client-side content became ready, or the site loads content lazily. Wait for a meaningful selector or application readiness signal. Use full-page capture where appropriate; inspect whether scrolling or interaction is required to load content.
Navigation times out The server is slow, the page waits on long-lived requests, or the timeout is too short for this workload. Use a navigation state suited to the site, set a reasoned finite timeout, and wait separately for the content needed in the screenshot.
Output file is overwritten Filenames are derived from a non-unique label or URL fragment. Use a stable unique index or sanitized URL plus a collision-safe suffix; retain a manifest that maps each output to its source URL.
Credentials appear in logs or repository history Secrets were embedded in code or passed unsafely on a command line. Remove them from source and logs, rotate exposed credentials, and inject secrets through protected runtime settings.
Captures differ between runs Different viewport, browser version, rendering environment, page data, or readiness timing. Pin the capture environment where feasible, standardize viewport and wait conditions, and distinguish expected dynamic content from capture failures.

10. FAQ

Can Playwright handle a login form with httpCredentials?

No. That option is for HTTP authentication credentials. A form-based login requires browser interactions or a separately established authenticated state.

Should I create a new context for every URL?

Usually not when the URLs share an origin and credentials. Reuse a context for that group; use separate contexts when credential or origin boundaries differ.

Does full-page mode create multiple screenshots?

No. It captures the full scrollable page as one tall image. Use viewport captures when you need fixed-size images.

Can I safely use Basic authentication over plain HTTP?

No. Basic credentials are Base64 encoded, not encrypted. Use HTTPS to protect them in transit, as described in RFC 7617.

Primary references