ScreenshotNeo

BlogHow-to

How to Save Bulk Website Screenshots to Azure Blob Storage

Capture an approved list of websites with Playwright and upload full-page screenshots to Azure Blob Storage using Entra ID, bounded concurrency, and a results manifest.

By the ScreenshotNeo team4 October 202613 min read

To save screenshots of multiple websites to Azure Blob Storage, read an approved URL list, capture each page with Playwright, and upload the returned image buffer with the Azure Storage JavaScript SDK. Authenticate with Microsoft Entra ID through DefaultAzureCredential, give the workload only the blob data permissions it needs, and use bounded concurrency plus unique blob names so one slow site or duplicate URL does not derail the batch.

This guide uses Node.js, Playwright, and Azure Blob Storage. Playwright can capture the viewport, a full scrollable page, or an element, and can return image bytes directly; the Azure SDK accepts buffers for block blob uploads. See the Playwright screenshot guide and Azure JavaScript upload guide.

1. Decide what the batch is allowed to capture

Start with a finite list of URLs that you own or are authorized to capture. Decide whether to process only that list or discover URLs from a sitemap or links; this example processes only the supplied list. Set a maximum count, a concurrency limit, and a timeout before running it. Respect the target sites’ access terms and controls. There is no universal crawl rate established by the APIs: choose a conservative policy for your targets and reduce concurrency if their responses indicate overload or blocking.

Choose the output before capture. A viewport screenshot is smaller and easier to compare consistently. A full-page capture includes the scrollable page and can consume substantially more memory for very tall pages. PNG is lossless and often larger; JPEG and WebP are options when smaller image files matter, subject to your required format and downstream compatibility. Playwright’s screenshot options include format, quality, clipping, scale, and full-page behavior; consult the Page screenshot API for the installed version.

2. Create the project and configure Azure access

  1. Install Node.js, create a project, and add Playwright and the Azure packages:
mkdir screenshot-batch
cd screenshot-batch
npm init -y
npm install playwright @azure/storage-blob @azure/identity
npx playwright install chromium
  1. Create a storage account and a private container, then grant the identity running the script the Storage Blob Data Contributor role scoped as narrowly as practical. The role assignment may take time to propagate.
  2. For local development, sign in using the Azure CLI with the identity that has that role: az login. In an Azure-hosted workload, configure its managed identity or other supported Entra credential. Set the account and container names as environment variables.

DefaultAzureCredential allows the SDK to use an available Entra credential chain. Microsoft recommends passwordless authentication for Blob Storage; account keys can provide broad access. See the JavaScript quickstart for setup and authorization guidance. Do not put account keys, SAS tokens, or other secrets in source control.

export AZURE_STORAGE_ACCOUNT="yourstorageaccount"
export AZURE_STORAGE_CONTAINER="website-shots"
# Sign in locally if using the Azure CLI credential:
az login

3. Prepare the URL list

Create urls.txt with one absolute HTTP or HTTPS URL per line. Blank lines and lines beginning with # are ignored. Keep this input under operator control: accepting arbitrary user-supplied URLs in a service can introduce server-side request forgery risks, including requests to internal addresses.

https://example.com/
https://www.microsoft.com/azure
# One URL per non-comment line

4. Run a bounded capture and upload batch

Save the following as capture.mjs. It launches one Chromium browser, reuses a browser context, and runs no more than the configured number of workers at once. Each worker creates and closes its own page. Screenshots are returned as buffers and uploaded with an image content type. The script writes a JSON Lines manifest with one success or failure record per input URL.

import { createHash } from 'node:crypto';
import { readFile, appendFile } from 'node:fs/promises';
import { chromium } from 'playwright';
import { DefaultAzureCredential } from '@azure/identity';
import { BlobServiceClient } from '@azure/storage-blob';

const account = process.env.AZURE_STORAGE_ACCOUNT;
const containerName = process.env.AZURE_STORAGE_CONTAINER;
if (!account || !containerName) {
  throw new Error('Set AZURE_STORAGE_ACCOUNT and AZURE_STORAGE_CONTAINER');
}

const concurrency = Number(process.env.CONCURRENCY ?? 3);
const timeoutMs = Number(process.env.PAGE_TIMEOUT_MS ?? 45000);
const fullPage = process.env.FULL_PAGE !== 'false';
const runId = new Date().toISOString().replaceAll(':', '-');
const manifestPath = `manifest-${runId}.jsonl`;
const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map(line => line.trim())
  .filter(line => line && !line.startsWith('#'));

if (!Number.isInteger(concurrency) || concurrency < 1 || concurrency > 16) {
  throw new Error('CONCURRENCY must be an integer from 1 to 16');
}

const parsedUrls = urls.map(value => {
  const url = new URL(value);
  if (!['http:', 'https:'].includes(url.protocol)) {
    throw new Error(`Unsupported URL protocol: ${value}`);
  }
  return url.href;
});

const service = new BlobServiceClient(
  `https://${account}.blob.core.windows.net`,
  new DefaultAzureCredential()
);
const container = service.getContainerClient(containerName);
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width: 1440, height: 1000 } });
let nextIndex = 0;
let failed = 0;

async function record(item) {
  await appendFile(manifestPath, `${JSON.stringify(item)}\n`);
}

async function worker() {
  while (true) {
    const index = nextIndex++;
    if (index >= parsedUrls.length) return;
    const url = parsedUrls[index];
    const page = await context.newPage();
    const capturedAt = new Date().toISOString();
    try {
      // domcontentloaded avoids waiting indefinitely for analytics and other background requests.
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: timeoutMs
      });
      if (!response) throw new Error('Navigation returned no main document response');
      if (!response.ok()) throw new Error(`Main document returned HTTP ${response.status()}`);

      const image = await page.screenshot({
        type: 'png',
        fullPage,
        animations: 'disabled',
        timeout: timeoutMs
      });
      const digest = createHash('sha256').update(url).digest('hex');
      // Run-specific names preserve separate captures and avoid concurrent writes to one blob.
      const blobName = `captures/${runId}/${digest}.png`;
      const blob = container.getBlockBlobClient(blobName);
      await blob.uploadData(image, {
        blobHTTPHeaders: { blobContentType: 'image/png' }
      });
      await record({ url, capturedAt, status: 'success', httpStatus: response.status(), blobName });
      console.log(`Saved ${url} -> ${blobName}`);
    } catch (error) {
      failed++;
      await record({
        url, capturedAt, status: 'failure',
        error: error instanceof Error ? error.message : String(error)
      });
      console.error(`Failed ${url}:`, error);
    } finally {
      await page.close();
    }
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, parsedUrls.length) }, worker));
} finally {
  await context.close();
  await browser.close();
}
console.log(`Processed ${parsedUrls.length}; failed ${failed}; manifest ${manifestPath}`);
if (failed) process.exitCode = 1;

Run it with node capture.mjs. Set FULL_PAGE=false for viewport screenshots, CONCURRENCY=2 to lower the worker count, or PAGE_TIMEOUT_MS=60000 to allow more time for slow pages. The example caps concurrency at 16 as a guardrail, not as a recommended throughput setting.

What the code does and what to change

  • Unique names: each URL is hashed rather than embedded in the blob path, avoiding raw query strings and credentials in object names. A run ID preserves separate runs. If you want one latest image per URL, define an explicit overwrite policy and prevent simultaneous workers from updating the same name.
  • Readiness: domcontentloaded waits for document parsing, not every image or API call. Some pages need a known selector via page.waitForSelector(), or a short explicit delay. Network-idle waits can hang on sites that keep connections open. Pick readiness per target rather than assuming one signal fits all sites.
  • Full-page: Playwright may need to expand or stitch a tall page; full-page output can be large and some pages use sticky elements, virtualization, or infinite scroll that do not represent one stable page. For infinite scroll, define a bounded scroll-and-load policy or capture only the viewport.
  • Response status: this sample records non-2xx main-document responses as failures. Remove that check if you deliberately want to archive error pages. Redirects are followed by navigation; the manifest retains the requested URL. Add the final page URL if redirect provenance matters.
  • Manifest: JSON Lines makes partial batches recoverable: each finished URL has a separate record. For durable production reporting, store the manifest in a durable location and consider writing records to a database or queue rather than one local file.
  • Content type: the sample uses PNG and image/png. If changing the format, update both the screenshot type, blob extension, and blobContentType. JPEG also supports a quality option; quality is not used for PNG.

5. Choose the right capture and storage options

Need Playwright / Azure choice Considerations
Entire scrollable page page.screenshot({ fullPage: true }) Can create very tall, memory-heavy images. Lazy content may need scrolling or site-specific readiness first.
Consistent screenshot dimensions Viewport screenshot and fixed newContext({ viewport }) Good for thumbnails and comparisons; content below the fold is omitted.
One component page.locator(selector).screenshot() Wait for a unique, visible element and handle missing selectors explicitly.
JPEG or WebP output type: 'jpeg' or 'webp', where supported Use matching file extensions and MIME types; check downstream consumers.
Higher pixel density Set context deviceScaleFactor Increases pixel dimensions and image size. Keep it consistent across the batch.
Image buffer upload BlockBlobClient.uploadData(buffer, options) Convenient for the screenshot buffer; memory use includes active page and image buffers.
Capture to disk first page.screenshot({ path }) then uploadFile(path) Useful for local inspection or resuming uploads, but requires temporary-file cleanup and disk capacity.
Large payloads or streams uploadStream() or SDK transfer options Tune buffer size and per-transfer concurrency against representative workloads. SDK examples are not universal performance recommendations.

Azure’s JavaScript SDK supports uploadData, uploadFile, and uploadStream. Set blob HTTP headers so downloaded objects have the right image content type. The SDK documentation states that its storage client libraries do not support concurrent writes to the same blob; use unique names or explicit conditional/overwrite behavior. See upload options and same-blob concurrency guidance.

6. Retries, idempotency, and recovery

The sample isolates errors per URL and continues the batch, then returns a nonzero process exit code if any capture failed. It does not automatically retry. Add retries only for errors that may recover, such as transient network failures or service throttling. Use a small retry limit with exponential backoff and jitter; do not repeatedly retry permanent navigation failures, access-denied responses, or invalid input. For Azure upload failures, retry the upload of the existing buffer when the error is transient, while preserving the same unique blob name for that run.

For restartable jobs, assign a stable run ID, record a URL’s status as soon as it finishes, and skip successful entries on resume. If deterministic naming is used, a retry may overwrite its own prior result; that is safe only if one worker owns that name and the overwrite behavior is intentional. For multiple independent runs, include a run identifier to prevent races and retain provenance.

7. When to use a managed service or Azure Playwright

Choose based on how URLs are supplied, how much browser control is needed, where results must land, and who will operate retries and retention.

Approach Good fit Tradeoffs and checks
ScreenshotNeo Need a screenshot API without operating a browser fleet; especially useful when clean captures and explicit billing outcomes matter. One GET request returns an image or PDF. It removes supported consent banners, newsletter popups, and chat widgets before capture; only clean shots are billed. Confirm how to move API results into your Blob workflow.
Playwright plus Azure Blob SDK Need control of browser context, URL input, blob naming, metadata, and capture settings. You operate browser lifecycle, queueing, retries, access policy, retention, and monitoring.
Managed bulk screenshot API Want a service to process URL lists, sitemaps, or domains and deliver images to cloud storage. AddScreenshots API documentation describes Azure Blob as a possible repository. Verify current features, terms, security, retention, region, URL acceptance, and price with the provider before adopting it.
Azure Playwright reporter Screenshots are artifacts of a Playwright end-to-end test workflow. This is a testing and report workflow, not a general-purpose domain crawler. Reporting has workspace, storage, RBAC, CORS, authentication, and version prerequisites.

Microsoft documents that the Azure Playwright reporter uploads test reports and related artifacts to workspace storage. Its setup requires reporting to be enabled and a storage account selected in workspace settings, Storage Blob Data Contributor access for runners, and Entra authentication; trace viewing also calls for a CORS rule allowing https://trace.playwright.dev with GET and OPTIONS. The documentation specifies Playwright 1.57 or later for reporting, so recheck the current reporter setup when implementing.

For visual regression comparisons, browser host operating systems can change screenshot snapshots. Microsoft recommends running comparisons in the service environment and configuring service-specific snapshot paths when required. See Azure Playwright visual comparison guidance.

Or skip the browser setup

ScreenshotNeo takes screenshots through one API request and can return PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing outcome. An MCP server gives AI agents tools to take screenshots, inspect page info, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Use the API to fetch an image, then upload the returned bytes to Blob Storage using the SDK pattern above. See the ScreenshotNeo API documentation for supported parameters and response behavior.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

Performance, reliability, and cost

  • Concurrency: more workers can improve batch completion time, but increase browser CPU and memory, network load, and pressure on target sites. Start low, observe failures and resource use, and tune for the actual URL set. Limit upload concurrency as well if storage requests or memory become a bottleneck.
  • Memory: each in-memory screenshot occupies space until its upload completes. Full-page captures and high device scale factors increase image size. Bound both active pages and queued work rather than launching one page per URL.
  • Readiness and timeouts: avoid global waits for every network request unless the target is known to settle. A navigation timeout is a per-URL outcome; preserve it in the manifest and decide whether to retry.
  • Storage cost: cost depends on the Azure account, region, stored image size, retention, and operations. The dossier does not establish a comparable price or a universal throughput figure; consult current Azure pricing for your account and workload. Apply a retention or lifecycle policy if captures are temporary.
  • Capture sensitivity: screenshots may contain personal, account, or other sensitive information visible to the browser. Use appropriate access controls and retention, avoid capturing authenticated pages unless authorized, and avoid placing secrets in URLs or blob names.
  • Visual repeatability: fonts, browser versions, host OS, animation, viewport, and page content can alter screenshots. Keep the environment and capture options stable when comparing images; the example disables animations but cannot freeze changing remote content.

Troubleshooting

Symptom Likely cause Fix
DefaultAzureCredential cannot get a token No supported credential is configured or local CLI is not signed in. Run az login locally, or configure the Azure-hosted workload’s identity. Confirm the account and tenant context.
Blob upload returns authorization failure The identity lacks blob data-plane permissions, role scope is wrong, or role assignment has not propagated. Assign Storage Blob Data Contributor at the appropriate scope and allow time for propagation. Management-plane access alone does not grant data writes.
Container not found Container name/environment variable is wrong or it has not been created. Check AZURE_STORAGE_CONTAINER and create the private container before running. This script intentionally does not create infrastructure.
Browser executable missing Playwright’s Chromium browser has not been installed for this environment. Run npx playwright install chromium; install operating-system dependencies as required by the Playwright installation documentation.
Navigation timeout or no response Slow host, blocked request, DNS/TLS issue, or page behavior that does not reach the selected readiness condition. Inspect the URL from the runner, choose an appropriate readiness event or selector, and adjust timeout for known slow targets. Record and selectively retry transient failures.
HTTP error page marked failed The sample treats a non-2xx main document as a failure. Remove or adapt the status check if error pages are intended archive targets.
Screenshot is blank or incomplete Capture ran before client rendering or lazy content; target blocks automation or requires interaction. Wait for a stable page-specific selector, add a deliberate bounded wait or authorized interaction, and inspect the response and rendered page. Some pages cannot be captured reliably.
Images or fonts are missing Resources have not loaded, are blocked, or load after the chosen readiness event. Wait for required image selectors or use an appropriate site-specific readiness check. Avoid assuming network idle is safe for every page.
Repeated runs overwrite images The blob naming scheme is deterministic and excludes a run ID. Include a run ID or establish a deliberate latest-version overwrite policy. Do not allow concurrent writes to the same blob without concurrency control.
Memory pressure or slow uploads Too many pages, very tall full-page images, or excess transfer concurrency. Lower worker count, capture viewport only, reduce device scale factor, or save files and upload sequentially. Tune SDK transfer settings using representative captures.
Azure Playwright visual comparisons fail between local and remote runs Browser host OS or rendering environment changed the expected snapshot. Run comparisons in the service environment and configure snapshot paths for the service OS as described in Microsoft’s visual comparison guidance.

FAQ

Can I save the screenshots directly to Azure without local files?

Yes. Playwright’s page.screenshot() returns a buffer when no path is supplied, and BlockBlobClient.uploadData() accepts that buffer. The example uses this in-memory flow.

Should I use a container per website?

Usually organize captures with blob name prefixes and a single container unless access boundaries or retention requirements call for separate containers. Blob prefixes are naming conventions rather than directories.

Does full-page mean every item on an infinite-scroll site?

No. It captures the page’s scrollable content at capture time; virtualized or infinite-scroll content may not have been rendered. Define and bound any scroll-and-load behavior you need.

Can Azure Playwright replace this batch script?

It fits when captures belong to Playwright tests and you want managed test execution and uploaded reports. For an arbitrary supplied URL list, the browser-plus-Blob pipeline or a purpose-built bulk capture service is a closer fit.