ScreenshotNeo

BlogHow-to

How to Capture Bulk Screenshots of URLs and Email a Report

Capture a list of URLs with Playwright, build an email report that shows each result, and handle failures, large batches, and delivery limits.

By the ScreenshotNeo team4 October 20269 min read

To capture screenshots of many URLs and email a report, read and validate the URL list, visit each page with a browser automation library, save one image per URL, record successes and failures, then send an HTML email with the images attached. The Node.js example below uses Playwright for capture and Nodemailer for delivery. It processes URLs sequentially to keep browser and target-site load predictable.

For a small report, attaching the images is convenient. For a large batch, email a link to an authorized file-storage location instead: mail providers impose different message-size limits, and those limits are not established here.

1. Prepare the input and install dependencies

Create a project and install Playwright and Nodemailer. Install the browser binary Playwright uses as well:

mkdir url-report
cd url-report
npm init -y
npm install playwright nodemailer
npx playwright install chromium

Save one HTTP or HTTPS URL per line in urls.txt:

https://example.com
https://www.wikipedia.org
https://playwright.dev

Run the script in an environment that can reach both the target sites and your configured SMTP server. Keep SMTP credentials in environment variables or a secrets manager, not in the script or source control.

2. Capture URLs and email the report with Node.js

Save this as report.mjs. It writes full-page PNGs to captures/, includes each URL and outcome in the email, and attaches successful images inline. A failed navigation stays visible in the report. The script sends the email even if some pages fail; it exits with a nonzero status after sending when any capture failed.

import { readFile, mkdir } from 'node:fs/promises';
import path from 'node:path';
import { chromium } from 'playwright';
import nodemailer from 'nodemailer';

const inputPath = process.env.URLS_FILE || 'urls.txt';
const outputDir = process.env.OUTPUT_DIR || 'captures';
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS || 45000);

const required = ['SMTP_HOST', 'SMTP_PORT', 'SMTP_USER', 'SMTP_PASS', 'MAIL_FROM', 'MAIL_TO'];
const missing = required.filter((key) => !process.env[key]);
if (missing.length) {
  throw new Error(`Missing required environment variables: ${missing.join(', ')}`);
}
if (!Number.isFinite(timeoutMs) || timeoutMs <= 0) {
  throw new Error('NAVIGATION_TIMEOUT_MS must be a positive number');
}

const urls = (await readFile(inputPath, 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line && !line.startsWith('#'));

const entries = urls.map((originalUrl, index) => {
  let parsed;
  try {
    parsed = new URL(originalUrl);
  } catch {
    return { index, originalUrl, error: 'Invalid URL' };
  }
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    return { index, originalUrl, error: 'Only HTTP and HTTPS URLs are supported' };
  }
  return { index, originalUrl, parsedUrl: parsed.href };
});

if (entries.length === 0) throw new Error(`No URLs found in ${inputPath}`);
await mkdir(outputDir, { recursive: true });

const browser = await chromium.launch({ headless: true });
const results = [];
try {
  for (const entry of entries) {
    if (entry.error) {
      results.push({ ...entry, status: 'Failed' });
      continue;
    }
    const filePath = path.resolve(outputDir, `page-${String(entry.index + 1).padStart(4, '0')}.png`);
    const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
    try {
      const response = await page.goto(entry.parsedUrl, {
        waitUntil: 'domcontentloaded',
        timeout: timeoutMs,
      });
      if (response && response.status() >= 400) {
        throw new Error(`Navigation returned HTTP ${response.status()}`);
      }
      await page.screenshot({ path: filePath, fullPage: true, type: 'png' });
      results.push({ ...entry, filePath, status: 'Captured' });
    } catch (error) {
      results.push({ ...entry, status: 'Failed', error: error.message });
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

function escapeHtml(value) {
  return String(value).replace(/[&<>"']/g, (char) => ({
    '&': '&amp;', '<': '&lt;', '>': '&gt;',
    '"': '&quot;', "'": '&#39;'
  })[char]);
}

const attachments = [];
const rows = results.map((result) => {
  const url = escapeHtml(result.originalUrl);
  if (result.status !== 'Captured') {
    return `<tr><td>${url}</td><td>Failed: ${escapeHtml(result.error || 'Capture failed')}</td></tr>`;
  }
  const cid = `capture-${result.index}@url-report`;
  attachments.push({ filename: path.basename(result.filePath), path: result.filePath, cid });
  return `<tr><td>${url}</td><td>Captured<br><img src="cid:${cid}" alt="Screenshot of ${url}" style="max-width:100%;height:auto"></td></tr>`;
}).join('\n');

const captured = results.filter((result) => result.status === 'Captured').length;
const failed = results.length - captured;
const port = Number(process.env.SMTP_PORT);
if (!Number.isInteger(port) || port <= 0) throw new Error('SMTP_PORT must be a positive integer');
const transporter = nodemailer.createTransport({
  host: process.env.SMTP_HOST,
  port,
  secure: process.env.SMTP_SECURE === 'true',
  auth: { user: process.env.SMTP_USER, pass: process.env.SMTP_PASS },
});

await transporter.sendMail({
  from: process.env.MAIL_FROM,
  to: process.env.MAIL_TO,
  subject: `URL screenshot report: ${captured} captured, ${failed} failed`,
  html: `<h1>URL screenshot report</h1><p>${captured} captured; ${failed} failed.</p><table border="1" cellpadding="6" cellspacing="0"><thead><tr><th>URL</th><th>Result</th></tr></thead><tbody>${rows}</tbody></table>`,
  attachments,
});

console.log(`Report sent: ${captured} captured, ${failed} failed`);
if (failed > 0) process.exitCode = 1;

Set the environment values for your SMTP provider, then run node report.mjs. The sample expects a host, port, username, password, sender, and recipient. Set SMTP_SECURE=true when your provider requires a direct TLS connection; otherwise leave it unset and follow the provider’s SMTP instructions. Some providers require an application password or other account-specific setup.

export SMTP_HOST='smtp.example.net'
export SMTP_PORT='587'
export SMTP_USER='reporter@example.net'
export SMTP_PASS='replace-with-secret'
export MAIL_FROM='reporter@example.net'
export MAIL_TO='recipient@example.net'
node report.mjs

3. Choose capture and report settings

Decision Default in the example When to change it
Viewport or full page Full-page PNG Use viewport captures for a consistent above-the-fold review. Full pages can be very tall and produce large attachments.
Wait condition domcontentloaded Choose a different wait condition or wait for a known selector when the page needs more time to render. networkidle can be unsuitable for pages with ongoing requests.
Viewport dimensions 1440 × 1000 Set dimensions to match the intended desktop or mobile review. Use a browser context with a device preset when device emulation is needed.
Image format PNG Playwright supports screenshot options such as JPEG and image buffers. JPEG can reduce size for photographic pages, with a quality trade-off.
Report delivery HTML email with inline image attachments Use a PDF or ZIP plus index when recipients need a portable report. For large payloads, upload to authorized storage and email a link.

Playwright documents screenshots to a file, full-page capture, image buffers, and locator-level screenshots in its screenshot guide and Page API. To capture one element instead of the whole page, use a locator such as await page.locator('main').screenshot({ path: filePath }) after confirming that selector exists.

4. Scale the batch safely

  • Validate before launching browsers. Reject malformed entries and non-HTTP schemes. If the URL list comes from untrusted users, add an allowlist or network controls to prevent requests to internal services.
  • Start sequentially. The sample uses one page at a time. If you add concurrency, limit it deliberately; browser memory, target-site load, rate limits, and email size all affect a sensible limit. There is no universal batch size or throughput guarantee.
  • Set a timeout and record each result. Keep the original URL, capture status, and error in the report. For scheduled runs, also write a persistent manifest with a timestamp and file path.
  • Retry selectively. A transient timeout may justify one retry, but avoid retrying every HTTP error or repeatedly hitting a site that is blocking the request. Keep attempts bounded and report the final outcome.
  • Plan for idempotency. Use a run identifier or clear the output directory before a run so an old image cannot be mistaken for a new capture. Avoid sending duplicate reports if a scheduler restarts a job after delivery.
  • Protect data. Screenshots can contain private page content. Restrict access to output files, recipients, and SMTP credentials, and clean up captures according to your retention needs.

The code uses file paths for Nodemailer attachments. Nodemailer also supports streams, which can avoid loading large files entirely into memory. See its documentation for attachment options, SMTP transport, and the Nodemailer overview.

5. Troubleshoot common failures

Symptom Likely cause What to do
net::ERR_NAME_NOT_RESOLVED or connection failure DNS, network access, proxy, or a mistyped URL Check the URL and connectivity from the machine running the script. Configure the required proxy or network access for that environment.
Navigation timeout The page is slow, blocked, or waiting on resources beyond the chosen condition Increase NAVIGATION_TIMEOUT_MS within a reasonable bound, use an appropriate wait condition, or wait for a page-specific selector. Keep the failure in the report if it still times out.
Screenshot is blank or incomplete Content renders after DOM readiness, requires scrolling, or depends on authentication or client-side scripts Wait for the relevant selector or a short page-specific delay, load authorized cookies or authentication state, and confirm the content is available in the browser context.
HTTP error is reported as failed The example treats a final response status of 400 or greater as a failed capture Inspect the response and target site. If your workflow intentionally captures error pages, change the status policy and label those captures accurately.
SMTP connection or authentication error Incorrect host, port, TLS mode, credentials, or provider policy Compare the settings with the mail provider’s current SMTP documentation. Check whether an application password or approved sender is required.
Email is rejected or too large Message size, attachment policy, recipient restrictions, or sending limits Reduce image dimensions or use fewer attachments, or store the report in an authorized location and email a link. Check the specific provider’s current limits.
Images do not display in the email The recipient’s mail client blocks inline images or strips CID references Include ordinary attachments or a link to the report as an alternative, and test with the recipient’s mail client.
Old images appear in a new report Output files or report inputs were reused across runs Use a fresh run-specific directory or remove previous captures before starting, and verify each manifest entry maps to the current run.

6. Cost and reliability considerations

A self-hosted Playwright workflow has no per-screenshot service charge in this design, but it uses your compute, browser installation, network, storage, and email service. Full-page images, concurrency, and retained history can increase resource use and message size. Email delivery and arbitrary-site capture are not guaranteed: target pages can change, block automation, require login, or fail to load. Validate the workflow against the actual URL set and monitor both capture outcomes and mail delivery.

Or skip the browser setup

ScreenshotNeo can return a screenshot from one GET request per URL, and its API also supports bulk capture of up to 100 URLs per call. For a bulk email report, you still assemble and send the report with your preferred email workflow. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; response headers say which outcome occurred. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for free and get 1,000 screenshots a month with no card.

FAQ

Can I run this on a schedule?

Yes. Run it from a scheduler in an environment with the browser, network access, and SMTP configuration available. Add logging and a run identifier so scheduled reports can be traced.

Can the report be a PDF?

Yes. Generate a PDF from the report HTML or use a PDF generation workflow, then attach that file. Account for the resulting size and preserve a per-URL status list.

Does this guarantee every URL will produce a screenshot?

No. Sites can be unavailable, require authentication, block automation, or render differently in the capture environment. Keep failures explicit and validate against the pages you need to monitor.