ScreenshotNeo

BlogHow-to

How to Capture Screenshots of URLs from a JSON File with Puppeteer

Read URLs from JSON, capture each page with Puppeteer, and save reliable screenshots with clear naming, readiness, and error handling.

By the ScreenshotNeo team4 October 20269 min read

Use Node.js to read and validate a JSON file, then use Puppeteer to open each URL and save a uniquely named screenshot. This guide assumes an input file shaped like {"urls":["https://example.com","https://example.org"]}. It processes URLs sequentially so failures are isolated and browser use stays predictable.

1. Install Puppeteer and prepare the input

Create a project and install Puppeteer, which downloads a compatible Chrome for Testing browser:

mkdir url-shots
cd url-shots
npm init -y
npm install puppeteer

Save this as urls.json:

{
  "urls": [
    "https://example.com",
    "https://example.org"
  ]
}

The schema is an explicit choice for this script, not a Puppeteer requirement. You can instead use a top-level array or objects with IDs, but update validation and naming to match.

2. Capture each URL with a runnable script

Save the following as capture.js and run node capture.js. It checks input before launching Chrome, creates the output directory, uses one browser, and closes it in a finally block. Individual navigation or screenshot errors are reported while later URLs continue.

const fs = require('node:fs/promises');
const path = require('node:path');
const puppeteer = require('puppeteer');

const INPUT_FILE = path.resolve('urls.json');
const OUTPUT_DIR = path.resolve('screenshots');
const NAVIGATION_TIMEOUT_MS = 30_000;
const WAIT_UNTIL = 'domcontentloaded'; // See readiness choices below.
const FULL_PAGE = true;
const VIEWPORT = { width: 1440, height: 900, deviceScaleFactor: 1 };

function validateUrls(value) {
  if (!value || typeof value !== 'object' || !Array.isArray(value.urls)) {
    throw new Error('Expected a JSON object with a "urls" array.');
  }
  const valid = [];
  const rejected = [];
  value.urls.forEach((entry, index) => {
    if (typeof entry !== 'string' || entry.trim() === '') {
      rejected.push({ index, reason: 'URL must be a non-empty string' });
      return;
    }
    let parsed;
    try {
      parsed = new URL(entry.trim());
    } catch {
      rejected.push({ index, reason: 'URL is malformed or missing a scheme' });
      return;
    }
    if (!['http:', 'https:'].includes(parsed.protocol)) {
      rejected.push({ index, reason: 'Only http and https URLs are supported' });
      return;
    }
    valid.push({ index, url: parsed.href, host: parsed.hostname });
  });
  return { valid, rejected };
}

async function main() {
  let raw;
  try {
    raw = await fs.readFile(INPUT_FILE, { encoding: 'utf8' });
  } catch (error) {
    throw new Error(`Cannot read ${INPUT_FILE}: ${error.message}`);
  }
  let parsed;
  try {
    parsed = JSON.parse(raw);
  } catch (error) {
    throw new Error(`Invalid JSON in ${INPUT_FILE}: ${error.message}`);
  }

  const { valid, rejected } = validateUrls(parsed);
  for (const item of rejected) {
    console.error(`Skipping urls[${item.index}]: ${item.reason}`);
  }
  if (valid.length === 0) {
    throw new Error('No valid http or https URLs to capture.');
  }

  await fs.mkdir(OUTPUT_DIR, { recursive: true });
  const browser = await puppeteer.launch({ headless: true });
  const results = [];
  try {
    const page = await browser.newPage();
    await page.setViewport(VIEWPORT);
    page.setDefaultNavigationTimeout(NAVIGATION_TIMEOUT_MS);

    for (const item of valid) {
      // Index-based names are unique even when URLs share a hostname.
      const output = path.join(
        OUTPUT_DIR,
        `${String(item.index + 1).padStart(4, '0')}-${item.host.replace(/[^a-z0-9.-]/gi, '_')}.png`
      );
      try {
        const response = await page.goto(item.url, { waitUntil: WAIT_UNTIL });
        if (response && response.status() >= 400) {
          console.warn(`${item.url} returned HTTP ${response.status()}; saving the rendered error page.`);
        }
        // Optional for apps that render important content after navigation:
        // await page.waitForSelector('[data-page-ready]', { timeout: 10_000 });
        await page.screenshot({ path: output, fullPage: FULL_PAGE });
        results.push({ url: item.url, output, ok: true });
        console.log(`Saved ${output}`);
      } catch (error) {
        results.push({ url: item.url, ok: false, error: error.message });
        console.error(`Failed ${item.url}: ${error.message}`);
      }
    }
  } finally {
    await browser.close();
  }

  const failures = results.filter(result => !result.ok);
  console.log(`Finished: ${results.length - failures.length} succeeded, ${failures.length} failed.`);
  if (failures.length) process.exitCode = 1;
}

main().catch(error => {
  console.error(error.message);
  process.exitCode = 1;
});

The script treats a non-success HTTP status as a rendered page worth saving and warns about it; change that policy if error pages should count as failures. It does not print full URLs in its normal success log beyond failures, which can matter if URLs contain tokens or private query parameters. Avoid putting credentials in input URLs where possible.

3. Choose when a page is ready

page.goto() can wait for a navigation lifecycle event. The choice affects speed and reliability; no single event proves that every website has finished rendering.

Option Useful when Tradeoff
domcontentloaded Most markup is present and you will wait for a specific element or a short app signal. Images, fonts, and client-rendered content may still be loading.
load You need the page load event and ordinary dependent resources. It does not guarantee delayed application data is rendered.
networkidle2 A mostly static page settles its network requests. Polling, analytics, streaming, or long requests can delay or prevent idleness.

Puppeteer’s screenshot guide demonstrates waitUntil: 'networkidle2', and its network idle API waits until the network is idle for at least the configured idle time. Treat that as a signal, not proof of visual completeness. For a client-rendered page, use a meaningful selector:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.waitForSelector('[data-page-ready]', { visible: true, timeout: 10_000 });
await page.screenshot({ path: output, fullPage: true });

Choose a selector that appears only when the content you need is ready. If you control the application, a dedicated readiness marker is more dependable than guessing with a fixed delay. A delay can help with an unavoidable animation or delayed widget, but adds the same wait to every URL.

4. Configure the screenshot

The example sets a 1440×900 viewport and captures the full page. Use viewport capture when you want the visible screen only; set FULL_PAGE = false. Full-page output can be much taller and larger. Puppeteer supports screenshot options including path, type, clipping, background transparency, and image quality. See the Page.screenshot API for current option details.

// Viewport screenshot:
await page.screenshot({ path: 'viewport.png' });

// Full page screenshot:
await page.screenshot({ path: 'whole-page.png', fullPage: true });

// A clipped region of the page:
await page.screenshot({
  path: 'region.png',
  clip: { x: 0, y: 0, width: 800, height: 600 }
});

// JPEG with quality (quality applies to JPEG):
await page.screenshot({ path: 'page.jpg', type: 'jpeg', quality: 80 });

// Transparent background where the page permits it:
await page.screenshot({ path: 'transparent.png', omitBackground: true });

The file extension can be used to infer the image type. Keep the output extension consistent with the chosen type. Set a consistent viewport and device scale factor when comparing captures. Pages can still vary because of content updates, fonts, animations, consent prompts, and remote assets.

5. Input, naming, and batch choices

  • Input validation: this sample skips blanks, malformed values, and unsupported schemes, then reports their array indexes. It stops if no usable URLs remain.
  • Output naming: a stable input index plus hostname avoids collisions and unsafe raw URL characters. For repeatable jobs, clear or version the output directory so old files are not mistaken for the latest run.
  • Failure policy: the sample continues after a URL fails and exits nonzero at the end. For all-or-nothing jobs, stop on the first failure and persist results before exiting.
  • Concurrency: sequential processing is simple and keeps browser resource use predictable. To increase throughput, use a bounded worker pool and measure memory and page stability on your target sites; there is no universal safe parallel page count.
  • Large JSON files: readFile buffers the complete input. This is appropriate for ordinary lists. For unusually large input, use a streaming parser or line-delimited JSON and process records incrementally.
  • Reuse: this example reuses one page sequentially. If pages leak state through cookies or local storage, create a fresh browser context or page per URL and close it after capture.

6. Browser setup options

The puppeteer package manages a compatible Chrome for Testing browser, which is the straightforward setup for a local script. If you use puppeteer-core, configure an installed browser explicitly:

npm install puppeteer-core
const puppeteer = require('puppeteer-core');
const browser = await puppeteer.launch({
  executablePath: '/path/to/chrome',
  headless: true
});

Alternatively configure a supported channel where available. Puppeteer says it works best with its downloaded Chrome for Testing and does not guarantee compatibility with arbitrary browser versions. On a server or CI runner, ensure the chosen browser executable and required system libraries are present.

7. cURL, Python, and ScreenshotNeo alternative

For a single request outside your Puppeteer script, cURL and Python can call a screenshot service. ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. Its one-call endpoint accepts a URL and returns an image or PDF. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

8. Node.js ScreenshotNeo call

If you want the same one-call option from Node.js, this uses the endpoint directly:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

9. Or skip the browser setup

With ScreenshotNeo, make one GET request instead of installing and managing a browser. Cookie banners are accepted like a visitor; 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture, and each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

10. Troubleshooting

Symptom Likely cause Fix
JSON parse error Trailing comma, comments, or malformed JSON. Validate the file as strict JSON and inspect the reported line and position.
Expected a urls array The file uses a different schema, such as a top-level array. Change the validator to match the actual schema or wrap values in {"urls": [...]}.
Navigation timeout The site remains active, is slow, or has long-running requests. Use a suitable lifecycle event, raise the timeout for that site, or wait for a specific selector instead of network idle.
Screenshot is blank or incomplete Capture occurred before client rendering, images, or fonts were ready. Wait for a page-specific selector or app signal; check that the target content is visible in the browser.
Browser executable missing puppeteer-core has no configured browser, or managed Chrome was not installed. Use puppeteer with its browser download, or set executablePath/channel correctly.
Permission denied writing output The output path is not writable. Choose a writable directory and ensure it exists before capture.
Files overwrite each other Names are derived from a non-unique host or path. Include the input index or a stable unique ID in every filename.
Process hangs after a failure Browser closure is skipped by error handling. Keep browser.close() in finally and avoid swallowing close errors silently.

11. Performance, reliability, and cost

One browser launch amortized across a sequential batch avoids the overhead of launching Chrome for every URL. Full-page screenshots and high device scale factors can increase image size and memory use. Bounded concurrency may improve throughput but also increases browser resource use and can make pages less stable; choose a limit by measuring on your own workload. Retries can help with transient network failures, but use a small cap and record failed URLs so persistent errors do not loop indefinitely.

Local Puppeteer has no per-screenshot API charge, but you provide the machine, browser installation, runtime, and maintenance. A managed screenshot API trades that setup for service pricing and request limits. ScreenshotNeo lists a Free plan with 1,000 shots/month and no card, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed, while bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Check current plan details in its documentation before adopting it.

12. FAQ

Can the JSON contain objects instead of strings?

Yes. Define a schema such as {"pages":[{"id":"home","url":"https://example.com"}]}, validate each object, and use its stable ID for the output filename.

Does Puppeteer save a screenshot automatically after navigation?

No. Call page.screenshot() after navigation and any required readiness checks, and provide a path if you want Puppeteer to save the file.

Can I capture pages that require authentication?

Possibly, if your script has authorized access and supplies the needed session or headers. Keep credentials out of source control and avoid logging sensitive URLs or tokens.

Can I use this in an AI agent?

Yes. ScreenshotNeo also offers an MCP server with screenshot, page-info, and PDF capture tools for MCP clients such as Claude and Cursor; see its docs for setup.

Sources