ScreenshotNeo

BlogComparisons

Best Open Source Tools for Bulk Website Screenshots

Compare shot-scraper and Playwright for automated bulk website screenshots, with runnable batch examples, workflow guidance, and a hosted API alternative.

By the ScreenshotNeo team4 October 20269 min read

For a screenshot-focused command-line workflow with a documented multiple-capture configuration, start with shot-scraper. Choose Playwright when screenshots are one part of custom browser automation or visual regression tests. Both can capture many pages; neither documentation establishes a universal batch limit, throughput, or success rate, so test a representative sample in the environment where the job will run.

If you want a hosted API instead of managing a browser installation, ScreenshotNeo is the first service to try: it removes common consent banners, popups, and chat widgets before capture, bills only clean shots, and its lowest paid plan is $5.

1. Choose the tool that fits the batch

Need Good starting point Why
A repeatable list of URLs and a CLI shot-scraper It is a screenshot-focused command-line utility with a documented multiple-screenshot workflow.
Custom navigation, interactions, or browser logic Playwright Its browser automation APIs let code control pages, locators, and capture timing.
Visual regression assertions Playwright Test The test runner can create screenshot baselines and compare later runs against them.
A screenshot API without browser setup ScreenshotNeo One GET request returns an image or PDF; clean captures are billed, and failed loads and cache hits are not.

Evaluate setup effort, how you define the URL batch, viewport versus full-page output, authentication and interactions, CI integration, and whether output must be comparable across runs. The official guides cover these controls, but the reviewed sources do not provide a performance comparison.

2. Install and capture a batch with shot-scraper

shot-scraper is a Python package with a CLI. Install the package and its browser, then put one URL on each line of an input file.

python -m pip install shot-scraper
shot-scraper install

cat > urls.txt <<'EOF'
https://example.com/
https://www.python.org/
https://www.sqlite.org/
EOF

mkdir -p screenshots
while IFS= read -r url; do
  [ -z "$url" ] && continue
  name=$(printf '%s' "$url" | sed -E 's#https?://##; s#[^A-Za-z0-9._-]+#-#g; s#-+$##')
  shot-scraper "$url" -o "screenshots/${name}.png"
done < urls.txt

This shell loop is a simple sequential batch wrapper around the CLI. It writes a separate PNG for each URL, skips empty lines, and derives a filename from the URL. For production jobs, use a stable slug list or explicit mapping if URLs can normalize to the same filename. The shot-scraper documentation also describes taking multiple screenshots through its multi-capture configuration; use that documented workflow when you want the batch defined in a configuration file rather than a shell loop.

Useful shot-scraper capabilities documented by the project include viewport sizing, selector-based capture, authentication, waiting for delays or conditions, page interaction, output formats, and JavaScript. Consult the official documentation for the exact flags and configuration syntax for your installed version rather than guessing options.

3. Capture a batch with Playwright

Playwright is a browser automation library. Install it for Node.js, install a browser, and run a script that visits each URL and saves either a viewport or full-page screenshot.

npm init -y
npm install playwright
npx playwright install chromium

cat > capture.mjs <<'EOF'
import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter(Boolean);

await mkdir('screenshots', { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
  const context = await browser.newContext({
    viewport: { width: 1440, height: 900 },
    deviceScaleFactor: 1
  });
  const page = await context.newPage();
  for (let i = 0; i < urls.length; i++) {
    const url = urls[i];
    try {
      const response = await page.goto(url, {
        waitUntil: 'networkidle',
        timeout: 45000
      });
      const status = response?.status() ?? 'no response';
      await page.screenshot({
        path: `screenshots/${String(i + 1).padStart(4, '0')}.png`,
        fullPage: true
      });
      console.log(`${status}\t${url}`);
    } catch (error) {
      console.error(`FAILED\t${url}\t${error.message}`);
    }
  }
  await context.close();
} finally {
  await browser.close();
}
EOF
node capture.mjs

The example uses a single page and processes URLs sequentially, which keeps browser resource use simpler and makes failures attributable to a URL. It records an HTTP status when a response exists, but a saved screenshot does not itself prove that the page is semantically correct. Add assertions or an explicit failure policy if downstream consumers must reject error pages.

Playwright supports viewport screenshots, full-page screenshots, locator or element screenshots, and returning image buffers for further processing. For custom authenticated or interactive captures, use its browser APIs to establish the required context, navigate, perform actions, and then capture. Keep credentials in CI secrets or another secret store; do not commit them in the URL file or script.

4. Configure output and repeatability

Choose what each file represents

  • Viewport: captures the currently visible browser area. Use a fixed viewport when comparing pages or capturing a consistent above-the-fold view.
  • Full page: captures the page beyond the viewport. Long pages can produce large images, and lazy-loaded content may require scrolling or other page-specific preparation.
  • Element: captures a selected region. Playwright documents locator screenshots; shot-scraper documents CSS-selected capture areas.

Settle the page before capture

Pages with client-side rendering, animation, ads, or delayed content may not be ready at the same time. shot-scraper documents waiting for delays or conditions. In Playwright, choose an appropriate navigation wait condition, wait for a selector that represents readiness, or perform the interaction needed before capture. A network-idle condition can be unsuitable for pages that keep connections open; use an explicit readiness condition when possible.

Keep visual runs comparable

Rendering can vary with browser version, operating system, fonts, viewport, device scale, and page state. Use the same browser and execution environment for baseline and later runs. Control volatile content where possible; Playwright Test supports custom stylesheets for filtering volatile elements in visual assertions.

Use Playwright Test for regression checks

When the purpose is to detect visual changes rather than archive images, use Playwright Test screenshot assertions. The first run can generate reference images; subsequent runs compare captures against those baselines. Its documented options include tolerances for pixel differences and custom stylesheets. Review baseline changes deliberately: dynamic timestamps, rotating content, and third-party widgets can create diffs unrelated to an application change.

5. Add bulk capture to CI

  1. Pin the project dependencies and browser installation strategy used by the job.
  2. Start with a short representative URL list, including one authenticated or slow page if those occur in the real workload.
  3. Save artifacts with deterministic names and preserve logs that identify the URL and failure.
  4. Set a per-page timeout and decide whether one failed URL should fail the whole job or only be reported.
  5. Schedule or trigger the job as needed. shot-scraper documents GitHub Actions workflows and recurring screenshot examples.
  6. Keep visual baselines and the environment that creates them consistent across runs.

Do not select a concurrency level from an assumed throughput figure. The reviewed documentation does not establish a universal maximum batch size or speed for either tool. Increase parallelism only after observing memory use, target-site behavior, and failure patterns in your own workload.

6. Authentication, interactions, and difficult pages

Some pages require login, a consent choice, a menu click, or a form state before the desired image exists. shot-scraper documents authentication and page interaction workflows. Playwright is appropriate when the steps need custom code, such as entering a flow, selecting an account state, or waiting for a specific element. Keep each capture’s state explicit: reuse a context only when shared cookies and storage are intended, and isolate contexts when pages require different identities.

For full-page captures, consider sticky elements, lazy-loaded images, and pages with infinite scrolling. A full-page screenshot does not guarantee that content which appears only after scrolling has loaded. Add a deliberate scroll or readiness step for such pages, and define a stopping condition for unbounded feeds.

7. Troubleshooting

Symptom Likely cause What to do
Browser executable missing The package is installed but its browser has not been installed in this environment. Run the browser installation step for the chosen tool and ensure CI uses the same setup.
Navigation times out The site is slow, unreachable, or never reaches the chosen wait condition. Check the URL and network access, raise a suitable per-page timeout, and wait for a page-specific readiness selector rather than an unnecessarily strict global condition.
Screenshot is blank or incomplete The app has not rendered yet, a navigation failed, or content is lazy-loaded. Inspect the response and page state, wait for a meaningful element, and scroll where the page requires it before capture.
Login page appears instead of target Authentication state was not supplied or expired. Refresh the documented authentication setup, verify cookies or storage in the browser context, and keep secrets outside source control.
Files overwrite each other Different URLs normalized to the same filename. Use an index or stable unique ID in filenames and retain the original URL in a manifest.
Large images consume excessive storage Full-page output or high device scale creates large files. Use viewport or element captures where they satisfy the task, reduce the viewport or scale where appropriate, and define artifact retention.
Visual diffs are noisy Browser or host differences, animation, timestamps, or changing third-party content. Run in a consistent environment, wait for stable state, and filter volatile regions with a stylesheet where appropriate.
One bad URL stops the whole batch The script propagates a navigation or capture error. Catch failures per URL, log them, and choose explicitly whether the final job should fail when any URL fails.

8. Performance, reliability, and cost

Performance: browser startup, page load, rendering, and image size all contribute to batch duration. Reusing a browser is practical for a sequence, but parallel pages increase resource use. Measure a representative batch before choosing concurrency; no benchmark or universal capacity claim is established by the reviewed sources.

Reliability: treat each URL as an independent unit of work. Record URL, status, duration, and error; retry only transient failures with a bounded retry policy; and make output names deterministic. A screenshot may capture an HTTP error page, so decide whether HTTP status or page-specific checks determine success.

Cost: shot-scraper and Playwright are open-source software, but running them still consumes machine time, storage, and CI capacity. Large full-page images and repeated runs increase those costs. ScreenshotNeo has a hosted free tier of 1,000 shots monthly without a card; paid plans start at $5 for 3,000 shots. Yearly billing gives two months free, and every feature is available on every plan.

9. Or skip the browser setup

ScreenshotNeo accepts a URL in one GET request and returns an image or PDF. See the API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; response headers indicate the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

10. FAQ

Is there a documented maximum number of URLs per batch?

The reviewed shot-scraper and Playwright documentation does not establish a universal batch limit. Practical capacity depends on the pages and the machine running the capture.

Which tool is easier to put into an existing visual test suite?

Playwright Test is the direct fit when screenshot comparison and assertions belong in the test suite. shot-scraper is a CLI-oriented fit for repeatable capture jobs.

Can I capture pages that change after load?

Yes, provided the workflow waits for the relevant state or performs the required interactions before saving the screenshot. Choose a readiness condition specific to the page.