ScreenshotNeo

BlogHow-to

How to Bulk Screenshot URLs with Playwright on a Windows Server in India

Capture a list of URLs with Playwright on Windows: install matching browsers, save viewport or full-page screenshots, and handle failures reliably.

By the ScreenshotNeo team4 October 20269 min read

To bulk screenshot URLs with Playwright on a Windows server, install Playwright and its matching browser binary, read URLs from a list, visit each URL in a headless browser, save a viewport or full-page image under a stable filename, and log each success or failure so failed URLs can be retried. The example below uses Node.js and Chromium; the same workflow applies to other supported Playwright browsers.

The title mentions India, but the workflow does not require an India-specific Playwright setting. If a site varies its content by location, set and validate the browser context locale or timezone for the behavior you need. Do not assume those settings make a server appear to be physically located in India.

1. Install Playwright on Windows

Use a supported Node.js installation and install Playwright in the project that will run the scheduled job. Playwright browser binaries are tied to Playwright releases, so install the browser after installing or updating the package.

mkdir playwright-bulk-shots
cd playwright-bulk-shots
npm init -y
npm install playwright
npx playwright install chromium

For a headless-only Chromium job, Playwright documents a smaller headless-shell installation option. Use it only when it matches the way your script launches Chromium:

npx playwright install chromium --only-shell

On Windows, the documented default browser cache is %USERPROFILE%\AppData\Local\ms-playwright. A scheduled task or Windows service may run under a different account from the one used to install Playwright. Confirm that the job account can access the browser binaries, or configure a deliberate shared browser path. After changing the Playwright version, rerun the install command for the matching browser build. See the [Playwright browser installation guide](https://playwright.dev/docs/browsers).

2. Prepare the URL list and output directory

Store one absolute URL per line in a UTF-8 text file named urls.txt. Blank lines and lines beginning with # are ignored in this example.

https://example.com
https://www.microsoft.com
# Add one URL per line

Use a dedicated output directory and a stable naming scheme. Filenames based on a sequence number avoid collisions when two URLs have the same host or path. The script below also writes a JSON Lines log so you can identify and retry failed entries.

3. Run a sequential bulk screenshot script

Create capture.js in the project directory. This runnable version reuses one Chromium process, opens and closes a fresh page for each URL, applies explicit navigation and screenshot timeouts, and continues after individual failures. It defaults to full-page PNGs. Set FULL_PAGE=false in the environment to capture only the viewport.

const fs = require('node:fs/promises');
const path = require('node:path');
const { chromium } = require('playwright');

const INPUT_FILE = path.resolve('urls.txt');
const OUTPUT_DIR = path.resolve('screenshots');
const LOG_FILE = path.resolve('results.jsonl');
const NAVIGATION_TIMEOUT_MS = 45_000;
const SCREENSHOT_TIMEOUT_MS = 60_000;
const FULL_PAGE = process.env.FULL_PAGE !== 'false';

function readUrls(contents) {
  return contents
    .split(/\r?\n/)
    .map(line => line.trim())
    .filter(line => line && !line.startsWith('#'));
}

async function appendResult(record) {
  await fs.appendFile(LOG_FILE, `${JSON.stringify(record)}\n`, 'utf8');
}

(async () => {
  await fs.mkdir(OUTPUT_DIR, { recursive: true });
  await fs.writeFile(LOG_FILE, '', 'utf8');

  const urls = readUrls(await fs.readFile(INPUT_FILE, 'utf8'));
  if (urls.length === 0) {
    throw new Error(`No URLs found in ${INPUT_FILE}`);
  }

  const browser = await chromium.launch({ headless: true });
  let failures = 0;

  try {
    for (let index = 0; index < urls.length; index += 1) {
      const url = urls[index];
      const filename = `${String(index + 1).padStart(5, '0')}.png`;
      const outputPath = path.join(OUTPUT_DIR, filename);
      let page;

      try {
        const parsed = new URL(url);
        if (!['http:', 'https:'].includes(parsed.protocol)) {
          throw new Error(`Unsupported URL protocol: ${parsed.protocol}`);
        }

        page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
        page.setDefaultNavigationTimeout(NAVIGATION_TIMEOUT_MS);
        page.setDefaultTimeout(NAVIGATION_TIMEOUT_MS);

        const response = await page.goto(url, { waitUntil: 'load' });
        if (!response) {
          throw new Error('Navigation returned no main-resource response');
        }

        await page.screenshot({
          path: outputPath,
          fullPage: FULL_PAGE,
          type: 'png',
          timeout: SCREENSHOT_TIMEOUT_MS,
        });

        await appendResult({
          index: index + 1,
          url,
          status: 'ok',
          httpStatus: response.status(),
          file: outputPath,
        });
        console.log(`OK ${index + 1}/${urls.length} ${url} -> ${outputPath} (HTTP ${response.status()})`);
      } catch (error) {
        failures += 1;
        const message = error instanceof Error ? error.message : String(error);
        await appendResult({ index: index + 1, url, status: 'error', error: message });
        console.error(`FAIL ${index + 1}/${urls.length} ${url}: ${message}`);
      } finally {
        if (page) await page.close().catch(() => {});
      }
    }
  } finally {
    await browser.close();
  }

  console.log(`Finished: ${urls.length - failures} succeeded, ${failures} failed. See ${LOG_FILE}.`);
  if (failures > 0) process.exitCode = 1;
})().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it from the project directory:

node capture.js

For viewport screenshots in PowerShell:

$env:FULL_PAGE = 'false'
node .\capture.js

The script waits for the page’s load event, which is a starting point rather than a universal readiness signal. Pages with client-rendered content, delayed images, or ongoing background requests may need a site-specific condition before the screenshot. See the readiness section below.

4. Choose viewport, full-page, or element capture

Capture Use it when Trade-off
Viewport You need a consistent screen-sized snapshot. Content below the current viewport is omitted.
Full page You need the entire scrollable page. Long pages can create very tall, larger images and take more memory to render and encode.
Element You need one chart, card, or page section. The target element must exist and be visible when captured.

Playwright’s Page API supports full-page capture and locator screenshots. For element capture, replace the page screenshot call with a locator screenshot:

await page.locator('#main-content').screenshot({ path: outputPath, timeout: SCREENSHOT_TIMEOUT_MS });

For full-page output, use fullPage: true; for the viewport, use fullPage: false. The screenshot API also supports image options such as type and quality where applicable. Check the documentation for the Playwright version installed in the project before depending on an option. See [Page.screenshot](https://playwright.dev/docs/api/class-page#page-screenshot) and the [screenshot guide](https://playwright.dev/docs/screenshots).

5. Match readiness to the website

No single wait condition works for every dynamic site. Choose a signal that corresponds to the content you need:

  • Navigation event: load waits for the page load event, as in the example. It may still be too early for content fetched afterward.
  • Specific content: wait for a stable selector that marks the section you intend to capture.
  • Known delay: use a short delay only when the site behavior is understood and validated; fixed sleeps slow every capture and can still miss variable delays.
  • Network quiet: use a network-idle condition only when the site’s request behavior makes it suitable. Analytics, polling, and streaming can prevent quiet, while some pages render before all requests stop.

For a site-specific selector, add a wait after navigation:

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('[data-ready="true"]').waitFor({ state: 'visible', timeout: 15_000 });
await page.screenshot({ path: outputPath, fullPage: FULL_PAGE });

Replace the selector with a real readiness marker from the target site. If lazy-loaded images matter, verify that the page has loaded them before capture; a full-page screenshot should not be treated as proof that every site’s lazy content is ready.

6. Run reliably as a Windows scheduled job

  1. Install the project dependencies and Chromium under the same Windows account that will run the job, or verify access to the installed browser cache.
  2. Use absolute paths for the project, input file, output directory, and logs. Scheduled tasks may start with a different working directory than an interactive shell.
  3. Set an appropriate execution time limit and ensure the task does not overlap with a previous batch unless outputs and server capacity are managed for overlap.
  4. Keep the input list and result log so failed URLs can be isolated. Retry transient navigation failures separately rather than silently rerunning every successful capture.
  5. Rotate or archive screenshots and logs according to your storage policy. Full-page images can consume substantially more disk than viewport captures.

One browser reused sequentially keeps browser startup overhead out of every URL iteration and limits simultaneous page load. For higher throughput, add controlled concurrency only after observing CPU, memory, target-site behavior, and failure rates on your server. There is no universal safe concurrency or URLs-per-minute figure: page weight and site behavior vary.

7. Troubleshooting

Symptom Likely cause What to check or change
Executable doesn’t exist or browser cannot launch The browser binary is missing, mismatched with the installed Playwright package, or inaccessible to the job account. Run npx playwright install chromium for the project’s installed version as the runtime account; confirm its browser cache path and permissions.
Works interactively but fails in Task Scheduler The scheduled task uses a different account, working directory, environment, or profile. Set the task’s start-in directory, use absolute paths, and install or expose the browser binaries for that account.
Navigation timeout The server or target is slow, unreachable, blocked, or waiting on a page event that never occurs. Check DNS and outbound connectivity, inspect the URL and response behavior, choose an appropriate navigation event, and set a timeout that fits the workload.
Screenshot is blank or missing expected content The capture occurred before client-rendered content appeared, a selector was wrong, or the site returned an error or bot-check page. Log the response status, wait for a meaningful site-specific selector, and inspect a saved diagnostic capture when appropriate.
Full-page screenshot is huge or fails The page is extremely long or has complex content that consumes memory during layout and encoding. Use viewport or element capture if it meets the need, and reduce concurrency. Consider whether the page can be split into sections.
Some URLs overwrite the same output Names were derived from a non-unique host or path. Use a sequence number, a stable hash of the complete URL, or both.
Batch stops at the first bad URL An exception is not caught per URL. Keep navigation and capture in a per-URL try/catch, record the error, and continue the loop.

8. Performance, reliability, and cost

Sequential capture is easier to diagnose and puts a predictable limit on browser work, but each URL waits for earlier URLs to finish. Controlled concurrency can reduce wall-clock time while increasing memory, CPU, network use, and load on the sites being captured. Start sequentially, measure your own workload, then raise concurrency gradually with a cap.

Use viewport output when a full page is unnecessary. Keep PNG when lossless detail matters; choose another supported format and quality setting only after checking downstream requirements and the installed API version. Write output and logs to storage with enough capacity, and avoid retaining captures longer than the task requires. Playwright itself does not charge per screenshot; your costs are the Windows server, storage, and network usage.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. A single GET request takes a URL and returns a PNG, JPEG, WebP, or PDF. The API accepts parameter names used by other screenshot APIs, which can make migration simpler. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets can be removed; each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
  • An MCP server gives AI agents tools for screenshots, page information, and PDF capture.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does the server need to be in India?

No India-specific server setting is established for this Playwright workflow. Use locale or timezone context settings only when you need to test localized page behavior, and validate the result on your target site.

Can I use Firefox or WebKit instead of Chromium?

Playwright supports multiple browser engines. Install the matching browser and use it when your capture requirement or target-site behavior calls for that engine; keep the installed browser aligned with the Playwright package version.

Can the script capture PDFs too?

Playwright’s page API includes PDF generation for Chromium. Use it when a document is the required output; check the current API documentation for its format and page settings.

Why does a successful HTTP status still produce the wrong screenshot?

An HTTP response only describes the main navigation response. Client-side rendering, consent screens, anti-bot checks, or delayed content can still change what appears in the captured page.

Sources