How to capture screenshots of multiple URLs with Puppeteer on Windows in India
Use Puppeteer on Windows to capture a list of URLs, save distinct screenshots, handle failures, and troubleshoot browser setup.
To capture screenshots of multiple URLs with Puppeteer on Windows, install a supported Node.js version and the puppeteer package, then loop through your URL list: navigate a page to each URL, save a uniquely named image with page.screenshot(), and close the browser when the job ends. Puppeteer normally downloads a compatible Chrome for Testing browser during installation. The workflow below runs sequentially so one failed URL does not stop the rest.
The current Puppeteer system requirements list Node.js 22.12 or newer and Windows x64 for Chrome for Testing. The official installation guide gives an approximate Windows browser download size of 280 MB. These details can change; check the current system requirements and installation guide before setting up a new environment. The reviewed Puppeteer documentation does not identify a separate India-specific setup requirement, so the steps apply regardless of location, subject to your network and organization policies.
1. Install Node.js and Puppeteer
- Install a supported Node.js release for Windows x64.
- Open PowerShell or Windows Terminal and create a project directory:
mkdir puppeteer-batch-shots cd puppeteer-batch-shots npm init -y npm install puppeteer - Create a file named
capture.mjsin that directory. The.mjsextension lets Node.js use JavaScript modules without changing the generated package configuration.
The puppeteer package normally downloads a compatible Chrome for Testing browser. If you manage Chrome yourself and can provide its executable path or channel, puppeteer-core is an option, but it does not download a browser. For a first script, puppeteer avoids that extra browser-management step.
2. Capture a list of URLs
Save this runnable script as capture.mjs. Replace the example URLs with pages you are authorized to access. It creates a screenshots folder, visits each URL in order, records the HTTP status when available, and continues after individual navigation or screenshot errors.
import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
const urls = [
'https://example.com/',
'https://www.wikipedia.org/',
];
const outputDir = 'screenshots';
await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch();
try {
for (let i = 0; i < urls.length; i++) {
const url = urls[i];
const page = await browser.newPage();
try {
await page.setViewport({
width: 1365,
height: 900,
deviceScaleFactor: 1,
});
const response = await page.goto(url, {
waitUntil: 'networkidle2',
timeout: 45_000,
});
const status = response?.status();
const filename = `${outputDir}/page-${String(i + 1).padStart(3, '0')}.png`;
await page.screenshot({ path: filename, fullPage: true });
console.log(`${status ?? 'no response'} ${url} -> ${filename}`);
} catch (error) {
console.error(`Failed: ${url}`, error);
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
Run it from the project directory:
node capture.mjs
Each screenshot is named by list position rather than raw URL. This avoids invalid path characters and prevents two URLs on the same host from overwriting each other. If your input list may contain duplicate entries and you want names that reveal which page was captured, keep the index and add a safely normalized hostname or a short hash; do not use an unfiltered URL as a filename.
3. Choose the right wait and screenshot options
Puppeteer’s documented screenshot method is Page.screenshot(). Its screenshots guide shows capturing after navigation. The loop and its error handling are application code built from those documented calls, rather than a separate batch feature.
Wait until the page is useful
| Strategy | When it helps | Trade-off |
|---|---|---|
waitUntil: 'load' |
Pages whose important content is available at the normal load event. | Client-rendered content or late-loading images may not be ready. |
waitUntil: 'networkidle2' |
A reasonable starting point for pages that settle after network activity; it is also used in Puppeteer’s screenshot example. | Analytics, polling, or other persistent requests can delay or prevent settling. A settled network does not guarantee the exact visual state you want. |
await page.waitForSelector('main article') |
Known sites with a stable selector that appears when the content is ready. | The selector must exist on each relevant page, or the wait times out. |
await new Promise(resolve => setTimeout(resolve, 1500)) |
A small fixed delay for a known animation or delayed render. | Fixed waits can be unnecessarily long or still too short on a slow page. |
For selector-based readiness, place the wait after page.goto() and before the screenshot. Use a meaningful content selector when you control or understand the target pages. Avoid assuming the same wait strategy is reliable for every unrelated site.
Viewport, full-page, and element captures
- Viewport image: omit
fullPageor set it tofalseto capture the visible viewport. - Full-page image: use
fullPage: true, as in the example. Very long pages can create large images, and pages that load content only while scrolling may need additional handling before capture. - Repeatable dimensions: set the viewport before navigation or capture with
page.setViewport({ width, height, deviceScaleFactor }). A scale factor above 1 produces higher-density output and increases image dimensions and file size. - One element: select the target and call its screenshot method, for example:
const element = await page.$('.product-card'); if (!element) throw new Error('Product card was not found'); await element.screenshot({ path: 'screenshots/product-card.png' });Puppeteer documents both page and element screenshot capture in its screenshots guide.
Screenshot options include the output path and full-page capture used above. Consult the current ScreenshotOptions API for supported formats and options in your installed version. If you omit a path, Puppeteer returns screenshot data instead of writing that file for you.
4. Read URLs from a text file
For a larger list, put one URL per line in urls.txt. This variant trims whitespace, skips blank lines and comment lines, and still uses a sequence number for safe, unique output names. Replace the array and its declaration in the earlier example with the following; keep the rest of the capture loop.
import { readFile } from 'node:fs/promises';
const urls = (await readFile('urls.txt', 'utf8'))
.split(/\r?\n/)
.map(line => line.trim())
.filter(line => line.length > 0 && !line.startsWith('#'));
if (urls.length === 0) {
throw new Error('urls.txt contains no URLs');
}
Validate your inputs if the file comes from an external source. For a controlled workflow, you can reject anything other than HTTP or HTTPS URLs before launching the browser:
for (const value of urls) {
const parsed = new URL(value);
if (!['http:', 'https:'].includes(parsed.protocol)) {
throw new Error(`Unsupported URL scheme: ${value}`);
}
}
5. Handle failures and partial results
The example catches errors inside the loop, so one page timing out does not prevent later URLs from being attempted. The outer finally closes Chrome even if the script fails. Decide whether partial output is acceptable: if every page is required, count failed entries and set a nonzero process exit code after the loop.
let failures = 0;
// In the per-URL catch block, add:
failures++;
// After the loop, before the outer finally closes the browser:
if (failures > 0) process.exitCode = 1;
A saved PNG alone does not establish that the intended page loaded. Log the response status, consider checking for a required page element, and inspect outputs for error pages, consent dialogs, bot challenges, or incomplete content. A navigation can also have no main-document response, hence the optional status in the example.
6. Run responsibly and keep the batch manageable
- Start sequentially. One page at a time is easier to debug and limits simultaneous browser work. If you later use multiple pages concurrently, cap the number of open pages and tune it for your machine and target sites; there is no universally optimal concurrency figure.
- Close resources. The example closes every page and closes the browser in a
finallyblock. Keep that cleanup if you add more processing. - Expect variable duration. Browser startup, page size, network response, and the selected wait condition all affect run time. A 45-second navigation timeout is a configurable starting point, not a completion guarantee.
- Plan for output size. Full-page and high-density captures use more memory and disk space than viewport images. Check available storage for large batches and choose image format and dimensions based on downstream use.
- Respect access limits. Automate only pages you are authorized to access, follow site terms and rate limits, and do not use screenshot automation to bypass authentication or anti-bot controls.
7. Windows troubleshooting
| Symptom | Likely cause | What to try |
|---|---|---|
Could not find Chrome or browser executable missing |
The package manager blocked Puppeteer’s install lifecycle script, so its browser download did not happen. | Run npx puppeteer browsers install, or configure the package manager to permit the Puppeteer install script. See the official installation guide. |
| Browser cache is missing or inaccessible | The browser was installed under a different user or cache location, or the default user-home cache cannot be accessed. | Check the installing and running accounts. Puppeteer supports selecting a cache location with PUPPETEER_CACHE_DIR or a configuration file; after changing download configuration, reinstall the browser if needed. See configuration. |
| Chrome will not start under an organization policy | A managed Chrome policy may enforce extensions while Puppeteer disables extensions by default. | Check the policy and Puppeteer’s troubleshooting guide. It documents enableExtensions: true for this environment-specific case; do not add it unless the policy requires it. |
Windows sandbox error or Access is denied |
Chrome’s sandbox setup or directory permissions may be blocked, particularly with older versions or restrictive managed environments. | Use a current Puppeteer version first. Its guide says it attempts sandbox permission setup with Chrome’s setup.exe starting in v22.14.0. For persistent issues, follow the documented Windows permissions guidance and your administrator’s policy rather than broadly loosening permissions. |
net::ERR_BLOCKED_BY_CLIENT on an HTTP URL |
Chrome for Testing may show an HTTPS-first warning for some remote HTTP pages. | Use HTTPS when available. If HTTP is required, inspect the resulting warning page and consult the Puppeteer troubleshooting notes. |
Navigation times out or networkidle2 never settles |
The site is slow, has persistent network activity, or does not reach the requested lifecycle state. | Try a more suitable lifecycle event, wait for a page-specific selector, or use a bounded delay for known behavior. Raise the timeout only when slower navigation is expected, and still check that the page content is correct. |
| Screenshot exists but content is blank, incomplete, or unexpected | The page may require client rendering, scrolling, a particular viewport, or interaction; it may also have returned an error or challenge page. | Wait for the relevant content selector, set the intended viewport, check response status, and inspect the image. Do not treat a file’s existence as proof of a successful capture. |
| Output is overwritten or filename creation fails | Names were derived from URLs without sanitizing characters, or repeated URLs collided. | Use an index-based name like page-001.png; add a sanitized hostname or stable suffix if needed. |
8. Or skip the browser setup
If you want the screenshot without installing and managing Chrome on Windows, ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL. The API accepts the parameter names used by other screenshot APIs, which can make switching straightforward. See the ScreenshotNeo API documentation for options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed before capture, along with supported newsletter popups and chat widgets. Each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server lets Claude, Cursor, and other MCP clients use
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; all features are on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
Frequently asked questions
Does Puppeteer require a different setup in India?
The reviewed official documentation does not specify a separate India setup. Use the documented Windows requirements and account for your own network and organization policies.
Can I save JPEG instead of PNG?
Yes. Check the current ScreenshotOptions documentation for the supported format and quality settings in the Puppeteer version installed in your project.
Can I use an already installed Chrome?
Yes. puppeteer-core is intended for a browser you manage; configure its executable path or channel and ensure the browser version is compatible. The regular puppeteer package is simpler when you want its managed Chrome for Testing download.
Why does the script take different amounts of time for different URLs?
Each page has different network activity and rendering behavior, and your wait condition determines when the script proceeds. Use a readiness condition tied to the content you need when pages have different load patterns.


