How to Generate Website Thumbnails for a List of Indian Coaching Institute Websites
Create consistent thumbnails from a list of coaching institute URLs with a runnable Playwright script, failure handling, and a hosted API option.
To generate website thumbnails for a list of Indian coaching institute websites, put the URLs in a file, open each URL in a browser, and save a screenshot at the same viewport and image settings. The Playwright script below does this one site at a time, records each result, and continues when an individual URL fails. No specific coaching institute websites were supplied or tested for this guide, so redirects, availability, layout, and access behavior will vary by site.
For a small or occasional list, Playwright gives you direct control over the browser and output files. If you prefer a hosted batch workflow, ScreenshotNeo accepts up to 100 URLs per bulk capture call; its API can return screenshots without requiring you to maintain browser infrastructure. See the ScreenshotNeo website and API documentation.
1. Prepare the URL list and output profile
Keep one URL per line in a UTF-8 text file named urls.txt. Include the scheme (https:// or http://) and preserve the original URL in the output manifest so you can match images to sites.
https://example.com
https://www.example.org
The domains above are placeholders. Replace them with the URLs you are authorized to capture. Decide on these settings before processing the list:
- Viewport: use one width and height for every page so the thumbnails are comparable. The example uses a desktop viewport of 1365 × 900.
- Image format: Playwright’s screenshot example uses PNG. PNG is lossless and convenient for inspection; convert to JPEG or WebP later if smaller files matter to your destination.
- Capture area: viewport screenshots work well for directory previews. Full-page screenshots include content below the fold, but may have much taller dimensions and inconsistent lengths.
- Wait strategy: use a bounded navigation timeout and a short settling delay. A fixed delay is simple, but dynamic sites may need a site-specific wait condition.
- Names: generate filenames from the URL’s host and add a numeric prefix. Keep a manifest because different URLs can share a hostname or redirect to another page.
Playwright supports capturing the viewport, a specific element, or the full scrollable page. See the Playwright screenshots guide and Playwright screenshot tool documentation.
2. Install Playwright
Use a current Node.js installation. In a new working directory, install Playwright and its Chromium browser:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as thumbnails.mjs in the same directory as urls.txt.
3. Run a batch capture with Node.js and Playwright
import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'thumbnails';
const manifestPath = path.join(outputDir, 'manifest.json');
const width = 1365;
const height = 900;
const navigationTimeoutMs = 45000;
const settleMs = 1200;
function safePart(value) {
return value.replace(/[^a-z0-9.-]+/gi, '-').replace(/^-+|-+$/g, '').slice(0, 100) || 'site';
}
function parseUrl(line, lineNumber) {
const trimmed = line.trim();
if (!trimmed || trimmed.startsWith('#')) return null;
let parsed;
try {
parsed = new URL(trimmed);
} catch {
throw new Error(`Line ${lineNumber}: invalid URL: ${trimmed}`);
}
if (!['http:', 'https:'].includes(parsed.protocol)) {
throw new Error(`Line ${lineNumber}: only http and https URLs are supported: ${trimmed}`);
}
return parsed.href;
}
const rawLines = (await readFile(inputFile, 'utf8')).split(/\r?\n/);
const urls = [];
for (let i = 0; i < rawLines.length; i++) {
try {
const url = parseUrl(rawLines[i], i + 1);
if (url) urls.push(url);
} catch (error) {
console.error(error.message);
process.exitCode = 1;
}
}
if (urls.length === 0) throw new Error(`No valid URLs found in ${inputFile}`);
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
for (let i = 0; i < urls.length; i++) {
const requestedUrl = urls[i];
const index = String(i + 1).padStart(4, '0');
const host = safePart(new URL(requestedUrl).hostname);
const file = `${index}-${host}.png`;
const outputPath = path.join(outputDir, file);
const page = await browser.newPage({
viewport: { width, height },
deviceScaleFactor: 1,
colorScheme: 'light',
reducedMotion: 'reduce'
});
page.setDefaultNavigationTimeout(navigationTimeoutMs);
let result;
try {
const response = await page.goto(requestedUrl, { waitUntil: 'domcontentloaded' });
await page.waitForTimeout(settleMs);
await page.screenshot({ path: outputPath, type: 'png', fullPage: false });
result = {
requestedUrl,
finalUrl: page.url(),
status: response?.status() ?? null,
image: file,
outcome: 'captured'
};
console.log(`Captured ${requestedUrl} -> ${file} (HTTP ${result.status ?? 'unknown'})`);
} catch (error) {
result = {
requestedUrl,
finalUrl: page.url(),
image: null,
outcome: 'failed',
error: String(error.message ?? error)
};
console.error(`Failed ${requestedUrl}: ${result.error}`);
process.exitCode = 1;
} finally {
await page.close();
}
results.push(result);
}
} finally {
await browser.close();
await writeFile(manifestPath, JSON.stringify({ width, height, format: 'png', results }, null, 2));
console.log(`Manifest written to ${manifestPath}`);
}
Run it with:
node thumbnails.mjs urls.txt thumbnails
The script validates URL syntax and protocol, skips blank lines and comments, uses a fixed viewport, captures the viewport rather than the full page, saves a JSON manifest, and moves on after a per-site error. It records the final URL and HTTP status when navigation yields a response. An HTTP error status does not necessarily mean the browser failed to render, so inspect the resulting image and manifest rather than treating status alone as a complete quality check.
4. Tune the capture for your directory
Viewport versus full-page images
For a visual directory of sites, a fixed viewport makes image dimensions predictable. To capture the full scrollable page, change fullPage: false to fullPage: true. Full-page output is useful when you need to inspect page length or below-the-fold content, but the images can be much taller and their content density will vary from site to site.
Mobile previews
For a mobile directory, choose a consistent mobile viewport, such as a width around 390 CSS pixels, and keep it fixed for every URL. Playwright can also emulate device settings; see its documentation for the current device emulation APIs. Do not mix desktop and mobile captures in one comparison grid unless the distinction is clear in your filenames or manifest.
Wait behavior and dynamic content
The script navigates until the DOM is available and then waits 1.2 seconds. This favors throughput over waiting for every external resource. If a site renders important content later, increase settleMs or wait for a stable page-specific selector. A network-idle strategy can be unsuitable for pages that continually poll or stream. Set a finite timeout so one slow page cannot stop the entire batch.
Specific elements and page cleanup
If the goal is to make a card preview of one component, locate that component and screenshot the element instead of the page. You can also use page-level CSS or script logic to hide a fixed overlay for your own capture. Check the page manually after such changes: hiding a banner can expose layout shifts or leave a blank area.
Validation and review
- Check that the manifest has one entry for each valid input line.
- Open a sample from the beginning, middle, and end of the list.
- Review failed rows and pages with unexpected redirects, blank output, interstitials, or consent prompts.
- Retry only the failures with a longer timeout or a tailored wait condition.
- Keep the source URL, final URL, capture profile, and capture date with the image set if the thumbnails need to be regenerated later.
5. Batch reliability, performance, and costs
The example processes pages sequentially and opens a fresh page per URL. Sequential capture is easier to debug and limits the load placed on your machine and target websites. For larger lists, you can add a small worker pool, but keep concurrency bounded: each browser page consumes resources, and many simultaneous requests can trigger rate limits or make results less predictable. No throughput benchmark is claimed here.
For reliability, retain the manifest and retry transient navigation failures selectively. Consider a maximum retry count with a short backoff, and avoid retrying invalid URLs indefinitely. Keep the browser and package versions controlled in repeatable jobs. A local workflow has no per-shot API charge, but it does require maintaining Node.js, Playwright, browser binaries, compute, storage, and the job runner.
For hosted services, check current pricing, request limits, data handling, retention, geography controls, and per-URL failure reporting before moving a production list. Those details vary and were not independently verified for the third-party providers mentioned in the research.
6. Hosted batch alternatives
If you need URL-list submission, queueing, or managed browser infrastructure, hosted capture APIs can reduce the amount of browser setup you maintain. AddScreenshots documents asynchronous bulk processing for URL lists and related inputs at its API Swagger UI. Site-Shot describes a screenshot tool and API, including device and country options on its service page. Capture 815 describes bulk website screenshots and an API at its service page. These are provider descriptions, not independent performance evaluations; confirm current limits and terms directly.
For this use case, ScreenshotNeo is the first hosted option to try: its bulk capture supports up to 100 URLs per call, and only clean shots are billed. It also reports page verdict and billing status in response headers, which helps distinguish a captured page from a bot check, blank result, failed load, or cache hit.
Or skip the browser setup
Send a GET request for a URL to ScreenshotNeo. For a URL list, use its bulk capture feature documented in the API documentation (up to 100 URLs per call). The one-URL call below shows the API’s basic request shape; replace the target URL and save the response as an image.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response indicates page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card required.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
invalid URL in the console |
A line lacks a valid URL or scheme. | Use a complete https:// or http:// URL. Remove accidental spaces and check the line reported by the script. |
| Navigation timeout | The server is slow, the page waits on resources, or network access is unavailable. | Retry that URL, raise navigationTimeoutMs for the affected job, and use domcontentloaded plus a bounded settling wait. Check the URL in a normal browser. |
| Image is blank or incomplete | The important content may be rendered after the fixed delay, blocked, or behind an interstitial. | Increase the settling wait or wait for a page-specific selector. Review manually; no behavior is assumed for any particular institute site. |
| Screenshot shows a consent dialog, popup, or chat overlay | The page presented an overlay during capture. | For a local workflow, handle only the overlay you have deliberately chosen to dismiss, using site-specific selectors and appropriate care. For automated cleanup, consider ScreenshotNeo’s consent and widget removal options. |
| Many pages fail when concurrency is increased | Local CPU or memory may be saturated, or target sites may limit rapid requests. | Reduce worker count, use sequential runs, or add a modest delay between requests. Retry only transient failures. |
| Repeated or overwritten filenames | Different inputs can share a hostname. | Keep the numeric prefix and manifest; if necessary, include a sanitized path segment or stable hash in each filename. |
| Node cannot launch Chromium | The browser binary may not be installed for the current Playwright package or environment. | Run npx playwright install chromium and confirm the runtime can launch headless browsers. |
| HTTP 403, CAPTCHA, or access-denied page | The site returned a restriction or bot check. | Do not assume a screenshot indicates the intended page. Record and review the result, respect the site’s terms, and do not attempt to bypass access controls. |
8. FAQ
Can I make a single contact sheet from the output?
Yes. The script creates one image per URL. Use an image processing tool of your choice to assemble a contact sheet, preserving the manifest so each tile remains traceable to its source.
Should thumbnails be full-page?
Usually not for a compact directory preview. Use a fixed viewport for comparable cards; choose full-page when inspecting page structure or below-the-fold content matters more than uniform image dimensions.
Does this guide verify that each institute website permits automated screenshots?
No. Check the relevant site’s terms and the rules that apply to your intended use. The research did not inspect specific URLs or determine site-specific permissions.


