How to Capture Website Thumbnails for a Directory of Indian Online Courses
Build a consistent course directory with Playwright screenshots, a CSV input list, image sizing guidance, and a workflow for reviewing failed or stale captures.
To capture website thumbnails for a directory of Indian online courses, automate a browser visit to each provider page, wait for meaningful content, capture either a consistent viewport or a specific course element, and save the image with the source URL and capture date. For most directory cards, use a fixed viewport or a clear hero section: a very tall full-page screenshot often becomes unreadable when reduced to card size.
This guide uses Playwright with Node.js. It reads provider URLs from a CSV file, captures WebP images, records per-URL success or failure, and supports either viewport or element screenshots. Playwright supports file screenshots, full-page screenshots, element screenshots, and screenshot buffers for post-processing. See the Playwright screenshot guide and Page screenshot API.
1. Choose what each directory card should show
Decide the image’s purpose before automating. A course directory thumbnail is a preview of a webpage view, not the provider’s favicon. For a card, consistency and legibility matter more than capturing every section of a landing page.
| Capture type | Use it when | Trade-off |
|---|---|---|
| Fixed viewport | You want a consistent first impression across many providers. | Content below the fold is omitted. |
| Element | A page has a clear course banner or provider mark. | You need a selector that identifies the right element on that site. |
| Full page | The entire page is itself the preview. | A long screenshot can become too small to read on a directory card. |
Use one directory card ratio and one output size throughout your catalog. The capture can be larger than the rendered card so it remains clear on high-density displays, but avoid mixing dimensions arbitrarily. These are design choices for your directory; Playwright provides the capture controls, not a standard thumbnail size.
A favicon serves a different purpose: it is a compact icon representing a site. Google’s Search guidance says a favicon for eligibility in Search must be square and at least 8×8 pixels, and recommends more than 48×48 pixels; that is not a specification for course-directory screenshots. Google’s favicon guidance
2. Set up the batch capture script
Install Node.js, create a project, install Playwright, and install its Chromium browser. The script below takes a viewport screenshot by default. You can optionally set a CSS selector for a page where the course hero is more useful than the entire visible viewport.
mkdir course-thumbnails
cd course-thumbnails
npm init -y
npm install playwright
npx playwright install chromium
Create courses.csv. Keep a stable slug for the output filename and a label editors can recognize. The optional third field is a CSS selector for an element screenshot.
slug,label,url,selector
provider-one,Provider One,https://example.com/courses/data-science,
provider-two,Provider Two,https://example.org/learn/design,.course-hero
The example domains are placeholders; replace them with course-provider URLs you are authorized to capture. Save the following as capture.mjs:
import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
const input = process.argv[2] ?? 'courses.csv';
const outputDir = process.argv[3] ?? 'thumbnails';
const mode = process.env.CAPTURE_MODE ?? 'viewport'; // viewport | full
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 900);
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 45000);
function parseCsv(text) {
// Handles quoted CSV fields and commas inside quoted values.
const rows = [];
let row = [], field = '', quoted = false;
for (let i = 0; i < text.length; i++) {
const ch = text[i];
if (quoted) {
if (ch === '"' && text[i + 1] === '"') { field += '"'; i++; }
else if (ch === '"') quoted = false;
else field += ch;
} else if (ch === '"') quoted = true;
else if (ch === ',') { row.push(field); field = ''; }
else if (ch === '\n') { row.push(field.replace(/\r$/, '')); rows.push(row); row = []; field = ''; }
else field += ch;
}
if (field.length || row.length) { row.push(field.replace(/\r$/, '')); rows.push(row); }
const [header, ...data] = rows;
return data.filter(r => r.some(Boolean)).map(r => Object.fromEntries(header.map((key, i) => [key.trim(), (r[i] ?? '').trim()])));
}
function safeSlug(value) {
return value.toLowerCase().replace(/[^a-z0-9-]+/g, '-').replace(/^-|-$/g, '') || 'course';
}
await mkdir(outputDir, { recursive: true });
const courses = parseCsv(await readFile(input, 'utf8'));
const browser = await chromium.launch({ headless: true });
const results = [];
try {
for (const course of courses) {
const slug = safeSlug(course.slug || course.label || new URL(course.url).hostname);
const result = { slug, label: course.label ?? '', url: course.url ?? '', capturedAt: new Date().toISOString() };
let page;
try {
const parsedUrl = new URL(course.url);
if (!['http:', 'https:'].includes(parsedUrl.protocol)) throw new Error('Only HTTP and HTTPS URLs are allowed');
page = await browser.newPage({
viewport: { width, height },
deviceScaleFactor: 1,
colorScheme: 'light',
locale: 'en-IN',
timezoneId: 'Asia/Kolkata'
});
page.setDefaultNavigationTimeout(timeoutMs);
const response = await page.goto(course.url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
result.httpStatus = response?.status() ?? null;
// Wait for the document to finish loading where possible, but do not fail solely
// because analytics, chat, or other long-lived requests keep running.
await page.waitForLoadState('load', { timeout: 15000 }).catch(() => {});
await page.evaluate(() => document.fonts?.ready);
// Scroll in steps to trigger common lazy-loaded images, then return to the top.
await page.evaluate(async () => {
const step = Math.max(300, Math.floor(innerHeight * 0.8));
for (let y = 0; y < document.body.scrollHeight; y += step) {
window.scrollTo(0, y);
await new Promise(resolve => setTimeout(resolve, 120));
}
window.scrollTo(0, 0);
});
await page.waitForTimeout(250);
const file = path.join(outputDir, `${slug}.webp`);
if (course.selector) {
const target = page.locator(course.selector).first();
await target.waitFor({ state: 'visible', timeout: 10000 });
await target.screenshot({ path: file, type: 'webp', quality: 82, animations: 'disabled' });
result.captureType = 'element';
result.selector = course.selector;
} else {
await page.screenshot({
path: file,
type: 'webp',
quality: 82,
fullPage: mode === 'full',
animations: 'disabled'
});
result.captureType = mode === 'full' ? 'full-page' : 'viewport';
}
result.file = file;
result.status = response && response.status() >= 400 ? 'http-error-captured' : 'captured';
} catch (error) {
result.status = 'failed';
result.error = String(error?.message ?? error);
} finally {
await page?.close().catch(() => {});
}
results.push(result);
console.log(`${result.status}: ${result.url}${result.file ? ` -> ${result.file}` : ` (${result.error})`}`);
}
} finally {
await browser.close();
await writeFile(path.join(outputDir, 'manifest.json'), JSON.stringify(results, null, 2));
}
const failures = results.filter(result => result.status === 'failed').length;
console.log(`Done: ${results.length - failures} captured, ${failures} failed. Manifest: ${path.join(outputDir, 'manifest.json')}`);
process.exitCode = failures ? 1 : 0;
Run it with defaults:
node capture.mjs courses.csv thumbnails
Capture the whole scrollable page instead of the initial viewport:
CAPTURE_MODE=full node capture.mjs courses.csv thumbnails
Change the common viewport dimensions or navigation timeout without editing the script:
VIEWPORT_WIDTH=1365 VIEWPORT_HEIGHT=768 NAVIGATION_TIMEOUT_MS=60000 node capture.mjs
The manifest records each source URL, capture date, capture type, HTTP status when available, and failure message. Store it with the generated files so editors can trace a thumbnail back to the provider page and decide when to refresh it.
3. Make captures consistent and useful
Wait for content without waiting forever
The script waits for the DOM, then gives the page a chance to finish its load event and lets fonts settle. It scrolls in steps to trigger common lazy-loaded images. It deliberately does not require every network request to become idle: analytics, chat, and streaming requests can remain open indefinitely. A fixed viewport and a bounded timeout make a batch more predictable, but they cannot guarantee that every provider page has finished rendering its meaningful content.
If a particular course page renders late, add a per-site selector and wait for it before capture. For example, immediately before the screenshot call, use await page.locator('main h1').waitFor({ state: 'visible', timeout: 15000 });. Prefer a meaningful visible selector to an arbitrary long sleep. A delay can help with a known animation or delayed asset, but it adds time to every capture and still does not prove the correct content appeared.
Keep browser conditions stable
The script fixes viewport, device scale, color scheme, locale, and timezone. Playwright documents format, clipping, quality, scale, masking, and stylesheet controls. Browser rendering can vary with operating system, browser version, settings, hardware, power state, and headless mode; for stable refreshes, run captures in the same environment. Playwright visual comparison guidance
For pages with distracting or changing elements, inject CSS before capturing. Example:
await page.addStyleTag({ content: `
.newsletter-modal, .support-chat, .cookie-banner {
display: none !important;
}
` });
Only use selectors you have checked against the target site. Hiding a consent prompt may change what a visitor sees; ensure the capture approach complies with the provider’s terms and your editorial requirements. You can also use Playwright screenshot masking for selected locators when a region should be visually obscured rather than removed.
Pick the output format and scale
- WebP: a practical lossy format for compact directory images; the example uses quality 82.
- JPEG: another lossy option, useful when your publishing pipeline expects it; select the type and quality in the screenshot options.
- PNG: lossless and useful for sharp edges or when avoiding lossy artifacts matters, usually with larger files.
- Scale: CSS-pixel scale produces smaller, uniform output; device-pixel scale captures more pixels on high-density devices and can make files larger or dimensions vary. Keep the device scale fixed for comparable output.
Playwright can return a screenshot buffer instead of writing directly to disk, allowing later processing such as a crop or image optimization step. Use one final card ratio across the directory; crop deliberately so important course titles and imagery remain visible.
4. Run a quality review before publishing
Automation can produce an image even when it is the wrong image. Review a sample from every batch and inspect all failures or unusual status results.
- Does the URL resolve to the intended provider and course listing?
- Is the useful content visible, or did a cookie banner, login wall, bot check, or error page dominate?
- Are images loaded and text legible at the directory’s rendered card size?
- Is the crop consistent with other cards, and does it omit irrelevant page sections?
- Is the provider branding current, and does the manifest include the source URL and capture date?
- Are you allowed to republish the captured page image for this directory?
The reviewed documentation does not establish that capturing a page grants permission to republish its contents. Check the provider’s terms and rights for your specific use; the legal position for publishing third-party course-site screenshots in India is not resolved here.
5. Troubleshooting common capture problems
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation timeout | The provider is slow, a request hangs, or the timeout is too short. | Increase NAVIGATION_TIMEOUT_MS; keep waiting bounded. The script already tolerates a load event timeout after DOM content is available. |
| Screenshot is blank or mostly empty | Content is client-rendered, a bot check intervened, or the page did not reach its meaningful state. | Inspect the URL in a browser, wait for a page-specific visible selector, and record the failed or blocked capture rather than publishing it as a thumbnail. |
| Course images are missing | Images are lazy-loaded or load after the screenshot. | Use the script’s scroll pass, wait for the relevant image or hero element, and re-capture. Some sites may require page-specific handling. |
| Selector timeout or empty element capture | The selector does not exist on that page, changed, or matches a hidden element. | Inspect the page, update the CSV selector to a stable visible element, and verify the selector manually before rerunning. |
| Cookie banner, newsletter popup, or chat widget covers the page | A consent or overlay component obscures content. | Review the site’s expected visitor flow. Where appropriate, accept consent as a visitor or inject narrowly targeted CSS after checking site terms; do not silently misrepresent the page. |
| Screenshot differs between machines | Browser version, OS, fonts, viewport, scale, or headless rendering differs. | Use a consistent browser and operating environment, pin your environment, and fix viewport and device scale. |
| WebP output is rejected downstream | Your CMS or image pipeline does not accept WebP. | Change screenshot type to jpeg or png and use the matching filename extension. |
| HTTP error was captured as an image | The server returned an error page with an HTTP response. | Check httpStatus in manifest.json; the script marks such captures http-error-captured so they can be excluded from publication. |
6. Performance, reliability, and cost
Browser capture has no per-screenshot API fee, but you operate the browser runtime, storage, and review process. Each URL takes time for navigation, rendering, and image writing. The sample processes URLs sequentially to keep resource usage simple and isolate failures. For a larger directory, you can add a small concurrency limit, but avoid launching an unbounded number of browser pages: provider throttling, memory pressure, and incomplete captures can undermine the time saved.
Make refreshes incremental. Keep a manifest, retry only transient failures, and avoid recapturing every provider on each run. For operational reliability, preserve the previous approved thumbnail until a replacement passes review. A successful navigation is not proof that the screenshot contains the intended course content; the batch status and human review serve different purposes.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One request returns an image or PDF, and its options include viewport or full-page capture, an element selector, device presets, CSS and JavaScript, waiting rules, custom headers and cookies, and more. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Replace the example URL with a provider URL. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month.
Frequently asked questions
Should I capture a course provider’s homepage or the exact course page?
Capture the page that supports the directory entry. If the listing describes a specific course, use that course’s landing page where possible and store its URL in the manifest.
Can I use a favicon as the directory thumbnail?
You can display a favicon as a small site identifier, but it is not a webpage preview. Use a screenshot for a visual page thumbnail and label the two assets distinctly.
How often should thumbnails be refreshed?
Set a refresh cadence that matches how often your directory needs current branding and course information. The capture date in the manifest lets editors identify old images; the sources reviewed do not prescribe a universal schedule.
Does capturing a public page mean I can republish the screenshot?
No permission conclusion follows from the capture method. Check the provider’s terms and obtain any permissions needed for your intended publication.


