How to Generate Website Thumbnails for a List of Indian NGO Websites
Build a consistent thumbnail set from NGO homepage URLs with Python and Playwright, including retries, file naming, and practical capture choices.
To generate thumbnails for a list of Indian NGO websites, put the reviewed homepage URLs in a CSV, use a browser to open each URL, and save one screenshot per site with the same viewport, scale, and image format. The Python and Playwright script below does that locally, records failures without stopping the batch, and names files from the NGO names in your input.
For a directory or report, use a viewport screenshot so each image shows a comparable first view. Use full-page capture only when the whole page matters; those images can be much taller and harder to compare. Playwright supports navigation and viewport or full-page screenshots through its Page API and screenshot guide.
1. Prepare the URL list
Create ngos.csv with a name and canonical homepage URL for each organization. Review the URLs first: use the official homepage, include https://, and avoid duplicate entries or campaign pages unless those are what you intend to catalog.
name,url
Example NGO,https://example.org/
Another NGO,https://another-example.org/
The sample names and domains above are placeholders. Replace them with URLs from your own reviewed list; this guide does not claim to have checked or captured any particular NGO website.
2. Install Playwright and its browser
Use Python 3.9 or newer in a virtual environment, then install Playwright and Chromium:
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install playwright
python -m playwright install chromium
3. Capture the list with Python
Save this as make_thumbnails.py. It creates an output directory, captures viewport PNGs at a fixed CSS viewport, retries transient failures, and writes a CSV report containing the final URL, status, and error (if any). Each URL is processed independently, so one failed site does not discard earlier captures or stop later ones.
import csv
import re
import sys
import time
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
INPUT = Path("ngos.csv")
OUTPUT = Path("thumbnails")
REPORT = OUTPUT / "results.csv"
WIDTH = 1280
HEIGHT = 800
NAVIGATION_TIMEOUT_MS = 30_000
ATTEMPTS = 3
def safe_filename(value):
value = re.sub(r"[^A-Za-z0-9._-]+", "-", value.strip()).strip(".-_")
return value[:100] or "website"
def valid_http_url(value):
parsed = urlparse(value)
return parsed.scheme in {"http", "https"} and bool(parsed.netloc)
def main():
OUTPUT.mkdir(parents=True, exist_ok=True)
rows = []
with INPUT.open(newline="", encoding="utf-8-sig") as source:
entries = csv.DictReader(source)
if not entries.fieldnames or not {"name", "url"}.issubset(entries.fieldnames):
raise SystemExit("ngos.csv must have name and url columns")
records = list(entries)
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(
viewport={"width": WIDTH, "height": HEIGHT},
device_scale_factor=1,
reduced_motion="reduce",
)
page = context.new_page()
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
for index, record in enumerate(records, start=1):
name = (record.get("name") or "").strip()
url = (record.get("url") or "").strip()
filename = f"{index:04d}-{safe_filename(name)}.png"
result = {"name": name, "input_url": url, "final_url": "", "file": "", "status": "error", "error": ""}
if not name:
result["error"] = "Missing NGO name"
rows.append(result)
continue
if not valid_http_url(url):
result["error"] = "URL must be an absolute http:// or https:// URL"
rows.append(result)
continue
for attempt in range(1, ATTEMPTS + 1):
try:
response = page.goto(url, wait_until="domcontentloaded")
# Give initial styles and images a brief opportunity to render.
page.wait_for_timeout(1000)
target = OUTPUT / filename
page.screenshot(path=str(target), type="png", full_page=False, animations="disabled")
result["final_url"] = page.url
result["file"] = str(target)
result["status"] = "ok"
if response is not None and response.status >= 400:
result["error"] = f"HTTP {response.status}; screenshot saved, inspect before publishing"
break
except PlaywrightTimeoutError as exc:
result["error"] = f"Navigation timeout on attempt {attempt}: {exc}"
except Exception as exc:
result["error"] = f"Attempt {attempt}: {type(exc).__name__}: {exc}"
if attempt < ATTEMPTS:
time.sleep(attempt * 2)
rows.append(result)
print(f"{result['status']}: {name} - {result['error'] or result['file']}")
context.close()
browser.close()
with REPORT.open("w", newline="", encoding="utf-8") as destination:
writer = csv.DictWriter(destination, fieldnames=["name", "input_url", "final_url", "file", "status", "error"])
writer.writeheader()
writer.writerows(rows)
failures = sum(row["status"] != "ok" for row in rows)
print(f"Finished {len(rows)} entries; {failures} need attention. Report: {REPORT}")
return 1 if failures else 0
if __name__ == "__main__":
sys.exit(main())
Run it from the folder containing the CSV and script:
python make_thumbnails.py
Files appear under thumbnails/, and thumbnails/results.csv lets you retry only the rows that need attention. The script deliberately uses an initial viewport and PNG. Change WIDTH and HEIGHT to match your card layout; choose one size for the whole set.
4. Choose capture settings for a useful thumbnail set
| Choice | Use it when | Trade-off |
|---|---|---|
| Viewport (script default) | You want comparable directory cards. | Content below the fold is omitted. |
| Full page | You need a page record or long-form review. | Image heights differ and may be unwieldy. |
| PNG | Text and sharp edges matter. | Files can be larger than lossy formats. |
| JPEG or WebP | Storage or transfer size matters and lossy output is acceptable. | Compression can soften small text or detail. |
| CSS scale | You want output dimensions to track CSS pixels. | It may look less sharp on high-density displays. |
| Device scale | You want high-density pixel output. | Files use more pixels and storage. |
Playwright screenshot options include full-page capture, file type, quality for JPEG/WebP, and CSS or device scale; consult the current API reference for the option details supported by your installed version. Keep the choice fixed across the list to make previews comparable.
Loading and dynamic pages
The sample waits for domcontentloaded and then one second. This is a practical compromise, not a guarantee that every site’s fonts, images, animations, or client-rendered content are ready. Increase the delay for a known slow site, or wait for a page-specific selector when you control the target and know which element signals readiness. Avoid relying on networkidle for every site: pages with ongoing analytics or long polling may never become idle. A missing lazy-loaded image may require scrolling it into view or using full-page capture; inspect the resulting files rather than assuming one wait rule works universally.
Consent banners and overlays
The local script captures what the browser renders. A cookie or consent dialog, newsletter prompt, chat widget, or other overlay can cover the page. Do not automatically accept consent on sites unless that action is appropriate for your collection process. For a repeatable directory, note affected rows and decide whether to keep the initial visitor view, configure a site-specific dismissal where permitted, or use a capture workflow that handles known overlays.
5. cURL, Python requests, and Node.js alternatives
These examples call Playwright directly in Node.js or show an HTTP request only where an API exists. Playwright itself is a browser automation library, not a screenshot REST endpoint, so cURL cannot invoke the local Playwright script. The cURL example below uses ScreenshotNeo, a hosted screenshot API.
Node.js with Playwright
Install the package and browser with npm install playwright and npx playwright install chromium. Save as capture.mjs; the loop reads the same CSV using a small CSV parser for clarity in production workflows.
import { chromium } from 'playwright';
import fs from 'node:fs/promises';
const urls = [
{ name: 'Example NGO', url: 'https://example.org/' },
{ name: 'Another NGO', url: 'https://another-example.org/' },
];
const browser = await chromium.launch();
const context = await browser.newContext({
viewport: { width: 1280, height: 800 },
deviceScaleFactor: 1,
reducedMotion: 'reduce',
});
const page = await context.newPage();
page.setDefaultNavigationTimeout(30_000);
await fs.mkdir('thumbnails', { recursive: true });
for (const [index, item] of urls.entries()) {
const filename = `thumbnails/${String(index + 1).padStart(4, '0')}.png`;
try {
await page.goto(item.url, { waitUntil: 'domcontentloaded' });
await page.waitForTimeout(1000);
await page.screenshot({ path: filename, type: 'png', fullPage: false, animations: 'disabled' });
console.log(`Saved ${item.name}: ${filename} (${page.url()})`);
} catch (error) {
console.error(`Failed ${item.name} (${item.url}):`, error.message);
}
}
await context.close();
await browser.close();
Python requests
requests can download a screenshot from an HTTP screenshot service, but it does not render web pages or create screenshots by itself. For local browser capture, use the Playwright script above. For a managed alternative, use the Python API request in the next section.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API takes one GET request per URL and returns an image or PDF. The code below saves a WebP screenshot; see the ScreenshotNeo API documentation for options and response details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.org/ -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.org/"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.org/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status} ${await res.text()}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
For a list, repeat the call for each URL or use the bulk capture option (up to 100 URLs per call). ScreenshotNeo removes cookie/consent banners from 60+ known platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000, and every feature is available on every plan. Sign up for 1,000 free screenshots a month with no card.
6. Batch reliability, performance, and cost
- Retries: retry timeouts and transient navigation failures with a small bounded count and backoff, as shown. A retry cannot fix a persistent DNS error, blocked site, invalid certificate, or access restriction.
- Concurrency: start sequentially for a modest list. Parallel browser pages can reduce elapsed time but consume more memory and create more requests to the sites. Add a conservative concurrency limit, handle per-item exceptions, and respect each site’s terms and load.
- Storage: PNG preserves crisp detail but can take more space. If storage matters, choose JPEG/WebP quality settings and check that text remains legible. Keep the results CSV with the images so failures and redirects remain traceable.
- Repeat runs: use stable names or a mapping file; index-prefixed filenames avoid collisions when NGO names normalize to the same text. For changing lists, consider a stable identifier column.
- Local cost: Playwright has no per-screenshot API charge in this workflow, but you provide compute, browser maintenance, storage, and operational time. A managed service trades local browser operations for provider pricing and data handling considerations.
- Managed-service review: before sending a list to any hosted provider, check current limits, pricing, retention, URL handling, regional behavior, and terms. A batch endpoint alone does not establish reliability or suitability for a particular collection.
ScreenshotNeo offers caching with a selectable TTL, bulk capture up to 100 URLs per call, and a usage API. Its current published plan prices are Free for 1,000 shots/month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Verify current details on its documentation before planning a recurring job.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Browser executable missing | The Python package is installed, but its browser was not downloaded. | Run python -m playwright install chromium. |
| Invalid URL or navigation error | Missing scheme, typo, DNS failure, or unreachable host. | Validate an absolute HTTP(S) URL and open it from the same machine. |
| Timeout | The server is slow, unreachable, or navigation waits longer than the configured limit. | Check reachability, raise the timeout selectively, retain bounded retries, and log failures. |
| Screenshot is blank or incomplete | Client-side rendering or assets were not ready after the short wait. | Wait for a known selector or add a measured delay; inspect final URL and page response. |
| Consent or chat overlay obscures page | The browser captured the normal rendered overlay. | Handle it only under your collection rules, or use a service whose documented cleanup matches your needs. |
| Repeated file overwritten | Names collapsed to the same safe filename. | Use the row index or a stable NGO identifier in every output filename. |
| HTTP error but image exists | The server returned an error page that was still screenshot-able. | Review the report and image; decide whether to keep it as an error-state preview. |
| Different-looking previews between runs | Responsive layouts, rotating content, animations, ads, or page updates changed. | Keep viewport and scale fixed, disable animations, and document unavoidable dynamic content. |
8. Before publishing the thumbnail directory
- Confirm every image corresponds to the intended organization and canonical URL.
- Review failures, redirects, error pages, consent overlays, and unexpected content manually.
- Keep dimensions and format consistent; crop only if the directory design requires it.
- Check image licensing, privacy, and site terms for your use case, especially before publishing captures or sending URLs to a third party.
- Do not interpret a successful screenshot as verification that the organization or its site content is current or authentic.
FAQ
Should I capture the homepage or a specific NGO page?
Use the homepage for a directory overview. Capture an about, programs, or contact page only when that page is the specific subject of your comparison, and record the chosen URL.
Can I create thumbnails from a spreadsheet?
Yes. Export columns for name and URL as CSV, or adapt the script to read your spreadsheet format. The important part is retaining a stable name or ID for each output.
Can I use these images as proof that a site was reviewed?
A screenshot records what the browser rendered at capture time. Keep the URL and timestamp in your own records if auditability matters; the sample report records URLs but does not add a timestamp.


