How to Make Website Thumbnails for a Curated List of Open Source Projects
Build a repeatable workflow for capturing, reviewing, and publishing consistent thumbnails for open source project websites.
Make a manifest of project slugs, homepage URLs, output paths, and optional capture selectors. Capture every site with the same browser viewport, save each image under a stable filename, then review it before publication. For directory cards, start with the first viewport or a deliberate hero element; full-page screenshots are usually too tall and dense for small cards.
This guide uses Playwright with Python. It also covers a command-line workflow, repeatable refreshes, image delivery, failures, and a hosted API option. Pick dimensions and crop rules from the actual directory design: there is no universal thumbnail size.
1. Choose a consistent thumbnail shape
Decide how thumbnails will appear in the directory before capturing them. Use the card’s aspect ratio and display size to choose a viewport and crop rule. A consistent viewport gives each project the same initial viewing area; a shared crop or fit rule in the directory’s image component keeps the cards aligned even when source sites use different layouts.
- First viewport: a good default for compact cards because it shows the homepage’s initial visual impression.
- Selected element: useful when a stable hero, logo area, or product preview better represents the project. Selectors can break when sites change, so inspect captures after refreshes.
- Full page: useful when the whole page is the subject. Long images often become difficult to read in small cards.
Keep desktop and mobile captures separate if both are needed; do not mix them in one thumbnail set unintentionally. Use stable, collision-free filenames such as project-slug.webp. Choose PNG, JPEG, or WebP based on your design and delivery needs, then resize or compress to suit the rendered card dimensions.
2. Keep the project list in a manifest
Store one row per project, including a stable slug and canonical homepage URL. An optional selector lets you target a particular page element. Keeping this data separate from the capture script makes additions and URL changes easy to review.
[
{
"slug": "datasette",
"url": "https://datasette.io/",
"output": "public/project-thumbnails/datasette.png"
},
{
"slug": "example-project",
"url": "https://example.org/",
"output": "public/project-thumbnails/example-project.png",
"selector": "main h1"
}
]
Validate that slugs are unique, URLs are valid HTTP or HTTPS addresses, and output paths stay inside the intended image directory. Treat the manifest as reviewed project data, not as arbitrary user input: a capture script that accepts uncontrolled URLs can be misused to make requests to internal services.
3. Capture thumbnails with Playwright and Python
Playwright can save page screenshots, full-page screenshots, screenshot buffers for post-processing, and selected-element screenshots. See the Playwright screenshots documentation. The script below reads the manifest, uses one viewport, captures either the page or an optional selector, and reports failures without deleting or replacing an existing image.
Install
python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
python -m playwright install chromium
Save as capture_thumbnails.py
import json
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
MANIFEST = Path("projects.json")
VIEWPORT = {"width": 1280, "height": 800}
NAVIGATION_TIMEOUT_MS = 30_000
def valid_http_url(value):
parsed = urlparse(value)
return parsed.scheme in {"http", "https"} and bool(parsed.netloc)
def main():
projects = json.loads(MANIFEST.read_text(encoding="utf-8"))
slugs = [project["slug"] for project in projects]
if len(slugs) != len(set(slugs)):
raise ValueError("Project slugs must be unique")
failures = []
with sync_playwright() as playwright:
browser = playwright.chromium.launch()
context = browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
page = context.new_page()
for project in projects:
slug = project["slug"]
url = project["url"]
output = Path(project["output"])
if not valid_http_url(url):
failures.append((slug, "URL must use http or https"))
continue
try:
response = page.goto(
url,
wait_until="domcontentloaded",
timeout=NAVIGATION_TIMEOUT_MS,
)
if response is not None and response.status >= 400:
failures.append((slug, f"HTTP {response.status}; existing image preserved"))
continue
# Allow initial page scripts to render. Adjust only for sites that need it.
page.wait_for_timeout(750)
selector = project.get("selector")
target = page.locator(selector).first if selector else page
if selector:
target.wait_for(state="visible", timeout=10_000)
output.parent.mkdir(parents=True, exist_ok=True)
target.screenshot(path=str(output), animations="disabled")
print(f"Saved {output}")
except (PlaywrightTimeoutError, Exception) as error:
failures.append((slug, f"{type(error).__name__}: {error}; existing image preserved"))
context.close()
browser.close()
if failures:
print("\nCapture failures:")
for slug, reason in failures:
print(f"- {slug}: {reason}")
raise SystemExit(1)
if __name__ == "__main__":
main()
Run it with python capture_thumbnails.py. The fixed viewport and device scale factor make the capture setup consistent. The short post-navigation delay is only a starting point: a site may need a selector wait or a site-specific delay. The script skips HTTP error responses and failed captures, leaving any previous approved file in place. Review whether that is the right policy for your publishing workflow.
For full-page captures, replace the screenshot call with page.screenshot(path=str(output), full_page=True). For a selector, the example already uses locator.screenshot(). Use a stable selector, and expect to update it if the site’s markup changes.
4. Add a command-line or scheduled workflow
shot-scraper’s Release 0.14 documentation describes multi-shot capture and a GitHub Actions workflow that can capture configured screenshots and write them back to a repository. Its documentation is version-specific; check the current project documentation before copying commands into a new setup.
A scheduled job is worthwhile when the curated list changes often enough to justify browser maintenance. A practical workflow is:
- Run captures on a manual trigger or schedule.
- Write outputs to predictable paths derived from reviewed slugs.
- Fail or flag individual capture problems rather than substituting a blank page.
- Review image diffs before committing or publishing refreshed thumbnails.
- Keep the previous approved image until a replacement has been checked.
For a small list that changes rarely, a manual run and review can be simpler and more reliable than a scheduled job. Repository automation should make image changes visible to maintainers rather than silently updating production assets.
5. Handle dynamic pages and visual noise
Pages can render at different speeds, show animations, lazy-load images, or display consent banners and popups. No single wait condition works for every site. Start with DOM content loaded, then add a selector wait or a modest delay for a specific page when inspection shows it is needed. Avoid waiting indefinitely for network idle on pages with persistent analytics or streaming requests.
- Blank or incomplete image: check whether the page needs a later render point, a specific selector, or a longer timeout.
- Cookie or newsletter overlay: determine whether the site provides a legitimate way to dismiss it; record the behavior so future captures remain predictable.
- Animation or carousel changes: disable animations where possible, and capture at a consistent time or stable state.
- Lazy images: capture the viewport containing the relevant content, or use full-page capture if the whole page is required and confirm images have loaded.
- Region-specific content: be aware that locale, consent, or geolocation may change the rendered page. Keep browser settings consistent and document intentional variations.
6. Review, publish, and refresh responsibly
Automation makes it easy to capture a page; it does not decide whether the result is accurate or appropriate to republish. Review each new or changed image for blank states, broken assets, overlays, layout shifts, and misleading content. Add descriptive alt text that identifies the project or explains the image’s role in the directory. Provide a neutral fallback when capture fails or reuse is uncertain.
The cited capture documentation does not grant permission to republish third-party website visuals, logos, or trademarks. Check applicable project and site terms before publication, and do not assume that public visibility alone grants reuse rights. If permission is unclear, use a neutral placeholder or another asset you are authorized to use.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable missing | Playwright package is installed but its browser was not downloaded. | Run python -m playwright install chromium in the same environment. |
| Navigation timeout | The site is slow, unreachable, or keeps network requests open. | Check the URL; use domcontentloaded, set a finite timeout, and wait for a specific visible element when appropriate. |
| Screenshot is blank or shows a loading shell | Capture happened before client-side content rendered, or the page failed. | Inspect the page and wait for a meaningful selector or a site-specific delay. Do not publish a failed refresh over a good image. |
| Selector not found | The selector is incorrect, not present at this viewport, or the markup changed. | Inspect the current page, update the selector, and consider reverting to a page screenshot for that project. |
| Thumbnail has inconsistent framing | Different viewport sizes, device scale factors, or crop behavior were used. | Use one capture configuration and apply the same image fit and crop rule in the card component. |
| Output overwritten by an error page | The script saved a browser-rendered error or challenge page as if it were a valid capture. | Check response status and inspect captures before publication; preserve the previous approved image when a request fails. |
| Image files are too large | Full-page images or lossless output exceed the card’s actual needs. | Prefer a viewport or element capture for cards, then resize and select a suitable image format and quality. |
8. Performance, reliability, and cost
Each project requires browser navigation and rendering, so the main time cost scales with the number and responsiveness of sites. Use a finite navigation timeout, capture one page at a time initially, and add bounded concurrency only after checking browser and network limits. Reuse a browser process across pages, as the example does, instead of launching a new browser for every URL.
For reliability, retain the manifest, log failures by slug, and keep a last-known-good image until a new capture passes review. A scheduled job can fail because of source-site changes, network conditions, bot checks, or changed markup. A failed run should be visible and should not silently turn into a published blank thumbnail.
Self-managed Playwright and shot-scraper avoid a hosted screenshot request charge, but require you to operate the browser environment and maintain automation. Hosted APIs reduce browser infrastructure to operate, but have service pricing, authentication, cache freshness, retention, and reliability considerations. Store long-term assets yourself: OpenGraph.io documents that its returned screenshot URLs expire after 24 hours, so download or cache outputs intended for long-term use. See its screenshot API documentation for its documented capture controls. The cited page does not establish a price, so check current terms directly before selecting it.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It can capture full pages or selected elements and supports viewport, device, timing, and image settings. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
For a list, call the endpoint once per manifest entry and save each response using its stable slug; check the response and preserve approved images on failure. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Should thumbnails show a project’s full homepage?
Usually not for directory cards. A consistent first viewport or a representative element remains more legible at small sizes; use full-page images when readers need to inspect the entire site.
How often should I refresh captures?
Refresh when a project site changes or when the directory’s editorial process calls for an update. A schedule is useful only if someone reviews the resulting changes.
Can I publish a screenshot just because the site is public?
Public access does not by itself establish reuse permission. Check applicable terms or use a neutral fallback when rights are uncertain.
What is the best image format?
It depends on the target design and delivery pipeline. Choose a format supported by your site, inspect visual quality at the rendered card size, and optimize file size there.


