How to Generate Website Thumbnails for a List of URLs with Selenium
Use Selenium WebDriver to capture a list of URLs as consistently framed PNG thumbnails, with per-page waits, unique filenames, and recoverable errors.
Use Selenium WebDriver to open each URL, wait for the page state you need, and call driver.save_screenshot(path) with a unique PNG path for each page. Set the browser viewport before the loop so the thumbnails share consistent framing. Catch failures per URL so one inaccessible page does not stop the rest of the batch.
1. Install Selenium and a browser
This example uses Python, Selenium WebDriver, and Chrome. Install Selenium in the environment that will run the job:
python -m pip install selenium
Install Chrome on the machine or container as well. Selenium Manager can commonly locate or obtain a compatible driver automatically, but the browser still needs to be available. In restricted or offline environments, install and configure the browser driver explicitly. See the Selenium WebDriver documentation.
2. Capture each URL as a PNG
Save this as thumbnails.py. It reads one URL per line from urls.txt, creates the output directory, names files with an index and a URL-derived label, and records failures without aborting the batch.
from pathlib import Path
from urllib.parse import urlparse
import re
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
INPUT_FILE = Path("urls.txt")
OUTPUT_DIR = Path("thumbnails")
VIEWPORT_WIDTH = 1280
VIEWPORT_HEIGHT = 800
PAGE_TIMEOUT_SECONDS = 30
def safe_name(url: str) -> str:
"""Make a readable, filesystem-safe label from a URL."""
parsed = urlparse(url)
label = parsed.netloc + parsed.path
label = re.sub(r"[^A-Za-z0-9._-]+", "-", label).strip("-._")
return (label[:80] or "page")
def main() -> None:
urls = [line.strip() for line in INPUT_FILE.read_text(encoding="utf-8").splitlines()
if line.strip() and not line.lstrip().startswith("#")]
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
options = Options()
options.add_argument("--headless")
options.add_argument(f"--window-size={VIEWPORT_WIDTH},{VIEWPORT_HEIGHT}")
# In Linux containers without a display, headless mode avoids requiring X.
driver = webdriver.Chrome(options=options)
driver.set_page_load_timeout(PAGE_TIMEOUT_SECONDS)
failures = []
try:
for index, url in enumerate(urls, start=1):
filename = f"{index:04d}-{safe_name(url)}.png"
output_path = OUTPUT_DIR / filename
try:
driver.get(url)
# 'complete' is a useful baseline for ordinary pages, but does not
# guarantee that application data, fonts, or lazy content is ready.
WebDriverWait(driver, PAGE_TIMEOUT_SECONDS).until(
lambda d: d.execute_script("return document.readyState") == "complete"
)
saved = driver.save_screenshot(str(output_path))
if not saved:
raise RuntimeError("WebDriver reported that the screenshot was not saved")
print(f"OK {url} -> {output_path}")
except Exception as exc:
failures.append((url, repr(exc)))
print(f"FAIL {url}: {exc}")
finally:
driver.quit()
if failures:
failure_file = OUTPUT_DIR / "failures.txt"
failure_file.write_text(
"\n".join(f"{url}\t{error}" for url, error in failures) + "\n",
encoding="utf-8",
)
print(f"{len(failures)} URL(s) failed; details: {failure_file}")
print(f"Finished: {len(urls) - len(failures)} saved, {len(failures)} failed")
if __name__ == "__main__":
main()
Put URLs in urls.txt, one per line:
https://example.com/
https://www.selenium.dev/documentation/
https://developer.chrome.com/
Run the batch with:
python thumbnails.py
The script creates PNGs under thumbnails/. The numeric prefix prevents filename collisions, including when two different URLs normalize to the same label. The PNG is a screenshot of the current browser window at the chosen viewport size.
3. Choose page readiness deliberately
A screenshot records the rendered state available at capture time. The document’s readyState reaching complete is a baseline, not proof that a single-page app has fetched its data or that images, fonts, animations, and lazy-loaded sections are ready. Choose a condition that reflects the page and the thumbnail you want.
Wait for a page-specific element
For a known site, wait until its main content appears. Replace the baseline wait in the loop with a selector appropriate to that site:
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
WebDriverWait(driver, 20).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "main article"))
)
Use a condition that indicates useful content is present, rather than waiting for an element that appears before its content is populated. If the URL list spans unrelated sites, a single selector may not work for all of them; use per-site rules or a conservative fallback.
Wait for a short delay when there is no reliable selector
A fixed delay can help with pages that render shortly after navigation, but it adds the delay to every URL and still cannot guarantee completion. Prefer a meaningful selector or application-specific signal where possible:
import time
time.sleep(2)
Lazy images and full-page captures
save_screenshot captures the current browser window. It does not by itself turn the result into a full-page screenshot or guarantee that below-the-fold lazy images have loaded. If your goal is a full-page image, use a browser-specific full-page capture method or a tool that exposes one. If the target is a component, Selenium also documents taking a screenshot of a specific element with the element screenshot method. See Selenium’s screenshot documentation.
For window-sized thumbnails, keep the viewport fixed and avoid scrolling unless you intentionally want a different framing. For full-page output, test how the chosen browser and capture method handle sticky elements, very tall pages, and lazy loading before processing a large batch.
4. Options that affect the output
| Choice | What it changes | Practical guidance |
|---|---|---|
| Viewport width and height | Responsive layout and visible framing | Set it before navigation and keep it constant for comparable thumbnails. |
| Headless or headed browser | Browser execution environment | Headless is convenient for jobs without a desktop. Keep browser mode consistent when visual reproducibility matters. |
| Readiness condition | Which rendered state gets captured | Use a page-specific element or app signal when available; a timeout is only a bound. |
| Output format | save_screenshot writes PNG |
Convert afterward if you need JPEG or WebP; account for quality and transparency when choosing a conversion. |
| Window or element | Whole viewport versus one component | Use a WebElement screenshot for a component; use the window method for a page thumbnail. |
| Browser session lifetime | Startup cost, state reuse, and isolation | One session for a batch is simple. Restart between groups if cookies, memory growth, or site state make isolation important. |
Chrome’s headless command-line documentation also demonstrates choosing a viewport with --window-size and bounding a one-off capture with --timeout. A timeout bounds waiting; it does not prove that all asynchronous content has finished rendering. See Chrome Headless documentation.
5. Make filenames and batch behavior reliable
- Use unique names. The example includes an index so duplicate URLs and similar URL paths do not overwrite each other. For repeatable reruns, consider writing each run into a dated or uniquely named directory.
- Keep an input manifest. Save the original URL list alongside results so each image can be traced to its source.
- Handle failures per URL. Navigation errors, timeouts, access restrictions, and file permission errors can affect individual items. Log the URL and error, continue where safe, and review the failure list.
- Always close the driver. A
finallyblock callsquit()even if the loop or input processing raises an error. - Choose session reuse intentionally. Reusing a browser avoids repeated startup for every URL and keeps the example simple, but pages may share cookies or other session state. Separate sessions provide more isolation at the cost of more browser startups. The cited documentation does not benchmark either approach.
6. Alternative ways to capture a list
Selenium is a good fit when the task needs WebDriver integration, a scripted URL loop, or browser interactions before capture. Chrome Headless offers a command-line screenshot for uncomplicated one-off URLs. Puppeteer is a JavaScript browser automation library that can take screenshots and generate PDFs. The right choice depends on the surrounding application and required interactions; the available sources do not establish a speed or price winner.
For a single basic capture, Chrome documents a command shaped like this:
chrome --headless --screenshot --window-size=412,892 https://developer.chrome.com/
See Chrome Headless for command options and Chrome’s Puppeteer overview for the JavaScript automation alternative. These approaches still need per-URL orchestration if you are processing a list.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request captures a URL as an image or PDF; see the API documentation for parameters and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Replace the example URL with the URL you want to capture. To process a list, make one request per URL and use a unique output filename as in the Selenium loop. Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, with no card required.
Performance, reliability, and cost
Each Selenium capture requires browser navigation and rendering. Page load time depends on the target site, network, browser, and chosen readiness condition; there is no source-backed throughput figure for this workflow. For large lists, measure your own representative pages, choose a sensible page-load timeout, and decide whether a single reused browser session or isolated sessions better fit the job.
Reliability depends on the pages as well as the script. Some pages require authentication, block automated access, fail intermittently, or render content through site-specific JavaScript. Do not treat a saved PNG as proof that the page loaded correctly: log navigation status where relevant and inspect a sample of outputs. If the list matters operationally, retain the URL-to-file mapping and retry only failures with a bounded retry policy.
Selenium and browser automation do not charge per screenshot as part of the code shown here, but operating the job has costs: compute, browser runtime, storage, network use, and engineering time. A screenshot API shifts browser operation to a service and has its own plan limits and pricing; compare those costs against your volume and setup requirements.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Chrome or driver does not start | Browser missing, incompatible driver, or restricted container | Install Chrome and configure a compatible driver. Check container permissions and use headless mode where there is no display. |
| Navigation times out | The site is slow, unreachable, or keeps network activity open | Set a deliberate page-load timeout. If the page has useful content despite background requests, define a page-specific readiness condition and decide whether to continue after the navigation timeout. |
| Screenshot is blank or incomplete | Capture ran before app content appeared, or the site denied/failed access | Wait for a meaningful element or app signal. Check the page and logs in a headed browser for diagnosis; do not assume a longer fixed delay solves every case. |
| Images are missing | Lazy loading, blocked requests, or capture before images load | For visible images, wait for them to complete loading where the page allows it. For below-the-fold content, use a tested full-page workflow that triggers or accommodates lazy loading. |
| Different runs look different | Responsive layout, browser or OS changes, fonts, timing, or dynamic content | Keep browser version, viewport, settings, and capture environment consistent. Playwright’s visual comparison guidance describes these environment sources of variation; see its snapshot documentation. |
| Files overwrite each other | Filename derived only from a non-unique label | Add an index, stable ID, or URL hash to every output path. |
| One bad URL stops the list | Exception handling wraps the whole batch rather than each item | Catch and log errors inside the per-URL loop, then close the driver in finally. |
| Output directory error | Directory does not exist or process lacks write permission | Create the directory before capture and check the job’s filesystem permissions and available disk space. |
FAQ
Does Selenium save screenshots as PNG?
Yes. Selenium’s documented Python save_screenshot helper saves a PNG for the current browser window.
Can I make one image for every URL?
Yes. The loop handles each URL separately and saves one uniquely named PNG per successful capture.
Does a page-load timeout mean the screenshot is complete?
No. It limits how long navigation waits. It does not establish that every asynchronous component or lazy image has finished rendering.
Can I capture one element instead of the whole viewport?
Yes. Selenium documents an element screenshot method; locate the desired element and save its screenshot when a component image is the goal.
Sources
- Selenium WebDriver documentation — navigation and screenshot workflow.
- Selenium screenshot documentation — window and element screenshots.
- Chrome Headless documentation — command-line screenshots, viewport sizing, and timeout.
- Chrome Puppeteer overview — browser automation with screenshots and PDF generation.
- Playwright snapshot documentation — environment factors that can affect screenshot comparisons.


