How to Generate Website Thumbnails for a Chrome Bookmarks Export
Turn a Chrome bookmarks HTML export into a folder of website screenshots with a URL-to-image manifest, bounded waits, and a retry log.
To generate website thumbnails for a Chrome bookmarks export, read the exported HTML file, collect each bookmark’s title, URL, and optional folder path, then visit each URL in a browser and save a screenshot at one consistent viewport size. Keep a manifest that maps every bookmark to its image, cap each page’s wait time, and record failures so a slow or blocked site does not stop the batch.
For a single URL, Chrome Headless can capture a screenshot from the command line. For a whole export, a browser automation script is easier to repeat and gives you per-URL timeouts, stable filenames, and an error log. This guide uses Python with Playwright and the Python standard library to parse the bookmarks file. Playwright can launch bundled Chromium or installed branded Chrome or Edge; its screenshot API supports viewport or full-page captures and image options. Playwright Page API.
What the workflow produces
The script below creates an output/ directory containing one PNG per bookmark, a manifest.json that preserves title, URL, folder path, and image path, and a failures.json file for bookmarks that could not be captured. It skips bookmarks without web URLs. A thumbnail is a screenshot of the page’s visible state at capture time; it is not a guaranteed clean or representative preview of every site.
1. Export bookmarks from Chrome
- In Chrome, open the Bookmarks manager from the browser menu.
- Use the manager’s menu to export bookmarks and save the HTML file, for example as
bookmarks.html. - Keep that file in a working directory and inspect a few entries before running a large batch. Exports can include folders, separators, and non-HTTP bookmarks; the script handles folders and skips non-web URLs.
2. Install the browser automation package
Use Python 3.9 or later. In a new project directory, create a virtual environment and install Playwright plus its Chromium browser:
python -m venv .venv
# macOS or Linux
source .venv/bin/activate
# Windows PowerShell
# .venv\\Scripts\\Activate.ps1
python -m pip install playwright
python -m playwright install chromium
Playwright’s installation and browser selection options are documented in its Python getting started guide. If you prefer installed Chrome or Edge, Playwright supports branded browsers as described in its browser documentation; change the launch line in the script to use the corresponding channel.
3. Save and run the batch script
Save this as make_thumbnails.py. It uses Python’s built-in HTML parser rather than relying on a particular Chrome export parser. It tracks folders while parsing, sanitizes filenames, bounds navigation and screenshot waits, continues after failures, and writes JSON output in UTF-8.
#!/usr/bin/env python3
import argparse
import hashlib
import json
import re
from html.parser import HTMLParser
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright
class BookmarkParser(HTMLParser):
def __init__(self):
super().__init__(convert_charrefs=True)
self.items = []
self.folder_stack = []
self.pending_folder = False
self.current_anchor = None
self.in_anchor = False
self.in_h3 = False
self.h3_parts = []
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
if tag == "h3":
self.in_h3 = True
self.h3_parts = []
elif tag == "a":
self.in_anchor = True
self.current_anchor = {"url": attrs.get("href", ""), "title_parts": []}
elif tag == "dl":
# A DL immediately after an H3 is that folder's child list.
if self.pending_folder:
self.folder_stack.append(self._pending_name)
self.pending_folder = False
else:
self.folder_stack.append(None)
def handle_endtag(self, tag):
if tag == "h3" and self.in_h3:
name = "".join(self.h3_parts).strip()
self._pending_name = name or "Untitled folder"
self.pending_folder = True
self.in_h3 = False
elif tag == "a" and self.in_anchor and self.current_anchor is not None:
title = "".join(self.current_anchor["title_parts"]).strip()
self.items.append({
"title": title or self.current_anchor["url"],
"url": self.current_anchor["url"],
"folder": [name for name in self.folder_stack if name],
})
self.current_anchor = None
self.in_anchor = False
elif tag == "dl" and self.folder_stack:
self.folder_stack.pop()
def handle_data(self, data):
if self.in_h3:
self.h3_parts.append(data)
if self.in_anchor and self.current_anchor is not None:
self.current_anchor["title_parts"].append(data)
def safe_stem(title, url, index):
stem = re.sub(r"[^A-Za-z0-9._-]+", "-", title).strip(".-_")[:70]
digest = hashlib.sha256(f"{index}:{url}".encode("utf-8")).hexdigest()[:10]
return f"{index:05d}-{stem or 'bookmark'}-{digest}.png"
def main():
parser = argparse.ArgumentParser(description="Create website thumbnails from Chrome bookmarks HTML")
parser.add_argument("bookmarks_html", type=Path)
parser.add_argument("--output", type=Path, default=Path("output"))
parser.add_argument("--width", type=int, default=1280)
parser.add_argument("--height", type=int, default=800)
parser.add_argument("--timeout-ms", type=int, default=25000)
parser.add_argument("--settle-ms", type=int, default=800)
parser.add_argument("--full-page", action="store_true")
args = parser.parse_args()
if args.width < 1 or args.height < 1 or args.timeout_ms < 1:
parser.error("width, height, and timeout-ms must be positive")
source = args.bookmarks_html.read_text(encoding="utf-8", errors="replace")
reader = BookmarkParser()
reader.feed(source)
args.output.mkdir(parents=True, exist_ok=True)
manifest = []
failures = []
candidates = [item for item in reader.items if urlparse(item["url"]).scheme in ("http", "https")]
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(viewport={"width": args.width, "height": args.height}, device_scale_factor=1)
page = context.new_page()
for index, item in enumerate(candidates, start=1):
image_name = safe_stem(item["title"], item["url"], index)
image_path = args.output / image_name
try:
response = page.goto(item["url"], wait_until="domcontentloaded", timeout=args.timeout_ms)
# A short settle delay gives client-side rendering time without
# waiting indefinitely for analytics or long-lived connections.
if args.settle_ms:
page.wait_for_timeout(args.settle_ms)
page.screenshot(path=str(image_path), full_page=args.full_page, timeout=args.timeout_ms)
manifest.append({**item, "image": str(image_path), "final_url": page.url,
"http_status": response.status if response else None})
print(f"OK {index}/{len(candidates)} {item['url']} -> {image_path}")
except Exception as exc:
failures.append({**item, "error": str(exc)})
print(f"FAIL {index}/{len(candidates)} {item['url']}: {exc}")
context.close()
browser.close()
(args.output / "manifest.json").write_text(json.dumps(manifest, ensure_ascii=False, indent=2), encoding="utf-8")
(args.output / "failures.json").write_text(json.dumps(failures, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"Finished: {len(manifest)} captured, {len(failures)} failed, {len(reader.items) - len(candidates)} non-web entries skipped")
if __name__ == "__main__":
main()
Run it with the exported file name:
python make_thumbnails.py bookmarks.html --output thumbnails
# Optional: capture the full scrollable page instead of the initial viewport
python make_thumbnails.py bookmarks.html --output thumbnails-full --full-page
# Adjust a uniformly sized desktop preview and the per-site maximum wait
python make_thumbnails.py bookmarks.html --width 1440 --height 900 --timeout-ms 40000
The default capture waits for the initial document to be parsed and then allows a short settle period. This avoids making every bookmark wait for all network connections to go quiet: advertising, analytics, streaming, and other pages can keep connections open. Increase --settle-ms for sites that paint late, or change wait_until to networkidle for a known set of sites where that condition is useful. Playwright’s page navigation and screenshot behavior are described in the Page API.
4. Capture one URL with Chrome Headless
For a one-off thumbnail, Chrome’s command-line capture is enough. The Chrome Headless documentation says --screenshot saves a screenshot of the target page and shows using --window-size to choose the viewport. Use the installed Chrome executable name for your operating system:
chrome --headless --screenshot=thumbnail.png --window-size=1280,800 https://example.com
For a page that needs time to render, add a maximum timeout (milliseconds):
chrome --headless --screenshot=thumbnail.png --window-size=1280,800 --timeout=5000 https://example.com
See the Chrome Headless command-line reference for flags and current command syntax. The CLI is useful for checking a URL or debugging a single capture; scripting is more practical when you need to associate hundreds of output files with bookmark titles and track errors.
5. Choose the right capture settings
| Choice | Recommended default | When to change it |
|---|---|---|
| Capture area | Viewport screenshot | Use full page for long documents; it can create tall, large images and may expose lazy-load or sticky-header quirks. |
| Viewport | 1280 × 800 CSS pixels | Use a smaller viewport for compact bookmark tiles or a mobile width if that is how previews will be displayed. Keep it constant across the batch. |
| Readiness | domcontentloaded plus a short settle delay |
Wait for a known selector or network idle when the target set needs it. Network idle can hang or time out on sites with persistent requests. |
| Image format | PNG in the sample script | Use JPEG or WebP if your pipeline converts images and smaller files matter. Playwright supports PNG, JPEG, and WebP capture options in its API. |
| Device scale | 1 in the sample script | Increase for sharper high-density previews, knowing that pixel dimensions and storage grow. |
| Browser | Playwright’s installed Chromium | Use Chrome or Edge when matching a branded browser environment matters; keep browser and version consistent for repeatable output. |
6. Preserve bookmark identity and organize the output
The manifest keeps the source URL and title next to the generated image path, so your application can associate previews with bookmark records without decoding filenames. Folder names are included as an array. Filenames include a sequence number and a short URL-derived digest to reduce collisions when bookmarks have identical titles.
If you would rather create one subdirectory per folder, derive the directory from the folder array and sanitize every component before creating it. Avoid using raw bookmark titles or URLs as paths: titles can contain path separators, duplicate names, reserved characters, and long strings. The script keeps generated image names flat to avoid those filesystem edge cases.
7. Handle failures, redirects, and special cases
- Redirects: the manifest records both the original URL and the final browser URL. Keep the original as the bookmark identity; redirects can change over time.
- Non-web bookmarks:
chrome://,file:, bookmarklets, and other schemes are skipped by this version. Do not try to navigate to local files or browser-internal pages from an unattended batch unless that access is specifically intended. - Login and consent walls: the screenshot may show the sign-in or consent page. The browser context is fresh and does not reuse your personal Chrome profile or cookies.
- Bot checks and CAPTCHAs: a site may show a challenge or deny automation. Record it as a failure or accept that the resulting screenshot reflects the challenge. Do not treat automation as a way to bypass access controls.
- Lazy-loaded content: the viewport screenshot may omit images below the fold; full-page capture can trigger different page behavior. If full content matters, use a tailored scroll-and-wait strategy for the sites involved.
- Very long or animated pages: prefer viewport capture for bookmark cards. Full-page capture can use considerable memory and can make the output unwieldy.
- Duplicate URLs: this script captures every bookmark entry so folder and title relationships remain intact. To save time, deduplicate by URL, capture once, and point multiple manifest records at the shared image.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Executable doesn't exist or browser launch error |
Playwright’s Chromium browser was not installed for the active environment. | Run python -m playwright install chromium using the same Python environment that runs the script. |
| Navigation timeout | The site is slow, unreachable, blocked, or keeps requests active. | Keep a finite timeout, raise it for the relevant batch, retain the failure entry, and retry those URLs separately. Avoid an unbounded wait. |
| Blank or incomplete screenshot | The document navigated but the app rendered later, or a consent/login/challenge page is visible. | Increase settle delay or wait for a site-specific selector. Inspect the final URL and screenshot before changing global waits. |
| Connection refused, DNS, or certificate error | The target host is unavailable from the machine or network. | Open the URL in a regular browser from that environment, check DNS/proxy/firewall settings, and retry later if the site is temporarily down. |
| Wrong or duplicate filenames | Titles are empty, repeated, or contain filesystem-sensitive characters. | Keep the sanitization and digest scheme, and use manifest records rather than assuming a title uniquely identifies an image. |
| Memory use grows or browser crashes | Many pages or large full-page images are being retained or rendered. | The sample reuses one page sequentially and closes the browser afterward. Use viewport shots, reduce dimensions, and split very large exports into smaller runs. |
| Some bookmark folders appear misplaced | The HTML nesting differs from the parser’s simple folder tracking assumptions. | Check folder paths on a few manifest entries against the export. Adjust the parser for that file shape before processing the full set. |
9. Performance, reliability, and cost
This local workflow makes requests from your own machine and uses your installed browser automation package; there is no per-image service charge in the script. The main cost is elapsed time, bandwidth, and storage. A sequential run is gentler on browser memory and target sites, but total duration grows with the number of bookmarks and each page’s render time. A bounded timeout ensures one unresponsive URL cannot hold the batch indefinitely.
For reliability, preserve the input export, manifest, and failure list. Retry only failed rows instead of recapturing everything, and inspect a small sample at the intended display size before generating a large set. Screenshots depend on live site state, geography, personalization, consent UI, browser version, and load timing; the same URL can produce a different image on a later run.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Send a URL with one GET request and save the returned image. Its clean-shot steps can accept a consent banner like a visitor and remove 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report page verdict and billing status.
Use your API key in place of YOUR_API_KEY. See the ScreenshotNeo API documentation for request options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
For a bookmark export, call the endpoint once per URL and store each returned image beside the same title-and-URL manifest used above. ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
Frequently asked questions
Does this change my Chrome bookmarks?
No. It reads the exported HTML file and writes image and JSON files to the output directory.
Will the thumbnails stay current when a website changes?
No. They are point-in-time captures. Rerun the batch or retry selected manifest entries when you want refreshed previews.
Can I use the output in a bookmark dashboard?
Yes. Load manifest.json and associate each image path with its bookmark title, URL, and folder array.
Should every page be captured full-page?
Usually not for thumbnail tiles. A consistent viewport tends to produce more uniform previews; use full-page screenshots when the entire document is needed.


