How to archive website screenshots with filenames based on page titles
Use Playwright to read a page title, make it filename-safe, and save a screenshot with collision-resistant metadata. Includes CLI and API options.
Read the page title after navigation, sanitize it for your filesystem, and pass the resulting path to Playwright’s screenshot method. Add a timestamp or unique ID so repeated captures do not overwrite one another, and store the source URL and capture settings as metadata.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
function safeTitle(title) {
const safe = title
.normalize('NFKC')
.replace(/[\\/:*?"<>|\u0000-\u001F]/g, '-')
.replace(/\s+/g, ' ')
.trim()
.replace(/[. ]+$/g, '')
.slice(0, 120);
return safe || 'untitled-page';
}
const url = process.argv[2] ?? 'https://example.com';
const archiveDir = 'archive';
await mkdir(archiveDir, { recursive: true });
const browser = await chromium.launch();
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
// If the site updates document.title client-side, wait for that update first.
const title = await page.title();
const timestamp = new Date().toISOString().replace(/[:.]/g, '-');
const filename = `${safeTitle(title)}-${timestamp}.png`;
const screenshotPath = path.join(archiveDir, filename);
await page.screenshot({ path: screenshotPath, fullPage: true });
await writeFile(path.join(archiveDir, `${filename}.json`), JSON.stringify({
url,
capturedAt: new Date().toISOString(),
title,
screenshot: filename,
viewport: { width: 1440, height: 900 },
fullPage: true
}, null, 2));
console.log(screenshotPath);
} finally {
await browser.close();
}
Install the dependency with npm install playwright, install a browser with npx playwright install chromium, then run node archive.mjs https://example.com. This creates a PNG and a JSON sidecar. The sidecar preserves provenance even if the title is empty, duplicated, or later changed.
1. Choose what the archive should preserve
A screenshot is a visual record of a rendering at capture time. It does not preserve page links, source text as structured content, scripts, network resources, or interactive behavior. If you need a page that can be replayed or preserved as a website, use a purpose-built web archiving workflow and retain the screenshot as a visual companion.
Decide whether each record should be a viewport shot or a full-page image. Playwright’s default captures the current viewport; fullPage: true captures the full scrollable page. Full-page captures can be very tall and may be less convenient to inspect or store.
2. Make title-derived filenames safe and unique
- Normalize Unicode: normalization reduces equivalent Unicode representations that otherwise produce visually identical but distinct names.
- Replace unsafe characters: the example replaces common forbidden path characters and control characters. Adapt this for the target operating system and any storage service.
- Trim whitespace and trailing periods: this avoids confusing names and filesystem compatibility issues.
- Set a length limit: the example limits the title portion to 120 characters. Filesystems limit encoded component and full path lengths, so a shorter limit may be needed for long archive paths or multibyte titles.
- Use a fallback: blank or punctuation-only titles become
untitled-page. - Avoid overwrites: the timestamp in the example is sufficiently granular for ordinary sequential captures. For concurrent jobs, add a UUID or another unique identifier and use exclusive file creation if overwriting is unacceptable.
Sanitizing a title is not the same as validating a URL or making a page safe to visit. Treat input URLs as untrusted in automated systems, and apply your normal network access controls.
3. Wait for the title and page state you need
page.title() reads the document title. Some applications set it only after client-side rendering or navigation. If that applies, wait for a known page condition before reading the title. A fixed delay can work as a last resort, but a selector or application-specific readiness condition is usually more reliable.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.waitForSelector('main article');
const title = await page.title();
Choose navigation readiness deliberately. Waiting for every network connection to stop can hang on sites that keep analytics or streaming connections open. Conversely, DOM readiness alone may precede image loading or client-side content. For visual completeness, wait for the relevant content and, where appropriate, images or fonts before capture.
4. Capture with the Playwright CLI
The CLI can capture a screenshot to an explicit filename and supports full-page capture. A title-derived name needs an extra step to read the rendered title, so use the script above when the title must be computed dynamically. The CLI’s documented default filename includes a timestamp when no filename is supplied.
npx playwright screenshot --full-page --filename=archive/example-page.png https://example.com
Use the CLI when a fixed filename or a shell-driven capture is enough. Use the script when you need sanitization, metadata, collision handling, or title-aware naming. See the Playwright CLI screenshots documentation.
5. Use shot-scraper when a command-line workflow fits
shot-scraper accepts an output path with -o; its Release 1.3 documentation says omitting the height produces a full-page screenshot. You still need a script or another step to read the page title and construct the output filename. It follows redirects.
shot-scraper https://example.com -o archive/example-page.png
See the shot-scraper Release 1.3 documentation for documented command options. The same documentation describes dumping final rendered HTML after JavaScript runs; that HTML output should not be treated as a complete replayable archive.
6. Keep an index and preserve provenance
Keep the original URL and capture time alongside the image rather than relying on a title-derived name as the only record. Useful index fields include:
- Original URL and final URL after redirects, if available.
- Capture timestamp in UTC.
- Original page title and saved filename.
- Viewport dimensions, full-page setting, and browser or rendering options relevant to interpretation.
- Capture outcome, such as success or navigation failure.
If the archive is shared or generated concurrently, store records in a database or append-only index and use unique filenames. Avoid using titles as directory paths: keep them as one sanitized filename component.
7. Troubleshoot common problems
| Symptom | Likely cause | Fix |
|---|---|---|
The file is named untitled-page |
The document title is blank when read. | Wait for the application’s title update or use a known selector/readiness condition before calling page.title(). |
| Two captures overwrite each other | The title is identical and the filename has no unique suffix. | Add a timestamp with enough precision, a UUID, or a stable page identifier. Ensure concurrent workers cannot choose the same path. |
| Invalid filename or path error | The title includes disallowed characters, trailing punctuation, or the resulting path is too long. | Sanitize separators and control characters, trim trailing periods/spaces, shorten the title, and keep the containing directory path short. |
| The screenshot is blank or incomplete | The page was captured before its relevant content rendered, or navigation failed. | Check navigation errors and final URL; wait for the main content selector and any required client-side update before capturing. |
| The page is cut off | The capture used the viewport instead of full-page mode. | Set fullPage: true. If the page is extremely long, consider viewport captures or section captures for practical review. |
| The full-page image omits lazy-loaded sections | Content loads only when scrolled into view. | Scroll through the page to trigger lazy loading, wait for relevant images, then take the full-page screenshot. |
| Navigation times out on an otherwise visible page | The selected load condition waits on persistent requests, or the site is slow. | Choose a less restrictive readiness condition, then wait explicitly for the content needed in the image. Set a suitable timeout and record failures instead of silently saving misleading output. |
| CLI output does not use the title | A static CLI filename does not read the page title automatically. | Use the Playwright script or add a separate title-reading step that passes a sanitized path to the CLI. |
8. Performance, reliability, and storage
Launching a browser for every URL adds overhead. For batches, reuse one browser and create a separate page or context per capture as appropriate; close pages and the browser in cleanup paths. Limit concurrency to what the host can support, since many full-page captures can consume substantial memory. Record failed navigations and retry only transient errors with a bounded retry policy.
Full-page PNGs can be large, especially at high viewport widths or on very tall pages. Choose a format and retention policy based on the archive’s use, and preserve capture settings so images remain interpretable. Keep stable metadata and do not infer that a successful image means the underlying page can be reconstructed.
Or skip the browser setup
ScreenshotNeo can return an image from one GET request, with URL and output options in the API documentation. Example using cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. The service returns the capture; for a title-based archive, save the response under a sanitized title-derived path and write your own URL and timestamp metadata.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Does the page title have to be unique?
No. Titles are labels, not unique identifiers. Add a timestamp or ID and retain the URL in metadata.
Does a screenshot archive preserve a website?
No. It records pixels at capture time. It does not preserve the page’s interactive behavior or all resources needed to reconstruct it.
Can I keep both the image and the rendered HTML?
Yes, if your workflow needs both. Treat rendered HTML as a separate artifact with its own limitations; the cited shot-scraper documentation does not describe its HTML output as a complete replayable archive.


