ScreenshotNeo

BlogHow-to

How to Build a Competitor SEO Screenshot Archive with URLs and Metadata

Build a repeatable competitor screenshot archive with a stable URL inventory, consistent captures, and a searchable metadata manifest.

By the ScreenshotNeo team4 October 202612 min read

A useful competitor SEO screenshot archive is a repeatable record of what a page visibly showed, at a specific time and under recorded capture conditions. Build it from a scoped list of competitors and page types, preserve both discovered and normalized URLs, capture pages with fixed settings, and index every success, redirect, and failure in a searchable manifest. Screenshots document visible page state; by themselves they do not explain why a competitor changed a page or whether the change affected rankings.

This guide builds a small archive with Python and Playwright, then shows how to maintain the URL inventory, capture metadata, review quality, and use public web archives as supplementary evidence.

1. Define what the archive should answer

Start with a question you can investigate by looking at pages. For example: “How have the pricing pages of these three competitors presented their plans over the last quarter?” A useful initial scope specifies:

  • Competitors: stable IDs and domains, including relevant subdomains.
  • Page classes: product, category, pricing, editorial, or another small set tied to your question.
  • Market context: language, region, and any location that can change the rendered page.
  • Device profile: desktop and mobile captures should be separate series.
  • Cadence: how often captures are useful for the question and feasible to review.

Keep the first batch small. Confirm that your URL rules and capture settings work before expanding it. Treat consent dialogs, personalization, animations, and asynchronous content as part of the observed state: record what appeared rather than silently removing it.

2. Discover and maintain the URL inventory

Seed the inventory from each competitor’s public XML sitemap and relevant internal links. Google documents sitemaps as a way to inform Google about pages, alongside guidance on URL structure, crawling, and page metadata. These sources help discover URLs; they do not guarantee that a competitor sitemap is complete or tell your archive which pages to select. See Google’s crawling and indexing documentation.

For every discovered URL, keep both the exact original and a normalized key. Normalization helps identify common variants such as scheme, host casing, trailing slash, and tracking parameters, while the original lets you inspect redirects and alternate forms. Do not overwrite the discovered URL with the normalized one. Do not use canonicalization signals as an instruction to discard a page from your research set.

Suggested inventory fields

  • competitor_id, domain, page_type, market, device_profile
  • requested_url and normalized_url
  • discovered_from (for example, sitemap or internal link) and discovered_at_utc
  • A review status such as queued, included, or excluded, with a short reason for exclusions

Normalization is a research convenience, not proof that two URLs serve the same page. Preserve query parameters if they may affect language, location, pagination, or page state; remove tracking parameters only from the deduplication key when you have checked that they are irrelevant. Keep the original URL in all cases.

3. Choose a capture method and fix its settings

For a custom archive, a browser automation script gives control over browser version, viewport, waiting, and file naming. Playwright supports viewport and full-page screenshots and screenshot options including format and scale. Use its Python screenshot guide and Page screenshot API reference for the current option details.

Keep the browser engine and version, viewport, device scale, capture mode, wait condition, and capture timing consistent within a comparison series. Record these settings on each attempt. A viewport shot records only the visible viewport; a full-page shot records the scrollable document. Full-page captures can be very tall, so choose them when below-the-fold content matters and keep the same choice across comparable captures.

Runnable Python example: URL queue, screenshot, and JSONL manifest

Install Python and Playwright, then install the Chromium browser with the commands below. Save the script as archive.py and run it. It reads one requested URL per line from urls.txt, writes captures under captures/, and appends one JSON record per attempt to manifest.jsonl. The example intentionally records navigation failures and does not invent page metadata when a field is unavailable.

python -m pip install playwright
python -m playwright install chromium
import asyncio
import json
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlsplit, urlunsplit

from playwright.async_api import async_playwright

INPUT = Path("urls.txt")
CAPTURE_DIR = Path("captures")
MANIFEST = Path("manifest.jsonl")
COMPETITOR_ID = "competitor-a"
PAGE_TYPE = "pricing"
MARKET = "en-US"
DEVICE_PROFILE = "desktop-1440x1000"
VIEWPORT = {"width": 1440, "height": 1000}
DEVICE_SCALE = 1
FULL_PAGE = True
NAVIGATION_TIMEOUT_MS = 45000


def utc_now():
    return datetime.now(timezone.utc).isoformat()


def normalize_url(url):
    """Conservative key: lowercase scheme/host, drop fragment, trim non-root slash."""
    parts = urlsplit(url.strip())
    host = (parts.hostname or "").lower()
    port = f":{parts.port}" if parts.port else ""
    netloc = host + port
    path = parts.path or "/"
    if path != "/":
        path = path.rstrip("/")
    return urlunsplit((parts.scheme.lower(), netloc, path, parts.query, ""))


def safe_key(url):
    parts = urlsplit(url)
    host = re.sub(r"[^a-zA-Z0-9.-]+", "_", parts.netloc)
    path = re.sub(r"[^a-zA-Z0-9.-]+", "_", parts.path.strip("/")) or "home"
    return f"{host}_{path}"[:140]


async def main():
    CAPTURE_DIR.mkdir(parents=True, exist_ok=True)
    requested_urls = [line.strip() for line in INPUT.read_text().splitlines()
                      if line.strip() and not line.lstrip().startswith("#")]
    seen = set()

    async with async_playwright() as p:
        browser = await p.chromium.launch()
        for requested_url in requested_urls:
            normalized_url = normalize_url(requested_url)
            if normalized_url in seen:
                continue
            seen.add(normalized_url)
            record = {
                "competitor_id": COMPETITOR_ID,
                "page_type": PAGE_TYPE,
                "market": MARKET,
                "device_profile": DEVICE_PROFILE,
                "requested_url": requested_url,
                "normalized_url": normalized_url,
                "final_url": None,
                "captured_at_utc": utc_now(),
                "http_status": None,
                "error": None,
                "page_title": None,
                "meta_description": None,
                "canonical_url": None,
                "browser_name": "chromium",
                "browser_version": browser.version,
                "viewport_width": VIEWPORT["width"],
                "viewport_height": VIEWPORT["height"],
                "device_scale": DEVICE_SCALE,
                "capture_mode": "full_page" if FULL_PAGE else "viewport",
                "wait_condition": "domcontentloaded",
                "screenshot_path": None,
                "notes": "Consent and page state captured as rendered; no elements removed.",
            }
            page = await browser.new_page(
                viewport=VIEWPORT,
                device_scale_factor=DEVICE_SCALE,
                locale=MARKET,
            )
            try:
                response = await page.goto(
                    requested_url,
                    wait_until="domcontentloaded",
                    timeout=NAVIGATION_TIMEOUT_MS,
                )
                record["final_url"] = page.url
                record["http_status"] = response.status if response else None
                record["page_title"] = await page.title()
                record["meta_description"] = await page.locator(
                    'meta[name="description"]'
                ).get_attribute("content", timeout=3000)
                record["canonical_url"] = await page.locator(
                    'link[rel="canonical"]'
                ).get_attribute("href", timeout=3000)
                filename = f"{safe_key(requested_url)}_{record['captured_at_utc'].replace(':', '').replace('-', '')}.png"
                output = CAPTURE_DIR / filename
                await page.screenshot(path=str(output), full_page=FULL_PAGE)
                record["screenshot_path"] = str(output)
            except Exception as exc:
                record["final_url"] = page.url
                record["error"] = f"{type(exc).__name__}: {exc}"
            finally:
                await page.close()
                with MANIFEST.open("a", encoding="utf-8") as manifest:
                    manifest.write(json.dumps(record, ensure_ascii=False) + "\n")
        await browser.close()


if __name__ == "__main__":
    asyncio.run(main())

The example uses domcontentloaded so it can capture pages that continue loading analytics or other background requests. If the content you care about appears later, wait for a specific selector or use a bounded delay and record that choice. Avoid an unbounded network-idle wait on pages with long-running requests. The metadata selectors may return an empty value; that is a valid observation, not a reason to guess.

Example URL queue

https://competitor.example/pricing
https://competitor.example/products/widget
https://competitor.example/category/analytics

Replace example domains with URLs you are authorized to access. The archive script is a capture example, not a high-volume crawler: add conservative concurrency, rate limits, retry limits, and operational review before scheduling large queues. Respect the target site’s access controls and applicable policies.

4. Store a searchable manifest

JSON Lines is convenient for appending one record per attempt. CSV works well when fields stay flat; SQLite is useful when you need indexed queries, joins, or concurrent readers. Keep the manifest as the authoritative index and the original screenshot as the evidence file. Never replace an original capture with an annotated comparison or diff.

  • competitor_id, domain, page_type, market, device_profile
  • requested_url, normalized_url, final_url
  • captured_at_utc, http_status, error
  • page_title, meta_description, canonical_url, when available
  • browser_name, browser_version, viewport_width, viewport_height, device_scale, capture_mode, wait_condition
  • screenshot_path, optional html_path or pdf_path, and notes for consent state or unusual rendering

Keep a separate record for each capture date and device profile. A path such as captures/competitor-a/pricing/2026-10-04T120000Z.png makes files understandable, while the manifest provides the searchable relationship between URL, time, settings, and outcome. Use UTC timestamps to avoid ambiguity across machines and daylight-saving changes.

5. Review archive quality before comparing pages

  1. Confirm every manifest row represents an attempt, including redirects, HTTP errors, timeouts, and other failures.
  2. For successful captures, confirm the image exists and the recorded final URL matches the browser destination.
  3. Check title and description fields only where the page exposed them; leave unavailable values empty.
  4. Compare captures with the same viewport, device scale, browser version, capture mode, wait condition, market, and approximate time context.
  5. Inspect consent dialogs, personalization, dynamic widgets, and partially loaded content. Record meaningful differences in notes.
  6. Keep raw images unchanged. Store side-by-side comparisons, annotations, or diffs separately.

When reviewing a change, describe what is visible: a headline changed, a pricing table moved, a CTA appeared, or a section was removed. Verify important details against the live URL and metadata. An image alone does not establish ranking movement, traffic, intent, or the outcome of an experiment.

6. Use the Wayback Machine as supplementary evidence

The Internet Archive can help find older captures, but its records can be missing or incomplete because of discovery gaps, access restrictions, robots.txt, owner exclusion, technical limits, or JavaScript-dependent pages. Its help documentation also explains that Save Page Now saves one specific page once; it does not archive a whole site, multiple pages, or schedule future captures. See the Wayback Machine help page.

For programmatic discovery, the Internet Archive documents an Availability API for checking for an accessible archived capture and a CDX server API for more complex capture queries. Check current behavior and usage constraints before depending on either in a production workflow. Treat a missing result as “no accessible capture found,” not proof that the page never existed.

7. Custom browser automation or a managed archive?

Choose based on control and operating effort. Custom Playwright capture provides control over browser, viewport, wait conditions, and local file handling, but you must build URL discovery and deduplication, scheduling, failure reporting, storage, retention, and change review. A managed service may reduce that implementation and retention work; compare the options against your required cadence, export needs, capture settings, portability, alerts, and cost.

Screenshot Archive advertises scheduled viewport and full-page captures, stored HTML/PDF and metadata, APIs, and change detection for competitor landing pages. Those are vendor claims and were not independently tested for this article. Verify its current plan limits, capture conditions, frequency, export, and retention directly before relying on or purchasing the service. For a screenshot API recommendation, ScreenshotNeo comes first here: it removes known consent banners, popups, and chat widgets before capture, bills only clean shots, and its paid plans start at $5 for 3,000 captures.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for parameters. This example captures a competitor URL to WebP; substitute a URL you are permitted to access.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

For a repeatable archive, save the returned image alongside your own URL inventory and manifest, and record the requested URL, final destination when available, capture time, device and other settings, and outcome. Sign up for 1,000 free screenshots a month with no card.

Performance, reliability, and cost

  • Throughput: Browser startup and page rendering usually dominate a small queue. Reuse a browser process, but isolate pages and record individual outcomes. Increase concurrency only after checking that it does not overload your runner or target sites.
  • Timeouts: Set navigation and selector timeouts. A timeout is an archive outcome; record it and decide whether a later retry is appropriate. Use bounded retries with a delay rather than retrying indefinitely.
  • Large pages: Full-page screenshots use more time, memory, and storage than viewport shots. Capture only the mode needed to answer the research question, and keep that mode consistent.
  • Reproducibility: Browser updates, geolocation, locale, page personalization, consent state, and dynamic content can change what appears. Pin or record browser versions and settings; annotate unavoidable differences.
  • Storage: Estimate cost from capture frequency, number of URLs, device profiles, image dimensions, and retention. Preserve original files and back up the manifest; do not assume a vendor’s retention or export behavior without checking its current terms.
  • Service billing: ScreenshotNeo bills only clean shots; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Its free tier is 1,000 monthly shots, then listed paid tiers are Starter $5/3,000, Growth $15/15,000, Pro $39/60,000, Scale $99/250,000, and Business $249/1,000,000; yearly billing gives two months free.

Troubleshooting

Symptom Likely cause Fix
Navigation times out The page is slow, keeps connections open, or waits on resources unrelated to the visible content. Use a bounded timeout and a wait condition suited to the page, such as domcontentloaded; wait for a specific content selector when needed. Record the wait choice and timeout.
Screenshot is blank or incomplete The page renders content after navigation, requires interaction, or blocks automated access. Inspect the final URL and page state, wait for the relevant selector, and record bot checks or failures rather than treating the capture as a normal page. Do not claim missing content was present.
Title or description is empty The page lacks that metadata, injects it later, or the selector did not resolve. Leave the field empty when unavailable; if it is client-rendered, wait for an appropriate signal and record the method.
Duplicate URL entries persist Query parameters, slash variants, or host variants have not been reviewed consistently. Compare normalized keys, decide which parameters matter to the research question, and retain every exact discovered URL for inspection.
Two captures look different despite unchanged content Viewport, browser version, device scale, consent state, personalization, animation, or capture timing differs. Compare recorded settings, use a consistent capture series, and note unavoidable state differences.
Wayback has no usable page The page may not have been captured, may be incomplete, or may depend on restricted or JavaScript content. Try the Availability or CDX APIs where appropriate and check the live page or your own subsequent captures; absence in Wayback is not proof of nonexistence.

FAQ

Does a screenshot prove a competitor changed its SEO strategy?

No. It records visible page state. It cannot establish intent, rankings, traffic, or the results of an experiment.

Should desktop and mobile captures share one archive record?

Keep each capture attempt as its own record and identify its device profile. That makes comparisons and queries unambiguous.

Is a sitemap a complete list of a competitor’s pages?

No. Use it as one discovery source alongside relevant internal links, and preserve the discovery source for each URL.

Can Save Page Now schedule recurring competitor captures?

No. The Internet Archive describes it as saving one specific page once, not scheduling future captures or archiving a whole site.