ScreenshotNeo

BlogHow-to

How to Extract Page Titles and Descriptions Across a Website

Build a reliable sitemap crawler that extracts, renders, audits, and exports every page title and meta description across your site.

By the ScreenshotNeo team29 September 202611 min read

How to Extract Page Titles and Descriptions Across a Website

To extract every page title and meta description, build a URL inventory from the XML sitemap, fetch each same-domain page, parse the HTML head, render pages whose metadata is injected by JavaScript, and export the raw values with quality flags. A dependable audit also records redirects, status codes, robots directives, canonical URLs, duplicate groups, and the source of each value.

The workflow below works for a small site or a large crawl. It starts with a runnable Python script, then explains rendering, discovery gaps, auditing rules, performance, reliability, cost, and failure recovery.

What you are extracting

The page title is the text inside the HTML <title> element. The description is the content value of <meta name="description">. Google documents both as page metadata and recommends specifying a title on every page (title-link documentation and snippet documentation).

A sitemap inventory feeds fetch, parse, render, and export stages.
A sitemap inventory feeds fetch, parse, render, and export stages.
Field What to store Why it matters
requested_url URL found in the sitemap or link graph Preserves discovery provenance
final_url URL after redirects Shows where metadata actually came from
status HTTP status code Separates valid pages from errors
content_type Response media type Prevents parsing PDFs, images, or feeds as HTML
title_raw, title_normalized Original and whitespace-normalized title Supports editorial review and duplicate grouping
description_raw, description_normalized Original and normalized description Reveals repeated or missing descriptions
metadata_source initial_html or rendered_dom Explains differences between HTTP and browser results
canonical, robots, lastmod Head directives and sitemap timestamp Helps explain indexing and crawl decisions
issue_flags Missing, duplicate, boilerplate, blocked, and other findings Creates an actionable audit queue

1. Discover the complete URL inventory

Start with /sitemap.xml. It may be a sitemap index that points to several child sitemaps. Parse every <loc>, retain <lastmod> when present, remove fragments, and keep only permitted same-domain URLs. Sitemaps help search engines discover pages but do not replace normal crawling; Google describes both sitemap discovery and the role of last-updated information in its sitemap documentation.

If the sitemap is missing or incomplete, seed the crawl with the home page and follow canonical internal links. Record whether each URL came from sitemap, internal_link, or manual_seed. That field makes omissions explainable later.

Normalize URLs before fetching

  • Remove URL fragments such as #reviews.
  • Lowercase the hostname and remove the default port.
  • Resolve relative sitemap entries against the sitemap URL.
  • Decide whether trailing slashes are equivalent for your site.
  • Drop tracking parameters such as utm_source when they do not identify a separate page.
  • Keep meaningful query parameters for faceted or search pages only if your audit scope includes them.
  • Reject non-HTTP schemes and external hosts unless explicitly allowed.

2. Run a server-rendered extraction in Python

This script reads a sitemap or sitemap index, fetches HTML with bounded concurrency, extracts metadata, and writes a CSV. It intentionally stores both raw and normalized values. Install dependencies with pip install requests beautifulsoup4.

Render only pages whose metadata appears after JavaScript runs.
Render only pages whose metadata appears after JavaScript runs.
import csv
import re
import time
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import parse_qsl, urlencode, urljoin, urlsplit, urlunsplit

import requests
from bs4 import BeautifulSoup

START = "https://example.com/sitemap.xml"
USER_AGENT = "MetadataAudit/1.0 (+https://example.com/contact)"
MAX_WORKERS = 8
TIMEOUT = (10, 45)
RETRIES = 3

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})

def normalize_url(url):
    p = urlsplit(url.strip())
    query = [(k, v) for k, v in parse_qsl(p.query, keep_blank_values=True)
             if not k.lower().startswith("utm_") and k.lower() not in {"fbclid", "gclid"}]
    host = p.hostname.lower() if p.hostname else ""
    port = p.port
    netloc = host
    if port and not ((p.scheme == "http" and port == 80) or (p.scheme == "https" and port == 443)):
        netloc = f"{host}:{port}"
    return urlunsplit((p.scheme.lower(), netloc, p.path or "/", urlencode(query), ""))

def fetch(url):
    last_error = ""
    for attempt in range(RETRIES):
        try:
            r = session.get(url, timeout=TIMEOUT, allow_redirects=True)
            content_type = r.headers.get("content-type", "").split(";", 1)[0].lower()
            row = {"requested_url": url, "final_url": r.url, "status": r.status_code,
                   "content_type": content_type, "fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
                   "title_raw": "", "description_raw": "", "metadata_source": "initial_html",
                   "canonical": "", "robots": "", "error": ""}
            if "html" not in content_type or r.status_code >= 400:
                row["error"] = f"not_html_or_http_{r.status_code}"
                return row
            soup = BeautifulSoup(r.text, "html.parser")
            row["title_raw"] = soup.title.get_text(" ", strip=True) if soup.title else ""
            tag = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "description"})
            row["description_raw"] = tag.get("content", "").strip() if tag else ""
            canonical = soup.find("link", rel=lambda v: v and "canonical" in v)
            row["canonical"] = canonical.get("href", "").strip() if canonical else ""
            robots = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "robots"})
            row["robots"] = robots.get("content", "").strip() if robots else ""
            return row
        except requests.RequestException as exc:
            last_error = str(exc)
            time.sleep(2 ** attempt)
    return {"requested_url": url, "final_url": "", "status": "", "content_type": "",
            "fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
            "title_raw": "", "description_raw": "", "metadata_source": "initial_html",
            "canonical": "", "robots": "", "error": last_error}

def parse_sitemap(url, seen=None):
    seen = seen or set()
    url = normalize_url(url)
    if url in seen:
        return []
    seen.add(url)
    r = session.get(url, timeout=TIMEOUT)
    r.raise_for_status()
    soup = BeautifulSoup(r.content, "xml")
    if soup.find("sitemapindex"):
        urls = []
        for loc in soup.find_all("loc"):
            urls.extend(parse_sitemap(urljoin(url, loc.get_text(strip=True)), seen))
        return urls
    return [normalize_url(loc.get_text(strip=True)) for loc in soup.find_all("loc")]

def norm(value):
    return re.sub(r"\\s+", " ", value or "").strip()

urls = sorted(set(parse_sitemap(START)))
rows = []
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
    jobs = [pool.submit(fetch, u) for u in urls]
    for job in as_completed(jobs):
        row = job.result()
        row["title_normalized"] = norm(row["title_raw"])
        row["description_normalized"] = norm(row["description_raw"])
        rows.append(row)

title_groups = defaultdict(list)
desc_groups = defaultdict(list)
for row in rows:
    if row["title_normalized"]:
        title_groups[row["title_normalized"].lower()].append(row["requested_url"])
    if row["description_normalized"]:
        desc_groups[row["description_normalized"].lower()].append(row["requested_url"])

for row in rows:
    flags = []
    title = row["title_normalized"]
    desc = row["description_normalized"]
    if not title: flags.append("missing_title")
    if not desc: flags.append("missing_description")
    if title and len(title_groups[title.lower()]) > 1: flags.append("duplicate_title")
    if desc and len(desc_groups[desc.lower()]) > 1: flags.append("duplicate_description")
    if title.lower() in {"home", "homepage", "untitled"}: flags.append("boilerplate_title")
    if "noindex" in row["robots"].lower(): flags.append("noindex")
    if row["error"]: flags.append(row["error"])
    row["issue_flags"] = ";".join(flags)

fields = ["requested_url", "final_url", "status", "content_type", "title_raw", "title_normalized",
          "description_raw", "description_normalized", "metadata_source", "canonical", "robots",
          "issue_flags", "fetched_at", "error"]
with open("metadata-audit.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=fields)
    writer.writeheader()
    writer.writerows(rows)

print(f"Wrote {len(rows)} rows to metadata-audit.csv")

The parser takes the first title and description values but preserves the full raw strings. If a malformed document contains multiple candidates, add a second extraction pass that stores every candidate for review rather than silently choosing one.

3. Handle JavaScript-rendered metadata

Run the initial HTTP parser first. If a page has a missing title or description, or your application is known to inject metadata after load, send only those URLs to a browser-rendering queue. Parse the rendered DOM and set metadata_source=rendered_dom. Google’s JavaScript SEO guidance explains why server-rendered HTML is straightforward while JavaScript can set or change the title and description (JavaScript SEO basics).

With Playwright, a minimal rendering worker looks like this:

from playwright.sync_api import sync_playwright

def rendered_metadata(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(url, wait_until="networkidle", timeout=60000)
        title = page.title()
        description = page.locator('meta[name="description"]').get_attribute("content") or ""
        browser.close()
        return title.strip(), description.strip()

Use a separate queue because browser rendering consumes substantially more CPU and memory than HTTP requests. Set a navigation timeout, wait for a useful selector when the site has one, and record whether rendering timed out. Do not treat a timeout as proof that metadata is absent.

4. Audit quality without arbitrary length rules

Flag missing values, duplicate or near-duplicate titles, boilerplate such as “Home,” titles that are vague or excessively verbose, duplicate descriptions, metadata that does not match the visible main content, JavaScript-only values, and pages blocked or excluded by robots directives. Google recommends descriptive, concise, distinct titles and page-specific descriptions; it also says snippets and title links may be truncated as needed. Therefore, use length as a review signal rather than claiming a universal character limit.

Near-duplicate detection

Normalize case and whitespace, remove punctuation, and compare token sets or edit distance. Grouping “Pricing | Example” and “Pricing – Example” is useful, while preserving the original strings for editors. For boilerplate, identify a shared prefix or suffix that appears on most pages and inspect the variable portion.

Content-match review

Automated checks cannot decide whether a description accurately represents a page. Add a manual queue for pages whose title is generic, whose description contains a product name from another template, or whose metadata source is rendered only after JavaScript. Include the final URL and a short visible-text sample in the review export.

5. Fill discovery gaps safely

Compare the sitemap set with URLs found in internal canonical links. Crawl only same-domain links, respect robots directives and access controls, and cap the number of pages per host. Keep discovery provenance so an editor can tell whether a URL was in the sitemap, linked internally, or added manually. A robots meta instruction can be read only when the crawler can access the page; Google documents this limitation in its robots meta-tag guidance.

Do not assume every sitemap URL is a page. Filter out media, feeds, APIs, and redirects according to your audit scope. Conversely, do not discard parameterized pages automatically if your site intentionally indexes them.

6. Choose the right tool

Approach Best for Trade-offs
HTTP client plus Beautiful Soup Small sites and server-rendered metadata Fast and inexpensive; misses client-injected values
Scrapy plus an HTML parser Large recursive crawls Provides scheduling, extraction, throttling, and retries; requires project setup
Browser renderer JavaScript-heavy applications Sees post-load metadata; higher CPU, memory, and failure surface
Commercial audit crawler Scheduled reports and team workflows Convenient exports; verify URL coverage and rendering behavior on a sample

Scrapy’s documentation describes combining its crawler with Beautiful Soup in callbacks (Scrapy overview). For a large site, begin with HTTP extraction and render only the pages that need it.

7. Reliability, performance, and cost controls

  • Concurrency: start with 4–8 workers, then increase only while the origin remains responsive. Bounded concurrency protects the site and your machine.
  • Retries: retry connection resets and 5xx responses with exponential backoff. Do not blindly retry permanent 4xx responses.
  • Timeouts: separate connect and read timeouts. Store timeout reason and attempt count.
  • Caching: cache successful responses during an audit and use ETag or Last-Modified for later runs.
  • Rate limits: honor robots rules, crawl-delay policies where applicable, and explicit site-owner limits.
  • Memory: stream exports and avoid retaining complete HTML for every page.
  • Reproducibility: save fetch time, user agent, redirect chain, and parser version.
  • Cost: HTTP requests are usually the cheapest path; browser sessions add compute. Rendering only missing or suspicious pages reduces both runtime and infrastructure spend.

8. Troubleshooting common failures

Symptom Likely cause Fix
Zero URLs from sitemap XML namespace, sitemap index, or blocked response Check status/content type, parse XML, and recurse through <sitemap> entries.
Title is empty but browser shows one JavaScript sets document.title Render the URL and mark the source as rendered_dom.
Description is empty No tag, wrong name case, or client injection Match the name case-insensitively, then use a renderer if needed.
Many duplicate URLs Fragments, tracking parameters, or slash variants Normalize before deduplication and define meaningful query parameters.
403 or 429 responses Access control or excessive request rate Respect site rules, lower concurrency, identify your user agent, and obtain permission where required.
Wrong page metadata Redirects, stale cache, or SPA route fallback Record final URL and status, disable stale cache, and wait for the route-specific render.
CSV has broken characters Encoding mismatch Write UTF-8 and open the file with an application that honors UTF-8.
Browser worker hangs Never-ending requests or heavy scripts Set navigation and overall job timeouts, block unnecessary resources, and capture a timeout flag.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API when you need the rendered page as an artifact while auditing metadata. One GET request returns a PNG, JPEG, WebP, or PDF; its browser handles the page instead of requiring you to maintain Playwright infrastructure. See the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. You get 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

9. Export for editors and engineers

Keep one machine-readable CSV or database table, then generate focused queues: missing metadata, duplicate titles, duplicate descriptions, boilerplate, JavaScript-only metadata, blocked pages, and content-match review. Include one sample URL for each duplicate group. A useful row is:

url, final_url, status, title_raw, title_normalized, description_raw, description_normalized, metadata_source, canonical, robots, lastmod, duplicate_group, issue_flags, fetched_at

For repeat audits, compare runs by normalized URL. Report newly missing values, changed titles, changed descriptions, and pages that moved from initial HTML to rendered DOM. This turns a one-time crawl into a regression check.

FAQ

Use the sitemap as the first inventory, then supplement it with internal-link discovery when coverage matters. Keep the source for every URL.

Do I need a browser for every page?

No. Parse initial HTML first and render only pages with missing metadata or known client-side generation.

Should I enforce a title character limit?

Use length to prioritize review. Search results can truncate titles and snippets, so there is no universal fixed cutoff in the cited guidance.

How do I audit pages behind authentication?

Use an authorized session with explicit cookies or headers, isolate those URLs from public crawl results, and protect the exported metadata as sensitive site data.

What is the most important duplicate check?

Group normalized, case-folded titles and descriptions, then inspect templates that leave only a token such as an ID or category name different.