How to Extract Page Titles and Descriptions Across a Website
Build a reliable sitemap crawler that extracts, renders, audits, and exports every page title and meta description across your site.

To extract every page title and meta description, build a URL inventory from the XML sitemap, fetch each same-domain page, parse the HTML head, render pages whose metadata is injected by JavaScript, and export the raw values with quality flags. A dependable audit also records redirects, status codes, robots directives, canonical URLs, duplicate groups, and the source of each value.
The workflow below works for a small site or a large crawl. It starts with a runnable Python script, then explains rendering, discovery gaps, auditing rules, performance, reliability, cost, and failure recovery.
What you are extracting
The page title is the text inside the HTML <title> element. The description is the content value of <meta name="description">. Google documents both as page metadata and recommends specifying a title on every page (title-link documentation and snippet documentation).

| Field | What to store | Why it matters |
|---|---|---|
| requested_url | URL found in the sitemap or link graph | Preserves discovery provenance |
| final_url | URL after redirects | Shows where metadata actually came from |
| status | HTTP status code | Separates valid pages from errors |
| content_type | Response media type | Prevents parsing PDFs, images, or feeds as HTML |
| title_raw, title_normalized | Original and whitespace-normalized title | Supports editorial review and duplicate grouping |
| description_raw, description_normalized | Original and normalized description | Reveals repeated or missing descriptions |
| metadata_source | initial_html or rendered_dom |
Explains differences between HTTP and browser results |
| canonical, robots, lastmod | Head directives and sitemap timestamp | Helps explain indexing and crawl decisions |
| issue_flags | Missing, duplicate, boilerplate, blocked, and other findings | Creates an actionable audit queue |
1. Discover the complete URL inventory
Start with /sitemap.xml. It may be a sitemap index that points to several child sitemaps. Parse every <loc>, retain <lastmod> when present, remove fragments, and keep only permitted same-domain URLs. Sitemaps help search engines discover pages but do not replace normal crawling; Google describes both sitemap discovery and the role of last-updated information in its sitemap documentation.
If the sitemap is missing or incomplete, seed the crawl with the home page and follow canonical internal links. Record whether each URL came from sitemap, internal_link, or manual_seed. That field makes omissions explainable later.
Normalize URLs before fetching
- Remove URL fragments such as
#reviews. - Lowercase the hostname and remove the default port.
- Resolve relative sitemap entries against the sitemap URL.
- Decide whether trailing slashes are equivalent for your site.
- Drop tracking parameters such as
utm_sourcewhen they do not identify a separate page. - Keep meaningful query parameters for faceted or search pages only if your audit scope includes them.
- Reject non-HTTP schemes and external hosts unless explicitly allowed.
2. Run a server-rendered extraction in Python
This script reads a sitemap or sitemap index, fetches HTML with bounded concurrency, extracts metadata, and writes a CSV. It intentionally stores both raw and normalized values. Install dependencies with pip install requests beautifulsoup4.

import csv
import re
import time
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import parse_qsl, urlencode, urljoin, urlsplit, urlunsplit
import requests
from bs4 import BeautifulSoup
START = "https://example.com/sitemap.xml"
USER_AGENT = "MetadataAudit/1.0 (+https://example.com/contact)"
MAX_WORKERS = 8
TIMEOUT = (10, 45)
RETRIES = 3
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
def normalize_url(url):
p = urlsplit(url.strip())
query = [(k, v) for k, v in parse_qsl(p.query, keep_blank_values=True)
if not k.lower().startswith("utm_") and k.lower() not in {"fbclid", "gclid"}]
host = p.hostname.lower() if p.hostname else ""
port = p.port
netloc = host
if port and not ((p.scheme == "http" and port == 80) or (p.scheme == "https" and port == 443)):
netloc = f"{host}:{port}"
return urlunsplit((p.scheme.lower(), netloc, p.path or "/", urlencode(query), ""))
def fetch(url):
last_error = ""
for attempt in range(RETRIES):
try:
r = session.get(url, timeout=TIMEOUT, allow_redirects=True)
content_type = r.headers.get("content-type", "").split(";", 1)[0].lower()
row = {"requested_url": url, "final_url": r.url, "status": r.status_code,
"content_type": content_type, "fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
"title_raw": "", "description_raw": "", "metadata_source": "initial_html",
"canonical": "", "robots": "", "error": ""}
if "html" not in content_type or r.status_code >= 400:
row["error"] = f"not_html_or_http_{r.status_code}"
return row
soup = BeautifulSoup(r.text, "html.parser")
row["title_raw"] = soup.title.get_text(" ", strip=True) if soup.title else ""
tag = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "description"})
row["description_raw"] = tag.get("content", "").strip() if tag else ""
canonical = soup.find("link", rel=lambda v: v and "canonical" in v)
row["canonical"] = canonical.get("href", "").strip() if canonical else ""
robots = soup.find("meta", attrs={"name": lambda v: v and v.lower() == "robots"})
row["robots"] = robots.get("content", "").strip() if robots else ""
return row
except requests.RequestException as exc:
last_error = str(exc)
time.sleep(2 ** attempt)
return {"requested_url": url, "final_url": "", "status": "", "content_type": "",
"fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
"title_raw": "", "description_raw": "", "metadata_source": "initial_html",
"canonical": "", "robots": "", "error": last_error}
def parse_sitemap(url, seen=None):
seen = seen or set()
url = normalize_url(url)
if url in seen:
return []
seen.add(url)
r = session.get(url, timeout=TIMEOUT)
r.raise_for_status()
soup = BeautifulSoup(r.content, "xml")
if soup.find("sitemapindex"):
urls = []
for loc in soup.find_all("loc"):
urls.extend(parse_sitemap(urljoin(url, loc.get_text(strip=True)), seen))
return urls
return [normalize_url(loc.get_text(strip=True)) for loc in soup.find_all("loc")]
def norm(value):
return re.sub(r"\\s+", " ", value or "").strip()
urls = sorted(set(parse_sitemap(START)))
rows = []
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
jobs = [pool.submit(fetch, u) for u in urls]
for job in as_completed(jobs):
row = job.result()
row["title_normalized"] = norm(row["title_raw"])
row["description_normalized"] = norm(row["description_raw"])
rows.append(row)
title_groups = defaultdict(list)
desc_groups = defaultdict(list)
for row in rows:
if row["title_normalized"]:
title_groups[row["title_normalized"].lower()].append(row["requested_url"])
if row["description_normalized"]:
desc_groups[row["description_normalized"].lower()].append(row["requested_url"])
for row in rows:
flags = []
title = row["title_normalized"]
desc = row["description_normalized"]
if not title: flags.append("missing_title")
if not desc: flags.append("missing_description")
if title and len(title_groups[title.lower()]) > 1: flags.append("duplicate_title")
if desc and len(desc_groups[desc.lower()]) > 1: flags.append("duplicate_description")
if title.lower() in {"home", "homepage", "untitled"}: flags.append("boilerplate_title")
if "noindex" in row["robots"].lower(): flags.append("noindex")
if row["error"]: flags.append(row["error"])
row["issue_flags"] = ";".join(flags)
fields = ["requested_url", "final_url", "status", "content_type", "title_raw", "title_normalized",
"description_raw", "description_normalized", "metadata_source", "canonical", "robots",
"issue_flags", "fetched_at", "error"]
with open("metadata-audit.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=fields)
writer.writeheader()
writer.writerows(rows)
print(f"Wrote {len(rows)} rows to metadata-audit.csv")
The parser takes the first title and description values but preserves the full raw strings. If a malformed document contains multiple candidates, add a second extraction pass that stores every candidate for review rather than silently choosing one.
3. Handle JavaScript-rendered metadata
Run the initial HTTP parser first. If a page has a missing title or description, or your application is known to inject metadata after load, send only those URLs to a browser-rendering queue. Parse the rendered DOM and set metadata_source=rendered_dom. Google’s JavaScript SEO guidance explains why server-rendered HTML is straightforward while JavaScript can set or change the title and description (JavaScript SEO basics).
With Playwright, a minimal rendering worker looks like this:
from playwright.sync_api import sync_playwright
def rendered_metadata(url):
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(url, wait_until="networkidle", timeout=60000)
title = page.title()
description = page.locator('meta[name="description"]').get_attribute("content") or ""
browser.close()
return title.strip(), description.strip()
Use a separate queue because browser rendering consumes substantially more CPU and memory than HTTP requests. Set a navigation timeout, wait for a useful selector when the site has one, and record whether rendering timed out. Do not treat a timeout as proof that metadata is absent.
4. Audit quality without arbitrary length rules
Flag missing values, duplicate or near-duplicate titles, boilerplate such as “Home,” titles that are vague or excessively verbose, duplicate descriptions, metadata that does not match the visible main content, JavaScript-only values, and pages blocked or excluded by robots directives. Google recommends descriptive, concise, distinct titles and page-specific descriptions; it also says snippets and title links may be truncated as needed. Therefore, use length as a review signal rather than claiming a universal character limit.
Near-duplicate detection
Normalize case and whitespace, remove punctuation, and compare token sets or edit distance. Grouping “Pricing | Example” and “Pricing – Example” is useful, while preserving the original strings for editors. For boilerplate, identify a shared prefix or suffix that appears on most pages and inspect the variable portion.
Content-match review
Automated checks cannot decide whether a description accurately represents a page. Add a manual queue for pages whose title is generic, whose description contains a product name from another template, or whose metadata source is rendered only after JavaScript. Include the final URL and a short visible-text sample in the review export.
5. Fill discovery gaps safely
Compare the sitemap set with URLs found in internal canonical links. Crawl only same-domain links, respect robots directives and access controls, and cap the number of pages per host. Keep discovery provenance so an editor can tell whether a URL was in the sitemap, linked internally, or added manually. A robots meta instruction can be read only when the crawler can access the page; Google documents this limitation in its robots meta-tag guidance.
Do not assume every sitemap URL is a page. Filter out media, feeds, APIs, and redirects according to your audit scope. Conversely, do not discard parameterized pages automatically if your site intentionally indexes them.
6. Choose the right tool
| Approach | Best for | Trade-offs |
|---|---|---|
| HTTP client plus Beautiful Soup | Small sites and server-rendered metadata | Fast and inexpensive; misses client-injected values |
| Scrapy plus an HTML parser | Large recursive crawls | Provides scheduling, extraction, throttling, and retries; requires project setup |
| Browser renderer | JavaScript-heavy applications | Sees post-load metadata; higher CPU, memory, and failure surface |
| Commercial audit crawler | Scheduled reports and team workflows | Convenient exports; verify URL coverage and rendering behavior on a sample |
Scrapy’s documentation describes combining its crawler with Beautiful Soup in callbacks (Scrapy overview). For a large site, begin with HTTP extraction and render only the pages that need it.
7. Reliability, performance, and cost controls
- Concurrency: start with 4–8 workers, then increase only while the origin remains responsive. Bounded concurrency protects the site and your machine.
- Retries: retry connection resets and 5xx responses with exponential backoff. Do not blindly retry permanent 4xx responses.
- Timeouts: separate connect and read timeouts. Store timeout reason and attempt count.
- Caching: cache successful responses during an audit and use
ETagorLast-Modifiedfor later runs. - Rate limits: honor robots rules, crawl-delay policies where applicable, and explicit site-owner limits.
- Memory: stream exports and avoid retaining complete HTML for every page.
- Reproducibility: save fetch time, user agent, redirect chain, and parser version.
- Cost: HTTP requests are usually the cheapest path; browser sessions add compute. Rendering only missing or suspicious pages reduces both runtime and infrastructure spend.
8. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Zero URLs from sitemap | XML namespace, sitemap index, or blocked response | Check status/content type, parse XML, and recurse through <sitemap> entries. |
| Title is empty but browser shows one | JavaScript sets document.title |
Render the URL and mark the source as rendered_dom. |
| Description is empty | No tag, wrong name case, or client injection |
Match the name case-insensitively, then use a renderer if needed. |
| Many duplicate URLs | Fragments, tracking parameters, or slash variants | Normalize before deduplication and define meaningful query parameters. |
| 403 or 429 responses | Access control or excessive request rate | Respect site rules, lower concurrency, identify your user agent, and obtain permission where required. |
| Wrong page metadata | Redirects, stale cache, or SPA route fallback | Record final URL and status, disable stale cache, and wait for the route-specific render. |
| CSV has broken characters | Encoding mismatch | Write UTF-8 and open the file with an application that honors UTF-8. |
| Browser worker hangs | Never-ending requests or heavy scripts | Set navigation and overall job timeouts, block unnecessary resources, and capture a timeout flag. |
Or skip the browser setup
ScreenshotNeo provides a website screenshot API when you need the rendered page as an artifact while auditing metadata. One GET request returns a PNG, JPEG, WebP, or PDF; its browser handles the page instead of requiring you to maintain Playwright infrastructure. See the ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. You get 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
9. Export for editors and engineers
Keep one machine-readable CSV or database table, then generate focused queues: missing metadata, duplicate titles, duplicate descriptions, boilerplate, JavaScript-only metadata, blocked pages, and content-match review. Include one sample URL for each duplicate group. A useful row is:
url, final_url, status, title_raw, title_normalized, description_raw, description_normalized, metadata_source, canonical, robots, lastmod, duplicate_group, issue_flags, fetched_at
For repeat audits, compare runs by normalized URL. Report newly missing values, changed titles, changed descriptions, and pages that moved from initial HTML to rendered DOM. This turns a one-time crawl into a regression check.
FAQ
Should I crawl the sitemap or follow links first?
Use the sitemap as the first inventory, then supplement it with internal-link discovery when coverage matters. Keep the source for every URL.
Do I need a browser for every page?
No. Parse initial HTML first and render only pages with missing metadata or known client-side generation.
Should I enforce a title character limit?
Use length to prioritize review. Search results can truncate titles and snippets, so there is no universal fixed cutoff in the cited guidance.
How do I audit pages behind authentication?
Use an authorized session with explicit cookies or headers, isolate those URLs from public crawl results, and protect the exported metadata as sensitive site data.
What is the most important duplicate check?
Group normalized, case-folded titles and descriptions, then inspect templates that leave only a token such as an ID or category name different.


