ScreenshotNeo

BlogHow-to

How to Check When a Website Was Last Updated

Check a page’s visible update date, structured metadata, sitemap, and Google crawl record—and learn what each signal can and cannot prove.

By the ScreenshotNeo team4 October 20262 min read

To check when a website page was last updated, first look for a visible “Last updated” or “Updated” label. Then compare it with the page’s structured date metadata and its entry in the site’s XML sitemap. If you have the site owner’s Google Search Console access, check the URL Inspection report too—but its last-crawl date is when Google observed the page, not when the publisher edited it.

These signals can disagree or be missing. Record what each one says and where it came from; do not present an inferred date as confirmed edit history. A homepage, article, sitemap, and search result can each show different dates because they describe different things.

1. Check the page itself

Open the exact page you care about, not just the site homepage. Look near the headline or byline, at the end of the article, and in the footer for a date labeled “Last updated,” “Updated,” or “Last modified.” A clearly labeled visible date is the publisher’s direct claim to readers, but it is still a claim.

Keep a publication date separate from an update date. “Published” usually describes when the page first appeared; it does not establish whether it was revised later. A date mentioned in the article may refer to an event, not to the page’s publication or revision.

2. Inspect the page’s structured date metadata

Open the page source or use your browser’s developer tools to search for datePublished and dateModified. They may appear in JSON-LD structured data, often alongside an article type. These values are useful to compare with the visible label, but machine-readable markup is still publisher-supplied information; its presence does not independently verify an edit.

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "datePublished": "2024-03-12T09:00:00Z",
  "dateModified": "2025-01-18T14:30:00Z"
}
</script>

Google recommends clear, consistent date information and says its systems use several factors to estimate a page’s date. It does not guarantee that a supplied byline date will appear in search results. See Google’s guidance on byline dates.

3. Find the page in the XML sitemap

Try the site’s conventional sitemap location, such as https://example.com/sitemap.xml, or inspect robots.txt for a sitemap reference. A sitemap can be an index that points to other sitemap files, so follow the relevant link if the page is not in the first file. Search within the XML for the exact page URL and inspect the <lastmod> value in that URL’s entry.

<url>
  <loc>https://example.com/articles/example/

The Sitemaps protocol defines <lastmod> as the modification date of the linked page, not the date the sitemap file was generated. Google says the value should reflect a significant page change, such as a change to main content, structured data, or links; a routine copyright-year change does not qualify. Google may use the value when it is consistently and verifiably accurate. See the Sitemaps protocol and Google’s sitemap guidance.

A sitemap value is still a site-provided signal. Do not assume it proves that the page’s meaningful content changed on that exact date.

4. Interpret search-result dates cautiously

A date displayed beside a search result is Google’s estimate of when it believes the page was published or significantly updated. Google considers multiple factors, and it does not promise to show the date supplied by the publisher. Treat a result date as a clue to compare with the page and sitemap, not as a definitive edit record. Google explains this in its byline date documentation.

5. Check Google’s last crawl if you have site access

In Google Search Console, open URL Inspection for the page. The report can show when Google last crawled that URL. This answers when Google observed the page, not when the publisher last edited it; the page could have changed before or after that crawl. Google documents the report’s crawl information in the URL Inspection tool help.

If you do not own or manage the site, you generally will not have access to its Search Console report. Do not substitute a search-result date for that private crawl record.

Compare the available signals

Signal What it tells you Access What to keep in mind
Visible “Last updated” label The update date the publisher presents to readers Ordinary page view It is a publisher claim; compare it when accuracy matters.
dateModified metadata A machine-readable modification date supplied in page markup Page source or developer tools Markup does not independently verify the date.
Sitemap <lastmod> The stated modification date for a particular URL XML sitemap Google’s use depends on consistent, verifiable accuracy and significant changes.
Search-result date Google’s estimate of publication or significant update Search results It is an estimate and may differ from a publisher-supplied date.
Search Console last crawl When Google last crawled the URL Site owner’s Search Console A crawl time is not an edit time.

Runnable ways to inspect a page and its sitemap

The following examples retrieve a page and search its HTML for common structured date fields. They also fetch a sitemap and print the entry for the exact URL. They do not prove that a date is correct; they help you collect the site’s published signals. Replace the example URL with the exact page. Sites can block automated requests, render metadata with JavaScript, or use sitemap indexes, so a missing result is not proof that no date exists.

cURL

# Fetch page HTML and show date-related markup.
curl -L --fail --silent --show-error 'https://example.com/articles/example/' \
  | grep -Eio '.{0,100}(datePublished|dateModified|Last updated|Updated).{0,180}'

# Fetch a sitemap and locate the page entry.
curl -L --fail --silent --show-error 'https://example.com/sitemap.xml' \
  | grep -A3 -B1 -F 'https://example.com/articles/example/'

This uses standard command-line tools. If grep finds nothing, inspect the full HTML or sitemap: the value may be encoded, formatted differently, or located in a child sitemap.

Python

Install the two dependencies with python -m pip install requests beautifulsoup4. Save as check_dates.py and run python check_dates.py.

import requests
from bs4 import BeautifulSoup
from urllib.parse import urlparse

page_url = "https://example.com/articles/example/"
site_root = f"{urlparse(page_url).scheme}://{urlparse(page_url).netloc}"
session = requests.Session()
session.headers["User-Agent"] = "Mozilla/5.0 (compatible; date-check/1.0)"

def get(url):
    response = session.get(url, timeout=20)
    response.raise_for_status()
    return response.text

html = get(page_url)
soup = BeautifulSoup(html, "html.parser")
print("Structured date fields:")
for tag in soup.find_all(attrs={"itemprop": ["datePublished", "dateModified"]}):
    print(tag.get("itemprop"), tag.get("content") or tag.get_text(" ", strip=True))
for script in soup.find_all("script", type="application/ld+json"):
    text = script.get_text(" ", strip=True)
    if "datePublished" in text or "dateModified" in text:
        print(text)

robots_url = site_root + "/robots.txt"
try:
    robots = get(robots_url)
    sitemap_urls = [
        line.split(":", 1)[1].strip()
        for line in robots.splitlines()
        if line.lower().startswith("sitemap:")
    ]
except requests.RequestException:
    sitemap_urls = []
if not sitemap_urls:
    sitemap_urls = [site_root + "/sitemap.xml"]

print("Sitemap matches:")
for sitemap_url in sitemap_urls:
    try:
        xml = get(sitemap_url)
    except requests.RequestException as error:
        print(f"Could not fetch {sitemap_url}: {error}")
        continue
    if page_url in xml:
        start = xml.find(page_url)
        print(xml[max(0, start - 100):start + len(page_url) + 180])
    elif "<sitemapindex" in xml or "

This lightweight script prints matching JSON-LD as text rather than fully parsing arbitrary nested JSON-LD graphs. For a robust crawler, parse XML with an XML parser, follow sitemap index entries, normalize trailing slashes and URL escaping carefully, and respect the site’s access rules.

Node.js

This example uses built-in Node.js modules and the built-in fetch available in current Node.js releases. Save as check-dates.mjs and run node check-dates.mjs.

const pageUrl = 'https://example.com/articles/example/';
const origin = new URL(pageUrl).origin;

async function get(url) {
  const response = await fetch(url, {
    headers: { 'user-agent': 'Mozilla/5.0 (compatible; date-check/1.0)' },
    signal: AbortSignal.timeout(20000),
  });
  if (!response.ok) throw new Error(`${url}: HTTP ${response.status}`);
  return response.text();
}

const html = await get(pageUrl);
console.log('Date-related markup:');
for (const match of html.matchAll(/(?:datePublished|dateModified|Last updated|Updated)[\\s\\S]{0,180}/gi)) {
  console.log(match[0].replace(/\\s+/g, ' '));
}

let sitemapUrls = [];
try {
  const robots = await get(`${origin}/robots.txt`);
  sitemapUrls = robots.split(/\\r?\\n/)
    .filter(line => /^sitemap:/i.test(line))
    .map(line => line.slice(line.indexOf(':') + 1).trim());
} catch {
  // Use the conventional location if robots.txt is unavailable.
}
if (sitemapUrls.length === 0) sitemapUrls = [`${origin}/sitemap.xml`];

console.log('Sitemap matches:');
for (const sitemapUrl of sitemapUrls) {
  try {
    const xml = await get(sitemapUrl);
    const index = xml.indexOf(pageUrl);
    if (index >= 0) {
      console.log(xml.slice(Math.max(0, index - 100), index + pageUrl.length + 180));
    } else if (xml.includes('

These examples are intentionally simple. HTML patterns vary, XML may be compressed, and a sitemap index needs its child files checked. If the date matters for an audit or citation, verify the result in the browser and preserve the exact page URL and the evidence you observed.

Record discrepancies instead of guessing

  1. Write down the exact page URL and when you checked it.
  2. Copy each date with its label and source: visible page, metadata, sitemap, search result, or Search Console crawl.
  3. Separate publication, modification, and crawl dates in your notes.
  4. If signals conflict, report the conflict and uncertainty. Do not select one date as definitive without independent edit-history evidence.

For example: “The page visibly says ‘Updated January 18, 2025’; its sitemap lists January 20, 2025; Google last crawled it February 2, 2025. These signals differ, and the crawl date does not establish the edit date.”

Save a visual record of the page

A screenshot can preserve what the page showed when you checked it, including a visible update label. It is useful documentation, but a screenshot does not establish when the publisher edited the page. Capture the relevant label and enough page context to identify the URL; keep the date and time of your observation separately.

For browser-based capture, use a full-page screenshot when the date or relevant content is below the fold. Check that lazy-loaded sections have appeared before saving, and retain the original URL alongside the image. Avoid treating a screenshot of a search result as direct evidence of the publisher’s edit date.

Or skip the browser setup

To keep a clean visual record while you inspect a page, ScreenshotNeo can return a screenshot with one request. It is a website screenshot API and MCP server from Yorker Media; it captures pages but does not determine or verify their update dates.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/articles/example/ -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/articles/example/"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/articles/example/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo and get 1,000 free screenshots a month, with no card.

Troubleshooting

Problem Likely cause What to do
No visible update date The publisher does not display one, or it is placed elsewhere. Check the byline, bottom of the article, and footer, then compare metadata and sitemap evidence.
No dateModified in the source The page may omit structured dates, use another format, or add content with JavaScript. Search for datePublished, inspect the rendered page with browser developer tools, and check the sitemap.
No sitemap at /sitemap.xml The site uses another path, has a sitemap index, or does not expose one. Check robots.txt for a Sitemap directive and follow any index files.
Sitemap has no matching URL The page may be omitted, use a canonical URL variant, or live in a child sitemap. Check URL spelling, scheme, trailing slash, redirects, and linked sitemap files.
Dates disagree Signals can be generated or updated at different times and describe different events. Preserve each date and label, distinguish crawl from edit, and state that the evidence conflicts.
Command or script gets HTTP 403/429 The site may restrict automated requests or rate-limit clients. Use the ordinary browser, slow down requests, and follow the site’s access policy. Do not evade access controls.
Script returns no metadata Markup may be dynamically rendered, encoded differently, or outside the simple pattern used. Inspect the rendered DOM or page source directly; treat the sample scripts as discovery aids, not complete parsers.

Performance, reliability, and cost

Manual inspection of one page is usually the quickest option and requires no tooling. Fetching a page and one sitemap is lightweight, but scripts can miss dynamically rendered markup, child sitemaps, or alternate URL forms. For a small number of pages, inspect the signals directly rather than building a crawler.

For repeatable research, save the exact URL, observation time, date labels, and a copy or screenshot of the page. A saved snapshot supports what you observed at that time; it does not prove the publisher’s underlying edit history. No signal in this workflow should be treated as a guaranteed record of every revision.

There is no special service cost to manually check the page or retrieve its public HTML and sitemap. Automated screenshot capture is optional and separate from date verification. ScreenshotNeo’s free allowance is 1,000 shots per month with no card; its paid plans start at $5 for 3,000 shots. Only use capture where a visual record is useful.

FAQ

Does “last updated” mean the whole website changed?

Usually it refers to the individual page where the label appears. Check the exact URL and label rather than assuming it describes the entire site.

Can I find the edit date for a page on a site I do not own?

Sometimes: the page or sitemap may publish a date. Search Console crawl details require access to the site’s property, and public signals do not necessarily reveal a verified edit history.

Which date should I cite if the signals conflict?

Cite the date together with its source and label, or explain the discrepancy. Do not turn a crawl date or an unverified metadata value into a confirmed edit date.

Does a screenshot prove when a page was updated?

No. It can preserve what was visible when you captured it. Record the observation time separately, and distinguish it from a publisher’s stated update date.