ScreenshotNeo

BlogHow-to

How to Audit Open Graph Tags Across a Website

Audit Open Graph tags site-wide: inventory URLs, extract and validate metadata, check image access, group template issues, and retest previews.

By the ScreenshotNeo team4 October 202610 min read

To audit Open Graph tags across a website, first build a complete inventory of the URLs you intend to publish, then crawl each URL, extract its Open Graph metadata, validate required fields and image accessibility, group issues by page template, fix shared causes, and retest representative pages with the relevant social preview tools.

The Open Graph Protocol identifies four foundational properties for each page: og:title, og:type, og:image, and og:url. It describes og:description and og:site_name as optional but generally recommended; when a page specifies og:image, it should also specify og:image:alt. See the Open Graph Protocol.

1. Define the audit scope and build a URL inventory

A site-wide audit has two separate jobs: find all pages that should be checked, then validate their output. Decide which domains, subdomains, language or regional variants, and page types are in scope. Gather URLs from the XML sitemap and internal links, then compare that set with an explicit list of important pages such as products, categories, articles, campaigns, and landing pages.

  1. Export the XML sitemap URLs.
  2. Crawl internal links from the pages in scope.
  3. Combine and deduplicate those URL lists.
  4. Add known priority URLs that discovery may miss.
  5. Label sitemap-only and orphaned URLs so you can decide deliberately whether to keep, fix, redirect, or exclude them.

For a small site, a spreadsheet and manual spot checks can be enough. At scale, a general crawler can discover pages, export page data, search source code, and extract custom HTML fields. A specialized Open Graph audit can package page-level checks, preview-image checks, issue rollups, and preview cards, but its documentation says it is not a complete technical SEO crawler. Compare tools on coverage controls, custom extraction, JavaScript rendering, image retrieval checks, preview simulation, exports, repeatability, and crawl limits.

2. Choose how to crawl and extract the tags

Make one report row per URL. Capture the HTTP status and final URL after redirects, plus the metadata and context needed to diagnose failures. A useful schema is:

Column Why it helps
Requested URL, final URL, HTTP status Identifies redirects, unreachable pages, and unexpected destinations.
HTML title and canonical link Provides context for comparing page identity and og:url.
og:title, og:type, og:image, og:url Checks the protocol’s four foundational page properties.
og:description, og:site_name Checks recommended optional context.
og:image:alt Checks the recommended image description when an image is set.
Duplicate instances of important properties Helps find conflicting output from templates, plugins, or components.
Image response, content type, dimensions, page type, notes Separates image retrieval and template issues from missing page tags.

These are recommended audit columns; a crawler may not provide all of them without custom extraction. For a general crawler, use its custom HTML extraction facility with CSS selectors, XPath, or regex as appropriate, and export the results. If metadata or links are generated client-side, compare raw HTML with rendered HTML. JavaScript rendering can reveal what a rendered browser sees, but it does not prove that every social crawler executes JavaScript in the same way.

Open Graph values are typically emitted as meta elements in the document head. For example:

<meta property="og:title" content="A useful page title">
<meta property="og:type" content="article">
<meta property="og:image" content="https://example.com/images/page-card.jpg">
<meta property="og:url" content="https://example.com/guides/page">
<meta property="og:description" content="A concise description of this page.">
<meta property="og:site_name" content="Example Site">
<meta property="og:image:alt" content="A description of the image.">

The sample uses an illustrative domain. In a real audit, extract the actual rendered values rather than assuming that the source template describes every published page.

3. Validate presence, values, and consistency

Flag missing foundational fields, then check whether values identify the page they describe. The protocol defines og:url as the object’s canonical URL and permanent graph identifier, so compare it with the URL strategy your site intends to publish. Check that titles, descriptions, and images are not accidentally repeated across unrelated pages.

  • Presence: flag missing og:title, og:type, og:image, or og:url.
  • Recommended metadata: note missing og:description and og:site_name separately from protocol-required properties.
  • Image description: when og:image is present, check for og:image:alt.
  • Identity: compare og:url with the intended canonical page address and investigate mismatches.
  • Uniqueness: identify pages with identical values where page-specific metadata is expected.
  • Duplicates: flag multiple declarations for the same property and inspect which component emits each one.

Do not treat one title length or image dimension as a universal protocol requirement. Platform displays and crops can differ. If you use a vendor’s image-size recommendation, label it as that vendor’s recommendation, then check actual previews for the services that matter to your audience.

4. Verify every referenced image separately

A present og:image tag does not establish that an external crawler can retrieve a usable image. Fetch each image URL independently and record the response status, redirects, content type, and whether the returned file is a valid image. Use an unauthenticated request where possible to approximate public retrieval. Opening the asset in your own logged-in browser is not conclusive evidence that an external crawler can fetch it.

Classify image issues so the responsible team can act on them:

  • Missing image property.
  • Relative or malformed image URL.
  • Non-success response or unexpected redirect.
  • Access restriction, such as authentication or request blocking.
  • Response content is not a usable image, or the image file is malformed.
  • Preview remains stale after the page or image changed.

Dimensions and aspect ratio are practical quality checks. OpenGraph.io’s audit documentation describes comparing declared dimensions and aspect ratio with its 1200×630 recommendation; that figure is a vendor recommendation, not a formal Open Graph Protocol requirement.

5. Group issues and fix their shared source

Sort findings by property, page template, CMS or plugin, and number of affected URLs. If dozens of pages share the same wrong title or image, investigate the template or metadata-generation logic before editing pages one by one. Check which template, CMS field, plugin, or component owns the output, make the shared correction, and then rerun the same URL inventory to compare results and catch regressions.

Keep the original report and the retest report. A simple before-and-after comparison makes it easier to verify that the correction reached every affected URL and did not break another page type.

6. Preview representative pages on the relevant platforms

Choose representative URLs for each page template and type. Inspect the extracted HTML, then use the relevant platform debugger or preview inspector when the displayed card disagrees with the metadata. Third-party checker guidance identifies Facebook Sharing Debugger and LinkedIn Post Inspector as retest options. Their current workflows and cache behavior can change, so consult the platform’s current official guidance for exact steps.

A crawler showing new values while a platform shows an old preview is consistent with separate preview caching, but does not by itself establish the cause. Confirm that the page and image are publicly retrievable, inspect the platform’s current preview tool, and retest after its available refresh process.

Approaches and tool tradeoffs

Approach Good fit Strengths Limits
General crawler such as Screaming Frog SEO Spider Technical teams needing a broad crawl and flexible extraction Crawl configuration, source search, custom extraction, exports, optional JavaScript rendering, and crawl comparison Open Graph-specific reporting may need custom extraction. The vendor page accessed in 2026 lists a free crawl limit of 500 URLs and a £199 annual license; limits and prices can change.
Specialized Open Graph audit such as OpenGraph.io Site Audit Teams wanting packaged Open Graph checks and reports Documented Open Graph and image checks, canonical checks, preview cards, issue rollups, and sitemap/internal-link discovery or a supplied URL list The vendor says it is not a replacement for a full technical SEO crawler; its documentation describes paid monitoring as weekly or monthly.
Manual source inspection plus a single-URL checker Small sites and focused debugging Quick view of tags and one preview Does not establish whole-site coverage unless every URL is inventoried and checked.

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It is useful alongside an Open Graph audit when you need screenshots of representative pages or want an AI agent to capture a page; it does not replace tag extraction or image URL validation. Learn about ScreenshotNeo.

Or skip the browser setup

For a visual record of a representative page, ScreenshotNeo can return a screenshot or PDF from one GET request. It captures the page, not an Open Graph audit report. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Replace the example URL with a page you want to capture. Cookie and consent banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

Troubleshooting common audit failures

Symptom Likely cause What to do
Tags are missing from the crawler export The extraction targets the wrong elements, or metadata appears only after client-side rendering. Inspect source and rendered HTML. Correct the CSS/XPath extraction or use the crawler’s rendered mode, then verify the relevant platform’s actual preview separately.
The same values repeat across many unrelated pages A shared template, CMS setting, plugin, or component emits generic or incorrect metadata. Group affected URLs by template, find the output owner, fix the common source, and recrawl the group.
og:image exists, but the preview has no image The image may be blocked, redirected unexpectedly, malformed, or served with unsuitable content. Fetch the asset without a logged-in session, check response and file type, and inspect the platform preview tool.
The crawler sees new tags but the platform card is old The platform may be showing a separately cached preview, or it may not be retrieving the same page output. Recheck page and image access, use the platform’s current debugger or inspector workflow, and allow for separate preview caching.
Important URLs are absent from the report Discovery missed orphaned pages, sitemap exclusions, or alternate host/language URLs. Compare crawl results with the sitemap and priority list; supply missing URLs explicitly or adjust crawl scope and discovery.
Two tags for the same property disagree Multiple templates, plugins, or components emit the property. Inspect the rendered head, identify each emitting component, and remove or correct the conflicting source.

Performance, reliability, and cost considerations

  • Control crawl load: set a scope and crawl configuration appropriate to the site, and avoid repeatedly crawling unrelated URLs while diagnosing one template.
  • Use repeatable coverage: keep a stable URL inventory and compare each retest against it; new discoveries should be reviewed and added deliberately.
  • Separate page checks from image checks: fetching every image adds work but catches failures a metadata-only extraction cannot reveal. Store image results so unchanged assets do not need needless repeat checks where your process supports it.
  • Choose rendering deliberately: rendered crawling can help expose client-generated metadata, but it costs more time and resources than checking raw HTML in many crawling workflows. Use it where source inspection shows it is needed.
  • Evaluate tool limits before a full run: check current crawl caps, monitoring frequency, export options, and pricing on the vendor’s own site. The cited Screaming Frog figures are vendor-published and may change.
  • Do not infer platform outcomes from tag presence: a successful crawl confirms what that crawler retrieved; it does not guarantee every platform retrieves or displays the same card.

FAQ

Do I need to check every URL or only templates?

Check every in-scope URL for coverage and values, then use representative URLs for deeper visual preview checks by template. Sampling previews alone cannot prove that URL inventory or page-specific metadata is complete.

Is a present og:image enough?

No. Independently verify that the image is publicly retrievable and usable, then inspect a platform preview if it does not appear as expected.

Does Open Graph require a particular title length or image size?

The protocol requirements summarized here do not set a universal title length or image size. Treat tool or platform dimensions as recommendations and validate how the target service displays the card.

Can a browser-rendered crawl prove what every social crawler will see?

No. Rendering helps inspect client-generated output, but platform crawlers can differ. Use the relevant platform’s preview or debugger for pages where the displayed result matters.