ScreenshotNeo

BlogHow-to

How to troubleshoot Open Graph images blocked by robots.txt

Trace a missing link preview image to its exact og:image URL, serving host, crawler, and robots.txt rule, then retest the asset.

By the ScreenshotNeo team4 October 20268 min read

If an Open Graph preview is missing its image, test the exact URL in the page’s og:image metadata and the robots.txt file for the host that serves that image. A rule on the page’s host does not necessarily apply when the image comes from a CDN or subdomain. Check rules for the crawler you care about; a Google or Bing pass does not prove a social preview fetcher can retrieve the image.

1. Find the exact og:image URL

Start with the HTML that an ordinary browser or command-line client receives. Find each og:image declaration, then establish which image URL the preview consumer is receiving. If there are multiple declarations, do not assume which one a particular platform will choose: verify the actual metadata and the platform’s current behavior.

Follow redirects and record the final URL, including scheme, hostname, nonstandard port, path, and query string. The final image may be hosted somewhere different from the page itself.

curl -sS -L https://www.example.com/article \
  | grep -iE '<meta[^>]+property=["'"']og:image["'"'][^>]*>'

This quick command is useful for inspection, but parsing HTML with a regular expression is not reliable for every document. For pages whose metadata is generated dynamically or whose markup is unusual, inspect the rendered document in a browser or use an HTML parser. Confirm that the metadata returned to an unauthenticated visitor contains the expected absolute image URL.

2. Check robots.txt on the image’s serving host

Use the final image URL’s scheme and host to locate that authority’s /robots.txt. For example, an image at https://cdn.example.net/assets/preview.png is governed by the robots file at https://cdn.example.net/robots.txt, not automatically by https://www.example.com/robots.txt.

curl -sS https://cdn.example.net/robots.txt
curl -sS -I 'https://cdn.example.net/assets/preview.png'

Robots rules are scoped to an authority: protocol, host, and port matter. If the image URL redirects across hosts, inspect the robots file for each relevant image host and test the final URL as well. A subdomain, CDN, and origin can each have separate rules. RFC 9309 describes this authority-based protocol and asks conforming crawlers that successfully retrieve robots.txt to follow parseable rules: RFC 9309.

3. Match the rule group to the crawler

Read the user-agent groups for the crawler whose access you are investigating, then evaluate the exact image path against the applicable allow and disallow rules. Do not assume that permission for one named crawler applies to another. Google selects the most specific matching user-agent group, and Google explains robots.txt as a way to control which resources its crawlers may access: Google’s robots.txt documentation.

If you are diagnosing a social preview, identify that platform’s current fetcher from its official documentation or tool. The research for this article did not verify Meta’s current preview user-agent or exact Sharing Debugger procedure, so do not substitute a guessed crawler name or claim that a search crawler’s result proves social access.

4. Test with crawler-specific tools

Use a tester that identifies the crawler and URL being evaluated. Google Search Console’s robots.txt report can test Google access and help identify the robots file that affects a URL. Bing Webmaster Tools’ robots.txt tester accepts a URL and a selected crawler, including Bingbot. These tests are scoped to their crawler contexts; neither is a universal social-preview tester.

  • Confirm the tester is evaluating the image asset URL, not just the article URL.
  • Confirm it is evaluating the right authority, especially for a CDN or image subdomain.
  • Confirm the selected crawler matches the question you are trying to answer.
  • Read which rule matched, rather than treating a pass/fail result as a complete diagnosis.

Official references: Google Search Console robots.txt report and Bing Webmaster Tools robots.txt tester.

5. Correct a rule that blocks the image

If the applicable group disallows the image path, update the robots configuration controlled by the site or image host so the intended crawler can fetch that path. Prefer a narrow rule that permits the intended asset path when that expresses the policy; then retest the same final image URL with the relevant crawler.

Robots.txt is a crawler request protocol, not access control. It does not make a resource private. If an image must be private, protect it with actual access controls rather than relying on a robots rule. RFC 9309 documents the protocol’s scope and behavior: RFC 9309.

6. If robots.txt allows the image, check the response

A robots pass only addresses crawler rules. Separately verify that an unauthenticated request can retrieve the image, redirects resolve, and the response contains image data rather than a login page or HTML error. Check the final URL you recorded, not just the original URL.

curl -sS -L -D response-headers.txt \
  -o downloaded-image \
  'https://cdn.example.net/assets/preview.png'
file downloaded-image

Review the response headers and downloaded content. This is a practical check, not a statement of any platform’s current status-code, format, size, cache, or firewall requirements. The sources collected for this guide do not establish those requirements for Meta’s current preview crawler.

What to record while troubleshooting

Check Record Why it matters
Page metadata Every og:image value received by an ordinary client Confirms the asset URL you need to diagnose
Redirects Original URL and final image URL The final asset may be served by another authority
Robots scope Scheme, host, port, robots.txt URL, and image path Rules are authority and path specific
Crawler The exact crawler selected in the tester or platform documentation Different crawlers can match different user-agent groups
Rule result Matched group and allow/disallow decision Shows whether a specific robots rule is the blocker
Asset response Redirect chain and whether the response is actually image data A robots pass does not prove successful image retrieval

Common errors and fixes

Symptom Likely cause Fix
The article URL passes a robots test, but the preview has no image The image URL is on a different host or path and was never tested Read og:image, follow redirects, and test the final asset URL against its serving authority.
Google’s tester allows the image, but a social preview still fails The test covered Google’s crawler, not necessarily the social platform’s fetcher Verify the platform’s current crawler identity and diagnostics in its official documentation or tool.
A rule appears to allow the path, but the result is blocked You may be reading a different user-agent group, robots file, host, port, or path than the tester uses Check the exact authority, selected crawler, matched group, and final image path.
The robots test passes, but no usable image appears The request may redirect, require authentication, or return an HTML error instead of image content Fetch the final URL without credentials, inspect redirects and headers, and examine the downloaded file.
A noindex directive is expected to fix blocked crawling A crawler blocked by robots.txt cannot see a noindex directive on that blocked resource Decide separately whether the goal is crawl access or search indexing. Google explains this distinction in its robots.txt introduction.

Crawl access is different from search indexing

Robots.txt controls crawling, while noindex is an indexing directive that a crawler must be able to fetch and see. Google says a crawler blocked by robots.txt cannot see a noindex directive on the blocked resource. If the goal is to keep an otherwise accessible resource out of search results, robots.txt is not a substitute for noindex: Google’s robots.txt introduction.

Or skip the browser setup

For a repeatable image capture of a page, ScreenshotNeo provides a website screenshot API and MCP server. It does not diagnose which robots rule a social crawler used, so keep the crawler-specific checks above for that diagnosis. Its API can capture a page after loading it and return an image or PDF; see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://www.example.com/article \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://www.example.com/article"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://www.example.com/article'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot, and each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. All features are on every plan.

Sign up free for 1,000 screenshots a month with no card.

Performance, reliability, and cost notes

  • Keep diagnosis focused: start with the exact image URL and relevant host. Testing only the page URL can miss a CDN rule entirely.
  • Retest after a change: use the same final asset URL and crawler context so the before and after results are comparable.
  • Use access logs as another check: when available, ask the site or CDN operator to verify whether requests for that image reached the serving layer. This complements a robots tester; it does not identify a platform’s crawler unless the request can be attributed reliably.
  • Avoid broad edits: change only the path access needed for the intended crawler, then verify the resulting rule match.
  • There is no benchmark or universal tool result here: Google and Bing tools answer for their crawler contexts; a social platform may behave differently. Check its current official tooling for platform-specific troubleshooting.

FAQ

Does robots.txt on my article host control an image on a CDN?

Not automatically. Check the robots.txt for the image’s actual serving authority, including any host reached after redirects.

Does allowing Googlebot prove a social platform can fetch the image?

No. The result applies to the crawler tested. Identify and test the social platform’s fetcher separately using current official guidance.

A blocked crawler cannot see a noindex directive on that resource. Crawling and indexing require separate decisions.

Can robots.txt protect an image from public access?

No. Robots.txt is a request to crawlers, not an access-control mechanism. Use authentication or other access controls for private assets.