ScreenshotNeo

BlogGuides

How to Find Websites That Publish an llms.txt File

A practical workflow for discovering llms.txt files through page metadata, scoped paths, documentation sites, and verified public directories.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: inspect the target page’s HTML and HTTP response headers for rel="describedby". If the site does not advertise a file, test the most specific path that could apply, then the origin root, and also check documentation subdomains or subpaths. A 404 at /llms.txt only proves that the root file is missing; a scoped file may still exist below /docs/ or another path.

1. Understand what you are looking for

llms.txt is a Markdown overview and curated link map for a site or URL scope. The proposal allows it at the origin root or under a path, where the path-scoped file applies to URLs beneath that path. The usual document has an H1, an optional blockquote summary, explanatory prose, and H2 sections containing links to useful resources. Read those links as the next step rather than treating the file as the content itself.

The convention is separate from robots.txt: robots.txt communicates crawler access preferences, while llms.txt provides on-demand content orientation. Publishing a file does not prove that a particular AI system will fetch or follow it. See the official proposal and the Chrome Lighthouse guidance.

2. Check the page’s HTML and HTTP headers first

A declared relation is more direct and less error-prone than guessing filenames. Look for an HTML link whose relation is describedby; its URL identifies the map relevant to the page. Also record rel="alternate" type="text/markdown", which advertises a Markdown representation of the individual page, not necessarily an llms.txt map.

curl -sS -D headers.txt -o page.html https://example.com/docs/api/auth

# Inspect HTTP Link headers
rg -i '(^|:) *Link:|describedby|text/markdown' headers.txt

# Inspect HTML link elements
rg -i 'rel=["'"']?[^"'"']*describedby|type=["'"']text/markdown|llms\.txt' page.html

Follow the URL in the describedby relation exactly. Resolve relative URLs against the page origin, and preserve a path-specific URL when one is supplied.

3. Test path-scoped locations, then the root

If no relation is declared, derive candidates from the page path. For https://example.com/docs/api/auth, check these in order:

  1. https://example.com/docs/api/llms.txt
  2. https://example.com/docs/llms.txt
  3. https://example.com/llms.txt

Use the most specific successful file. A server may return redirects, so follow them and record the final URL. Prefer a successful response with a text content type, but do not reject a valid file only because a server labels it application/octet-stream.

for url in \
  https://example.com/docs/api/llms.txt \
  https://example.com/docs/llms.txt \
  https://example.com/llms.txt; do
  printf '\n=== %s ===\n' "$url"
  curl -L -sS -o /tmp/llms.txt -w 'status=%{http_code} type=%{content_type} final=%{url_effective}\n' "$url"
  if test -s /tmp/llms.txt; then head -n 12 /tmp/llms.txt; fi
done

4. Check documentation hosts and subpaths

Software companies often publish documentation on a subdomain such as docs.example.com, or under a path such as example.com/docs/. Repeat the relation and candidate-path checks on that host. Do not infer that a main-domain 404 applies to the documentation host.

curl -sS -I https://docs.example.com/llms.txt
curl -sS -I https://example.com/docs/llms.txt

5. Discover many candidate sites with public directories

For broad discovery, use directories such as directory.llmstxt.cloud and llmstxt.site. Treat entries as leads: listings can become stale, move, or omit scoped files. Fetch the live URL, verify the status and content, and note the retrieval time. The community reference collects additional discovery resources.

Method Coverage Directness Freshness
HTML or HTTP relation One known page Highest Live
Scoped path guesses One host and hierarchy Medium Live
Public directory Many candidate sites Lower until verified May lag

6. Automate discovery in Python

This script checks a page’s HTML and Link headers, then tests increasingly broad path candidates. It reports only responses that are likely text documents.

from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup

s = requests.Session()
s.headers['User-Agent'] = 'llms-txt-discovery/1.0'
page_url = 'https://example.com/docs/api/auth'
r = s.get(page_url, timeout=20, allow_redirects=True)
r.raise_for_status()

soup = BeautifulSoup(r.text, 'html.parser')
found = []
for link in soup.find_all('link'):
    rel = {x.lower() for x in link.get('rel', [])}
    if 'describedby' in rel:
        found.append(urljoin(r.url, link.get('href', '')))

for raw in r.headers.get('Link', '').split(','):
    if 'describedby' in raw.lower():
        left = raw.split(';', 1)[0].strip()
        if left.startswith('<') and left.endswith('>'):
            found.append(urljoin(r.url, left[1:-1]))

p = urlparse(r.url)
parts = [x for x in p.path.split('/') if x]
if parts and '.' in parts[-1]:
    parts.pop()
for i in range(len(parts), -1, -1):
    path = '/' + '/'.join(parts[:i]) + '/llms.txt'
    found.append(f'{p.scheme}://{p.netloc}{path}')

seen = set()
for candidate in found:
    if not candidate or candidate in seen:
        continue
    seen.add(candidate)
    try:
        x = s.get(candidate, timeout=20, allow_redirects=True)
        kind = x.headers.get('content-type', '')
        if x.status_code == 200 and ('text' in kind or x.text.lstrip().startswith('#')):
            print(x.url, kind, len(x.content))
    except requests.RequestException as e:
        print('error', candidate, e)

7. Automate discovery in Node.js

const page = new URL('https://example.com/docs/api/auth');
const res = await fetch(page, {redirect: 'follow'});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const html = await res.text();
const candidates = new Set();

for (const m of html.matchAll(/<link[^>]+>/gi)) {
  const tag = m[0];
  if (/rel=["''][^"'']*describedby/i.test(tag)) {
    const href = tag.match(/href=["'']([^"'']+)/i)?.[1];
    if (href) candidates.add(new URL(href, res.url).href);
  }
}
const link = res.headers.get('link') || '';
for (const part of link.split(',')) {
  if (/describedby/i.test(part)) {
    const href = part.match(/<([^>]+)>/)?.[1];
    if (href) candidates.add(new URL(href, res.url).href);
  }
}
const u = new URL(res.url);
const pieces = u.pathname.split('/').filter(Boolean);
if (pieces.at(-1)?.includes('.')) pieces.pop();
for (let i = pieces.length; i >= 0; i--) {
  candidates.add(`${u.origin}/${pieces.slice(0, i).join('/')}${i ? '/' : ''}llms.txt`);
}
for (const candidate of candidates) {
  const x = await fetch(candidate, {redirect: 'follow'});
  const type = x.headers.get('content-type') || '';
  if (x.status === 200 && (type.includes('text') || (await x.clone().text()).trimStart().startsWith('#'))) {
    console.log(candidate, x.status, type);
  }
}

8. Validate and interpret a result

  • Confirm the response is current and belongs to the expected host or scope.
  • Read the H1 and summary to identify coverage.
  • Check H2 link groups for API references, guides, policies, and canonical pages.
  • Resolve relative links and test important destinations.
  • Record redirects, status, content type, and retrieval time for reproducibility.

Do not treat a file’s existence as evidence that every AI product consumes it. Chrome describes the convention as optional, and Lighthouse marks a missing file as Not Applicable. The community standards reference likewise classifies it as an emerging convention without a documented commitment from every engine.

9. Troubleshooting common failures

Symptom Likely cause Fix
Root URL returns 404 No root file Check the page relation, scoped paths, and docs host.
HTML search finds nothing Metadata is delivered in HTTP headers or injected differently Inspect Link headers with curl -I; check server-rendered HTML.
Redirect loop or wrong host Canonicalization or proxy rule Use -L, inspect the final URL, and stop if it leaves the expected site.
403 or 429 Access control or rate limiting Slow requests, use a descriptive user agent, and follow the site’s terms.
200 response is an app shell SPA fallback serves HTML for unknown paths Inspect content type and body; a status code alone is insufficient.
Directory entry is dead Stale listing Verify against the live URL and search the site’s current docs paths.
Several files succeed Different scopes Choose the most specific file for your target URL; retain broader files for navigation.

10. Performance, reliability, and cost

For one site, relation checks plus three or four targeted requests are usually faster and more reliable than crawling. Cache successful results with a timestamp, use connection reuse, set finite timeouts, and back off on 429 responses. For directory-scale work, queue hosts, cap concurrency per domain, and avoid repeatedly downloading unchanged files. A HEAD request is useful for triage, but some servers implement HEAD incorrectly; fall back to a small GET when headers are missing or unreliable.

The requests themselves are ordinary HTTP fetches. Your cost is bandwidth and runtime; public directories may save discovery time but add verification requests. Keep raw responses or hashes if you need an audit trail.

Or skip the browser setup

If you need a screenshot of an llms.txt page, documentation page, or verification result, ScreenshotNeo provides a single API request. The service removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses identify the page verdict and billing status in X-Page-Verdict and X-Billed headers. An MCP server also lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/llms.txt -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/llms.txt"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/llms.txt' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Where is llms.txt located?

It may be at the origin root, such as /llms.txt, or under a path such as /docs/llms.txt. A declared describedby relation is the authoritative discovery hint.

Does every website publish one?

No. The convention is optional, and a missing file is not an error.

Is llms.txt the same as robots.txt?

No. robots.txt expresses crawler access preferences; llms.txt is a curated orientation map.

Should I trust a directory listing?

Use it to find candidates, then fetch and verify the live URL because entries can become stale.

What if both root and scoped files exist?

Use the most specific file that covers the page, while keeping the root file as a broader navigation source.