ScreenshotNeo

BlogAI agents

Free llms.txt Generator: Create an AI-Readable Sitemap

Create a useful llms.txt file from your sitemap, publish it correctly, and keep it updated with free scripts and generator options.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: An llms.txt generator creates a curated Markdown file that explains your website to AI agents and links them to the pages that matter most. Start with your sitemap, keep the list focused, publish the result at https://your-domain.example/llms.txt, and review it whenever your important pages change.

llms.txt is a proposed convention for an agent-readable website overview. It complements robots.txt and sitemap.xml; it does not replace either one.

What llms.txt is

The proposal describes a Markdown file that gives an agent orientation before it fetches detailed pages. A useful file has:

  • A clear H1 title.
  • An optional blockquote with a short site summary.
  • Optional explanatory paragraphs.
  • H2 sections containing links and brief descriptions.

Keep the overview concise. Link to authoritative pages so an agent can retrieve detail only when it needs it. A file can live at the site root or at a subpath that covers the URLs below that path.

llms.txt vs robots.txt vs sitemap.xml

File Purpose Typical content
robots.txt Communicates crawler access rules Allow/disallow directives and sitemap locations
sitemap.xml Lists indexable, human-facing URLs URL records, modification dates and optional priorities
llms.txt Provides curated orientation for agents Markdown sections, selected links and context

Publish all three when they serve your site. A sitemap is comprehensive; llms.txt is selective and explanatory.

What a good llms.txt file looks like

# Acme Documentation

> Acme provides hosted payment APIs, SDKs and integration guides.

Use these pages to understand and integrate Acme:

## Start here
- [Documentation home](https://acme.example/docs/): Overview of the API and SDKs.
- [Quickstart](https://acme.example/docs/quickstart): Make your first test request.
- [Authentication](https://acme.example/docs/authentication): API keys, OAuth and permissions.

## API reference
- [REST API reference](https://acme.example/docs/api/): Endpoints, parameters and responses.
- [Webhooks](https://acme.example/docs/webhooks): Event types and signature verification.

## Policies
- [Changelog](https://acme.example/changelog/): Product and API changes.
- [Terms](https://acme.example/legal/terms/): Service terms.

Descriptions should tell an agent why a link is useful. Avoid dumping every low-value URL from your sitemap.

How to create an llms.txt file manually

  1. Choose the scope. A root file can describe the whole site; a subpath file can describe a documentation section.
  2. Write one sentence that identifies the site and its audience.
  3. Group links by task, such as “Start here,” “API reference,” “Guides” and “Policies.”
  4. Select canonical, maintained pages. Remove tag archives, duplicate URLs, search results and expired campaigns.
  5. Add a short description to every link.
  6. Save the file as plain UTF-8 Markdown named llms.txt.
  7. Deploy it at the site root or the relevant subpath.
  8. Request the public URL and check that it returns HTTP 200 with a text content type.

Build a free generator from your sitemap

The following script downloads a sitemap, removes common low-value paths, and writes a starter file. It is intentionally conservative: edit the generated Markdown before publishing.

Python generator

from urllib.parse import urlparse
from xml.etree import ElementTree as ET
from pathlib import Path
import urllib.request

SITEMAP_URL = 'https://example.com/sitemap.xml'
OUTPUT = Path('llms.txt')
MAX_LINKS = 100
SKIP_PARTS = ('/tag/', '/search', '/author/', '/feed', '?replytocom=')

def fetch(url):
    request = urllib.request.Request(url, headers={'User-Agent': 'llms-txt-generator/1.0'})
    with urllib.request.urlopen(request, timeout=30) as response:
        return response.read()

def sitemap_urls(xml_bytes):
    root = ET.fromstring(xml_bytes)
    return [node.text.strip() for node in root.iter() if node.tag.endswith('}loc') and node.text]

urls = []
for url in sitemap_urls(fetch(SITEMAP_URL)):
    parsed = urlparse(url)
    if parsed.scheme not in ('http', 'https'):
        continue
    if any(part in url for part in SKIP_PARTS):
        continue
    urls.append(url)
    if len(urls) >= MAX_LINKS:
        break

lines = [
    '# Example.com',
    '',
    '> Example.com documentation and guides.',
    '',
    '## Important pages',
]
for url in urls:
    path = urlparse(url).path.strip('/') or 'home'
    label = path.replace('-', ' ').replace('/', ' / ').title()
    lines.append(f'- [{label}]({url}): Read the {label.lower()} page.')

OUTPUT.write_text('\n'.join(lines) + '\n', encoding='utf-8')
print(f'Wrote {len(urls)} links to {OUTPUT}')

Node.js generator

import { writeFile } from 'node:fs/promises';

const sitemapUrl = 'https://example.com/sitemap.xml';
const response = await fetch(sitemapUrl, {
  headers: { 'user-agent': 'llms-txt-generator/1.0' }
});
if (!response.ok) throw new Error(`Sitemap request failed: ${response.status}`);

const xml = await response.text();
const urls = [...xml.matchAll(/<loc>([^<]+)<\/loc>/g)]
  .map(match => match[1].trim())
  .filter(url => !/\/tag\/|\/search|\/author\/|\/feed|replytocom=/i.test(url))
  .slice(0, 100);

const lines = [
  '# Example.com',
  '',
  '> Example.com documentation and guides.',
  '',
  '## Important pages',
  ...urls.map(url => {
    const path = new URL(url).pathname.replace(/^\/+|\/+$/g, '') || 'home';
    const label = path.replaceAll('-', ' ').replaceAll('/', ' / ');
    return `- [${label}](${url}): Read the ${label} page.`;
  })
];

await writeFile('llms.txt', lines.join('\n') + '\n', 'utf8');
console.log(`Wrote ${urls.length} links to llms.txt`);

cURL: inspect and validate the published file

curl -i https://example.com/llms.txt

Check for a successful status, readable Markdown, the expected hostname, and links that resolve. A simple link check can be run with:

curl -s https://example.com/llms.txt | grep -oE 'https?://[^ )]+' | while read url; do
  curl -L -s -o /dev/null -w '%{http_code} %{url_effective}\n' "$url"
done

Generator options and selection checklist

Whether you use a script, CMS plugin or hosted generator, evaluate these capabilities:

  • Sitemap import: Supports XML sitemaps and, if needed, sitemap indexes.
  • Selection controls: Lets you include, exclude, reorder and group URLs.
  • Proposal structure: Produces an H1, optional summary and H2 link sections.
  • Subpaths: Can create a file for a documentation or product section.
  • Markdown links: Supports clean Markdown page variants where your site provides them.
  • Deployment: Explains how to publish at the correct path and content type.
  • Automation: Regenerates after content changes without overwriting editorial choices unexpectedly.
  • Privacy: States whether URLs or site data leave your infrastructure.
  • Cost: Is free without signup if that is a requirement.

Where to put llms.txt

For a whole site, publish the file at:

https://your-domain.example/llms.txt

For a section, publish it beneath that section, such as /docs/llms.txt, and include links within that scope. Keep canonical URLs in the links. The proposal also discusses clean Markdown page variants such as .md URLs; only link those variants when your platform actually serves them.

Keeping the file accurate

  • Regenerate or review it when navigation, products or documentation structure changes.
  • Prefer stable landing pages over temporary announcements.
  • Remove redirected, deleted and duplicate URLs.
  • Keep descriptions factual and short.
  • Check that sensitive, private or unlaunched URLs are not exposed.
  • Record the generation date in your deployment process, not necessarily in the public file.

Adoption and implementation are evolving. A platform may generate the file, but individual AI tools can choose whether and how to use it. Verify current behavior before promising visibility or ranking outcomes.

Platform and integration choices

The proposal lists integrations and implementations including Mintlify, GitBook, Yoast SEO, AIOSEO, Wix, vitepress-plugin-llms, docusaurus-plugin-llms, Drupal LLM Support, llms-txt-php and server-llm-txt. Choose a native integration when it can preserve your editorial grouping and deployment workflow; use a script when you need full control or self-hosting.

Or skip the browser setup

If you need screenshots of the pages linked from your new llms.txt, ScreenshotNeo provides a website screenshot API. It accepts one GET request and returns PNG, JPEG, WebP or PDF. The request below captures the published file’s page (replace the URL with any page you want to document).

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/llms.txt -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://example.com/llms.txt'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/llms.txt' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for options. Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get started.

Troubleshooting

Problem Cause Fix
404 at /llms.txt The file was deployed to a theme or asset folder. Place it at the public site root and redeploy.
HTML is returned instead of Markdown A rewrite, SPA fallback or authentication layer intercepted the path. Exclude llms.txt from the rewrite and serve it as a static text file.
Generator finds no URLs The sitemap uses namespaces, a sitemap index or blocked access. Parse namespaced loc elements, fetch child sitemaps, and inspect the HTTP response.
Links point to staging The sitemap’s canonical base URL is wrong. Correct the sitemap and filter hosts before generating.
File is huge and unfocused Every sitemap URL was copied without curation. Set a limit, remove archives and group only important pages.
Private pages are exposed The generator imported URLs that should not be public. Exclude them, review the output, and enforce authentication separately.
Changes do not appear CDN or server caching is serving an old file. Purge the cache and check the response headers and deployment artifact.

Performance, reliability and cost

  • Performance: Keep the file small enough for a quick first fetch. A curated list is faster for agents to process than a full URL inventory.
  • Reliability: Generate from a stable sitemap, validate links, and deploy atomically so agents never receive a partial file.
  • Automation: Run generation after content releases, then apply an editorial allowlist or review step before publishing.
  • Cost: A local script and static file can run without a service fee. Hosted generators may add account or usage costs; verify their current terms.
  • Privacy: Do not include unpublished, personal or access-controlled URLs merely because they appear in an internal sitemap.

FAQ

Does llms.txt improve search rankings?

The proposal is an orientation aid for agents, not a ranking guarantee. AI tools decide independently whether to fetch or use it.

Must every page be listed?

No. Select the pages that best explain your products, documentation and policies, then link to deeper material from those pages.

Can I use Markdown page URLs?

Yes, when your site serves clean Markdown variants. Otherwise link to the normal canonical pages.

Can one site have more than one file?

Yes. A root file can describe the whole site and subpath files can describe focused sections.

Is a generator required?

No. A hand-written file is valid and may be better for a small site. A generator helps repeatable sitemap imports and updates.

Does llms.txt control crawling?

No. Use robots.txt for crawler access rules and authentication or server controls for private content.