How to Analyze and Generate llms.txt for a Website Using Python

Use the validate_llmstxt() and generate_llmstxt() functions in the Geo-SEO CLI to check existing files for specification compliance or crawl a site to create new llms.txt and llms-full.txt documents.

The zubair-trabzada/geo-seo-claude repository provides a lightweight Python utility that automates the creation and validation of llms.txt files. These files help large language models understand your site structure by providing concise, structured context about your content. According to the source code in [scripts/llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py), the tool handles both validation of existing files and generation of new ones through automated site crawling.

Core Functions for llms.txt Management

The [llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py) script exposes two primary entry points that handle different stages of the llms.txt lifecycle.

Validating Existing llms.txt Files

The validate_llmstxt(url) function performs a comprehensive format check on existing files hosted at /llms.txt and /llms-full.txt. It constructs absolute URLs using urlparse (lines 33-34), then downloads content via requests.get() with DEFAULT_HEADERS to mimic browser behavior and avoid bot detection (lines 58-60).

The validation logic enforces four required structural elements:

  • Title: First line must start with # (line 68)

  • Description: Must contain at least one blockquote line starting with > (lines 74-77)

  • Sections: Requires minimum one ## heading (line 82)

  • Links: Must include at least one markdown link matching the regex pattern r"- \[.+\]\(.+\)" (line 89)

The function returns a dictionary containing exists, format_valid, link counts, issues, and actionable suggestions for remediation.

Generating New llms.txt from Site Crawls

The generate_llmstxt(url, max_pages=30) function creates both concise and detailed versions by crawling your website. It fetches the homepage using BeautifulSoup with the lxml parser (line 45) to extract the site name and meta description, then discovers internal links while filtering external domains, static assets, and duplicate anchors (lines 75-84).

The crawler categorizes pages into five logical groups based on URL path patterns (lines 90-98):

  • Main Pages
  • Products & Services
  • Resources & Blog
  • Company Information
  • Support

The function outputs two distinct formats:

  • llms.txt: Concise version with title, description, section headings, and plain link lists
  • llms-full.txt: Enhanced version that fetches each page's <meta name="description"> content and appends it to link entries (lines 51-58)

Implementation Examples

Validating an Existing File Programmatically

Import the validation function to check compliance before deployment:

from scripts.llmstxt_generator import validate_llmstxt

result = validate_llmstxt("https://example.com")

if result["format_valid"]:
    print(f"✓ Valid format with {result['link_count']} links")
else:
    print("Issues found:")
    for issue in result["issues"]:
        print(f"  - {issue}")

Generating Files with Custom Crawl Depth

Control the crawl scope by adjusting the max_pages parameter and persist both output variants:

from scripts.llmstxt_generator import generate_llmstxt
import pathlib

# Crawl up to 50 pages

output = generate_llmstxt("https://example.com", max_pages=50)

# Write standard version

pathlib.Path("llms.txt").write_text(output["generated_llmstxt"])

# Write extended version with descriptions

pathlib.Path("llms-full.txt").write_text(output["generated_llmstxt_full"])

Command Line Interface

Invoke the script directly for quick validation or generation tasks:


# Check if existing llms.txt follows the specification

python scripts/llmstxt_generator.py https://example.com validate

# Generate fresh files for your domain

python scripts/llmstxt_generator.py https://example.com generate

Both commands output JSON-formatted results suitable for piping into CI/CD pipelines.

Key Source Files and Architecture

Understanding the repository structure helps when extending functionality or debugging issues.

File Purpose Link
scripts/llmstxt_generator.py Core engine with validate_llmstxt() and generate_llmstxt() scripts/llmstxt_generator.py
scripts/fetch_page.py Page retrieval helper supporting the crawl infrastructure scripts/fetch_page.py
tests/test_fetch_page_ssr.py Unit tests verifying BeautifulSoup parsing and header handling tests/test_fetch_page_ssr.py
requirements.txt Dependency manifest listing requests, beautifulsoup4, lxml requirements.txt
README.md Installation instructions and general CLI usage README.md

The implementation uses DEFAULT_HEADERS for all HTTP requests to prevent simple bot blocking, and the categorization logic relies on keyword matching against URL paths to organize content hierarchically.

Summary

  • Validation: Use validate_llmstxt() to programmatically verify that existing files contain required Markdown structure including H1 titles, blockquote descriptions, H2 sections, and properly formatted links.
  • Generation: Use generate_llmstxt() to automatically crawl websites, categorize pages into logical groups, and produce both concise and detailed file versions.
  • Compliance: The validator checks four specific structural elements defined in lines 68-89 of the source code.
  • Crawling: The generator filters external domains and static assets while respecting the max_pages parameter to manage server load.
  • Integration: Both functions support CLI invocation and return JSON-serializable dictionaries for automated workflow integration.

Frequently Asked Questions

What is the difference between llms.txt and llms-full.txt?

The llms.txt file provides a concise overview with page titles and URLs organized by category, while llms-full.txt includes the meta description for each page. According to the source code in [llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py), the full version requires additional HTTP requests to fetch each discovered page and extract its <meta name="description"> tag (lines 51-58), making it more comprehensive but slower to generate.

How does the validator determine if an llms.txt file is correctly formatted?

The validation function performs four specific checks: the first line must start with # (title), at least one line must start with > (description), at least one line must contain ## (section headings), and at least one line must match the regex pattern - \[.+\]\(.+\) (markdown links). Any missing element triggers specific issues and suggestions in the return dictionary.

Can I customize which pages are included when generating llms.txt?

Yes, you can control the crawl depth using the max_pages parameter in generate_llmstxt(), which defaults to 30 pages. The crawler automatically filters out external domains, static assets, and duplicate anchor links based on the logic in lines 75-84. For further customization, modify the categorization logic around lines 90-98 to adjust how URLs are grouped into sections.

What dependencies are required to run the llms.txt generator?

The script requires requests for HTTP operations, beautifulsoup4 for HTML parsing, and lxml as the parser backend. These are listed in [requirements.txt](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/requirements.txt). The tool also uses Python standard libraries including urlparse, re, and pathlib for URL manipulation and pattern matching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →