How to Analyze and Generate llms.txt for a Website Using Python
Use the validate_llmstxt() and generate_llmstxt() functions in the Geo-SEO CLI to check existing files for specification compliance or crawl a site to create new llms.txt and llms-full.txt documents.
The zubair-trabzada/geo-seo-claude repository provides a lightweight Python utility that automates the creation and validation of llms.txt files. These files help large language models understand your site structure by providing concise, structured context about your content. According to the source code in [scripts/llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py), the tool handles both validation of existing files and generation of new ones through automated site crawling.
Core Functions for llms.txt Management
The [llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py) script exposes two primary entry points that handle different stages of the llms.txt lifecycle.
Validating Existing llms.txt Files
The validate_llmstxt(url) function performs a comprehensive format check on existing files hosted at /llms.txt and /llms-full.txt. It constructs absolute URLs using urlparse (lines 33-34), then downloads content via requests.get() with DEFAULT_HEADERS to mimic browser behavior and avoid bot detection (lines 58-60).
The validation logic enforces four required structural elements:
-
Title: First line must start with
#(line 68) -
Description: Must contain at least one blockquote line starting with
>(lines 74-77) -
Sections: Requires minimum one
##heading (line 82) -
Links: Must include at least one markdown link matching the regex pattern
r"- \[.+\]\(.+\)"(line 89)
The function returns a dictionary containing exists, format_valid, link counts, issues, and actionable suggestions for remediation.
Generating New llms.txt from Site Crawls
The generate_llmstxt(url, max_pages=30) function creates both concise and detailed versions by crawling your website. It fetches the homepage using BeautifulSoup with the lxml parser (line 45) to extract the site name and meta description, then discovers internal links while filtering external domains, static assets, and duplicate anchors (lines 75-84).
The crawler categorizes pages into five logical groups based on URL path patterns (lines 90-98):
- Main Pages
- Products & Services
- Resources & Blog
- Company Information
- Support
The function outputs two distinct formats:
llms.txt: Concise version with title, description, section headings, and plain link listsllms-full.txt: Enhanced version that fetches each page's<meta name="description">content and appends it to link entries (lines 51-58)
Implementation Examples
Validating an Existing File Programmatically
Import the validation function to check compliance before deployment:
from scripts.llmstxt_generator import validate_llmstxt
result = validate_llmstxt("https://example.com")
if result["format_valid"]:
print(f"✓ Valid format with {result['link_count']} links")
else:
print("Issues found:")
for issue in result["issues"]:
print(f" - {issue}")
Generating Files with Custom Crawl Depth
Control the crawl scope by adjusting the max_pages parameter and persist both output variants:
from scripts.llmstxt_generator import generate_llmstxt
import pathlib
# Crawl up to 50 pages
output = generate_llmstxt("https://example.com", max_pages=50)
# Write standard version
pathlib.Path("llms.txt").write_text(output["generated_llmstxt"])
# Write extended version with descriptions
pathlib.Path("llms-full.txt").write_text(output["generated_llmstxt_full"])
Command Line Interface
Invoke the script directly for quick validation or generation tasks:
# Check if existing llms.txt follows the specification
python scripts/llmstxt_generator.py https://example.com validate
# Generate fresh files for your domain
python scripts/llmstxt_generator.py https://example.com generate
Both commands output JSON-formatted results suitable for piping into CI/CD pipelines.
Key Source Files and Architecture
Understanding the repository structure helps when extending functionality or debugging issues.
| File | Purpose | Link |
|---|---|---|
scripts/llmstxt_generator.py |
Core engine with validate_llmstxt() and generate_llmstxt() |
scripts/llmstxt_generator.py |
scripts/fetch_page.py |
Page retrieval helper supporting the crawl infrastructure | scripts/fetch_page.py |
tests/test_fetch_page_ssr.py |
Unit tests verifying BeautifulSoup parsing and header handling | tests/test_fetch_page_ssr.py |
requirements.txt |
Dependency manifest listing requests, beautifulsoup4, lxml |
requirements.txt |
README.md |
Installation instructions and general CLI usage | README.md |
The implementation uses DEFAULT_HEADERS for all HTTP requests to prevent simple bot blocking, and the categorization logic relies on keyword matching against URL paths to organize content hierarchically.
Summary
- Validation: Use
validate_llmstxt()to programmatically verify that existing files contain required Markdown structure including H1 titles, blockquote descriptions, H2 sections, and properly formatted links. - Generation: Use
generate_llmstxt()to automatically crawl websites, categorize pages into logical groups, and produce both concise and detailed file versions. - Compliance: The validator checks four specific structural elements defined in lines 68-89 of the source code.
- Crawling: The generator filters external domains and static assets while respecting the
max_pagesparameter to manage server load. - Integration: Both functions support CLI invocation and return JSON-serializable dictionaries for automated workflow integration.
Frequently Asked Questions
What is the difference between llms.txt and llms-full.txt?
The llms.txt file provides a concise overview with page titles and URLs organized by category, while llms-full.txt includes the meta description for each page. According to the source code in [llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py), the full version requires additional HTTP requests to fetch each discovered page and extract its <meta name="description"> tag (lines 51-58), making it more comprehensive but slower to generate.
How does the validator determine if an llms.txt file is correctly formatted?
The validation function performs four specific checks: the first line must start with # (title), at least one line must start with > (description), at least one line must contain ## (section headings), and at least one line must match the regex pattern - \[.+\]\(.+\) (markdown links). Any missing element triggers specific issues and suggestions in the return dictionary.
Can I customize which pages are included when generating llms.txt?
Yes, you can control the crawl depth using the max_pages parameter in generate_llmstxt(), which defaults to 30 pages. The crawler automatically filters out external domains, static assets, and duplicate anchor links based on the logic in lines 75-84. For further customization, modify the categorization logic around lines 90-98 to adjust how URLs are grouped into sections.
What dependencies are required to run the llms.txt generator?
The script requires requests for HTTP operations, beautifulsoup4 for HTML parsing, and lxml as the parser backend. These are listed in [requirements.txt](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/requirements.txt). The tool also uses Python standard libraries including urlparse, re, and pathlib for URL manipulation and pattern matching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →