# How to Analyze and Generate llms.txt for a Website Using Python

> Learn to analyze and generate llms.txt for your website using Python. The Geo-SEO CLI offers functions to validate compliance or crawl your site to create new llms.txt and llms-full.txt files.

- Repository: [Zubair Trabzada/geo-seo-claude](https://github.com/zubair-trabzada/geo-seo-claude)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Use the `validate_llmstxt()` and `generate_llmstxt()` functions in the Geo-SEO CLI to check existing files for specification compliance or crawl a site to create new [`llms.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms.txt) and [`llms-full.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms-full.txt) documents.**

The [zubair-trabzada/geo-seo-claude](https://github.com/zubair-trabzada/geo-seo-claude) repository provides a lightweight Python utility that automates the creation and validation of [`llms.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms.txt) files. These files help large language models understand your site structure by providing concise, structured context about your content. According to the source code in [[`scripts/llmstxt_generator.py`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py)](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py), the tool handles both validation of existing files and generation of new ones through automated site crawling.

## Core Functions for llms.txt Management

The [[`llmstxt_generator.py`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llmstxt_generator.py)](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py) script exposes two primary entry points that handle different stages of the [`llms.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms.txt) lifecycle.

### Validating Existing llms.txt Files

The **`validate_llmstxt(url)`** function performs a comprehensive format check on existing files hosted at [`/llms.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main//llms.txt) and [`/llms-full.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main//llms-full.txt). It constructs absolute URLs using `urlparse` (lines 33-34), then downloads content via `requests.get()` with `DEFAULT_HEADERS` to mimic browser behavior and avoid bot detection (lines 58-60).

The validation logic enforces four required structural elements:
- **Title**: First line must start with `# ` (line 68)

- **Description**: Must contain at least one blockquote line starting with `> ` (lines 74-77)
- **Sections**: Requires minimum one `## ` heading (line 82)

- **Links**: Must include at least one markdown link matching the regex pattern `r"- \[.+\]\(.+\)"` (line 89)

The function returns a dictionary containing `exists`, `format_valid`, link counts, `issues`, and actionable `suggestions` for remediation.

### Generating New llms.txt from Site Crawls

The **`generate_llmstxt(url, max_pages=30)`** function creates both concise and detailed versions by crawling your website. It fetches the homepage using `BeautifulSoup` with the `lxml` parser (line 45) to extract the site name and meta description, then discovers internal links while filtering external domains, static assets, and duplicate anchors (lines 75-84).

The crawler categorizes pages into five logical groups based on URL path patterns (lines 90-98):
- Main Pages
- Products & Services
- Resources & Blog
- Company Information
- Support

The function outputs two distinct formats:
- **[`llms.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms.txt)**: Concise version with title, description, section headings, and plain link lists
- **[`llms-full.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms-full.txt)**: Enhanced version that fetches each page's `<meta name="description">` content and appends it to link entries (lines 51-58)

## Implementation Examples

### Validating an Existing File Programmatically

Import the validation function to check compliance before deployment:

```python
from scripts.llmstxt_generator import validate_llmstxt

result = validate_llmstxt("https://example.com")

if result["format_valid"]:
    print(f"✓ Valid format with {result['link_count']} links")
else:
    print("Issues found:")
    for issue in result["issues"]:
        print(f"  - {issue}")

```

### Generating Files with Custom Crawl Depth

Control the crawl scope by adjusting the `max_pages` parameter and persist both output variants:

```python
from scripts.llmstxt_generator import generate_llmstxt
import pathlib

# Crawl up to 50 pages

output = generate_llmstxt("https://example.com", max_pages=50)

# Write standard version

pathlib.Path("llms.txt").write_text(output["generated_llmstxt"])

# Write extended version with descriptions

pathlib.Path("llms-full.txt").write_text(output["generated_llmstxt_full"])

```

### Command Line Interface

Invoke the script directly for quick validation or generation tasks:

```bash

# Check if existing llms.txt follows the specification

python scripts/llmstxt_generator.py https://example.com validate

# Generate fresh files for your domain

python scripts/llmstxt_generator.py https://example.com generate

```

Both commands output JSON-formatted results suitable for piping into CI/CD pipelines.

## Key Source Files and Architecture

Understanding the repository structure helps when extending functionality or debugging issues.

| File | Purpose | Link |
|------|---------|------|
| [`scripts/llmstxt_generator.py`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py) | Core engine with `validate_llmstxt()` and `generate_llmstxt()` | [scripts/llmstxt_generator.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py) |
| [`scripts/fetch_page.py`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/fetch_page.py) | Page retrieval helper supporting the crawl infrastructure | [scripts/fetch_page.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/fetch_page.py) |
| [`tests/test_fetch_page_ssr.py`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/tests/test_fetch_page_ssr.py) | Unit tests verifying BeautifulSoup parsing and header handling | [tests/test_fetch_page_ssr.py](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/tests/test_fetch_page_ssr.py) |
| [`requirements.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/requirements.txt) | Dependency manifest listing `requests`, `beautifulsoup4`, `lxml` | [requirements.txt](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/requirements.txt) |
| [`README.md`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/README.md) | Installation instructions and general CLI usage | [README.md](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/README.md) |

The implementation uses `DEFAULT_HEADERS` for all HTTP requests to prevent simple bot blocking, and the categorization logic relies on keyword matching against URL paths to organize content hierarchically.

## Summary

- **Validation**: Use `validate_llmstxt()` to programmatically verify that existing files contain required Markdown structure including H1 titles, blockquote descriptions, H2 sections, and properly formatted links.
- **Generation**: Use `generate_llmstxt()` to automatically crawl websites, categorize pages into logical groups, and produce both concise and detailed file versions.
- **Compliance**: The validator checks four specific structural elements defined in lines 68-89 of the source code.
- **Crawling**: The generator filters external domains and static assets while respecting the `max_pages` parameter to manage server load.
- **Integration**: Both functions support CLI invocation and return JSON-serializable dictionaries for automated workflow integration.

## Frequently Asked Questions

### What is the difference between llms.txt and llms-full.txt?

The [`llms.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms.txt) file provides a concise overview with page titles and URLs organized by category, while [`llms-full.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llms-full.txt) includes the meta description for each page. According to the source code in [[`llmstxt_generator.py`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/llmstxt_generator.py)](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/scripts/llmstxt_generator.py), the full version requires additional HTTP requests to fetch each discovered page and extract its `<meta name="description">` tag (lines 51-58), making it more comprehensive but slower to generate.

### How does the validator determine if an llms.txt file is correctly formatted?

The validation function performs four specific checks: the first line must start with `# ` (title), at least one line must start with `> ` (description), at least one line must contain `## ` (section headings), and at least one line must match the regex pattern `- \[.+\]\(.+\)` (markdown links). Any missing element triggers specific issues and suggestions in the return dictionary.

### Can I customize which pages are included when generating llms.txt?

Yes, you can control the crawl depth using the `max_pages` parameter in `generate_llmstxt()`, which defaults to 30 pages. The crawler automatically filters out external domains, static assets, and duplicate anchor links based on the logic in lines 75-84. For further customization, modify the categorization logic around lines 90-98 to adjust how URLs are grouped into sections.

### What dependencies are required to run the llms.txt generator?

The script requires `requests` for HTTP operations, `beautifulsoup4` for HTML parsing, and `lxml` as the parser backend. These are listed in [[`requirements.txt`](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/requirements.txt)](https://github.com/zubair-trabzada/geo-seo-claude/blob/main/requirements.txt). The tool also uses Python standard libraries including `urlparse`, `re`, and `pathlib` for URL manipulation and pattern matching.