How to Use the Python Scripts for Programmatic GEO-SEO Claude Analysis

The GEO-SEO Claude repository ships self-contained Python scripts that enable you to programmatically fetch web pages, validate and generate llms.txt files, audit brand presence across AI-impact platforms, and calculate citability scores for AI search optimization.

The zubair-trabzada/geo-seo-claude repository provides a modular toolkit for Generative Engine Optimization (GEO) that mimics how AI crawlers analyze web content. By leveraging these Python scripts, you can automate SEO audits that extract structured data, verify AI crawler guidance, and assess brand visibility across platforms like YouTube, Reddit, and Wikipedia. This guide demonstrates how to integrate these tools into your workflow through both command-line interfaces and direct Python module imports.

Core Scripts Overview

The repository contains five specialized scripts located in the scripts/ directory, each targeting specific GEO-SEO signals:

  • fetch_page.py – Retrieves web pages using realistic browser headers and extracts comprehensive SEO payloads including meta tags, heading structures, JSON-LD, and SSR (server-side render) detection.
  • llmstxt_generator.py – Validates existing llms.txt files or generates new llms.txt and llms-full.txt files to guide AI crawlers.
  • brand_scanner.py – Queries high-impact platforms (YouTube, Reddit, Wikipedia, LinkedIn) to generate brand presence reports.
  • citability_scorer.py – Combines page metrics and brand signals to calculate how likely content is to be cited by AI systems.
  • crm_dashboard.py – Flattens JSON outputs into CSV format for spreadsheet integration and reporting.

All scripts share a common DEFAULT_HEADERS constant that mimics modern Chrome browser headers, ensuring requests bypass typical anti-bot filters while maintaining only standard dependencies listed in requirements.txt.

Installation and Setup

Before using the scripts programmatically, install the required dependencies:

pip install -r requirements.txt

The repository requires only requests, beautifulsoup4, and lxml, making it lightweight for production environments.

Fetching Page Data with fetch_page.py

The scripts/fetch_page.py module serves as the foundation for page analysis, performing HTTP GET requests with browser-like headers and parsing HTML via BeautifulSoup.

Command-Line Usage

Fetch a single page and output structured JSON:

python scripts/fetch_page.py https://example.com

This outputs a JSON object containing title, meta_tags, heading_structure, word_count, internal_links, external_links, and an has_ssr_content boolean flag indicating whether the page relies on client-side rendering.

Programmatic Integration

Import the fetch_page function directly into your Python applications:

from scripts.fetch_page import fetch_page

data = fetch_page("https://example.com")
print(data["title"])
print(data["heading_structure"])
print(data["has_ssr_content"])

Managing AI Crawler Guidance with llmstxt_generator.py

The scripts/llmstxt_generator.py module handles the emerging standard for AI crawler guidance files, supporting both validation of existing files and generation of new ones.

Validating Existing llms.txt Files

Verify that an existing llms.txt meets formatting requirements:

python scripts/llmstxt_generator.py https://example.com validate

This returns a JSON response with format_valid status and specific issues or suggestions for improvement.

Generating New llms.txt and llms-full.txt

Create fresh AI guidance files by crawling and categorizing site content:

python scripts/llmstxt_generator.py https://example.com generate

For programmatic use, import the generate_llmstxt function:

from scripts.llmstxt_generator import generate_llmstxt

result = generate_llmstxt("https://example.com")

with open("llms.txt", "w") as f:
    f.write(result["generated_llmstxt"])
    
with open("llms-full.txt", "w") as f:
    f.write(result["generated_llmstxt_full"])

The function categorizes discovered pages into Products, Resources, Company, Support, and Main sections, producing both a concise llms.txt and a detailed llms-full.txt with optional page descriptions.

Auditing Brand Presence with brand_scanner.py

The scripts/brand_scanner.py module queries platforms that most influence AI citation graphs, including YouTube, Reddit, Wikipedia/Wikidata, and LinkedIn.

Running Brand Scans from the Terminal

Execute a comprehensive brand audit:

python scripts/brand_scanner.py "Acme Corp" acmecorp.com

This produces a JSON report containing platform-specific presence flags, direct search URLs, and prioritized overall_recommendations for improving visibility.

Importing as a Python Module

Integrate brand monitoring into automated workflows:

from scripts.brand_scanner import generate_brand_report

report = generate_brand_report("Acme Corp", "acmecorp.com")
print(report["overall_recommendations"][0])

Scoring and Exporting Results

After gathering raw data, use the remaining scripts to calculate metrics and format outputs for business intelligence tools.

Calculating Citability Scores

The scripts/citability_scorer.py module combines metrics from fetch_page.py (SSR flags, structured data presence, word count) with brand_scanner.py outputs to generate a numeric citability score and explanatory rationale.

Exporting to CSV for Reporting

Convert JSON outputs from any script into CRM-ready CSV format:

python scripts/crm_dashboard.py fetch_page_output.json > page_report.csv

This reads the JSON payloads, flattens key fields, and writes a CSV file suitable for import into spreadsheets or customer relationship management systems.

Typical Workflow for Automated GEO-SEO

To implement a complete programmatic pipeline using the GEO-SEO Claude scripts:

  1. Extract page signals – Run fetch_page.py to capture technical SEO data and SSR status.
  2. Verify AI guidance – Use llmstxt_generator.py in validate mode to check existing llms.txt files, or generate mode to create new ones.
  3. Audit brand presence – Execute brand_scanner.py to identify visibility gaps on AI-impact platforms.
  4. Calculate citability – Process outputs through citability_scorer.py to prioritize optimization efforts.
  5. Export and report – Use crm_dashboard.py to flatten results for stakeholder dashboards.

These scripts expose their core functions (fetch_page, validate_llmstxt, generate_llmstxt, generate_brand_report) for seamless integration into larger Python applications, supporting both one-off analyses and fully automated monitoring pipelines.

Summary

  • The zubair-trabzada/geo-seo-claude repository provides five specialized scripts for programmatic GEO-SEO analysis.
  • fetch_page.py extracts comprehensive SEO data including SSR detection and structured data parsing.
  • llmstxt_generator.py validates and generates AI crawler guidance files in both concise and detailed formats.
  • brand_scanner.py audits brand presence across YouTube, Reddit, Wikipedia, and LinkedIn.
  • All scripts support both CLI usage and Python module imports via functions like fetch_page() and generate_brand_report().
  • The DEFAULT_HEADERS constant ensures requests mimic legitimate browser traffic to avoid anti-bot blocking.

Frequently Asked Questions

What dependencies are required to run these scripts?

The scripts require only requests, beautifulsoup4, and lxml as specified in requirements.txt. These standard libraries handle HTTP requests and HTML parsing without additional heavy frameworks.

Can I use these scripts in a production pipeline?

Yes. The scripts are designed for both command-line usage and programmatic import. Core functions like fetch_page(), generate_llmstxt(), and generate_brand_report() return structured dictionaries that can be consumed by scheduling systems, CI/CD pipelines, or monitoring dashboards.

What is the difference between llms.txt and llms-full.txt?

According to the llmstxt_generator.py implementation, llms.txt provides a concise categorized list of URLs (Products, Resources, Company, Support, Main), while llms-full.txt includes the same structure with detailed page descriptions and additional metadata for AI crawlers requiring comprehensive context.

How does the citability scorer determine ratings?

The citability_scorer.py algorithm combines technical signals from fetch_page.py—including the has_ssr_content flag, JSON-LD presence, heading hierarchy quality, and word count—with brand authority metrics from brand_scanner.py to produce a composite score indicating likelihood of AI citation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →