How to Use the Python Scripts for Programmatic GEO-SEO Claude Analysis
The GEO-SEO Claude repository ships self-contained Python scripts that enable you to programmatically fetch web pages, validate and generate llms.txt files, audit brand presence across AI-impact platforms, and calculate citability scores for AI search optimization.
The zubair-trabzada/geo-seo-claude repository provides a modular toolkit for Generative Engine Optimization (GEO) that mimics how AI crawlers analyze web content. By leveraging these Python scripts, you can automate SEO audits that extract structured data, verify AI crawler guidance, and assess brand visibility across platforms like YouTube, Reddit, and Wikipedia. This guide demonstrates how to integrate these tools into your workflow through both command-line interfaces and direct Python module imports.
Core Scripts Overview
The repository contains five specialized scripts located in the scripts/ directory, each targeting specific GEO-SEO signals:
fetch_page.py– Retrieves web pages using realistic browser headers and extracts comprehensive SEO payloads including meta tags, heading structures, JSON-LD, and SSR (server-side render) detection.llmstxt_generator.py– Validates existingllms.txtfiles or generates newllms.txtandllms-full.txtfiles to guide AI crawlers.brand_scanner.py– Queries high-impact platforms (YouTube, Reddit, Wikipedia, LinkedIn) to generate brand presence reports.citability_scorer.py– Combines page metrics and brand signals to calculate how likely content is to be cited by AI systems.crm_dashboard.py– Flattens JSON outputs into CSV format for spreadsheet integration and reporting.
All scripts share a common DEFAULT_HEADERS constant that mimics modern Chrome browser headers, ensuring requests bypass typical anti-bot filters while maintaining only standard dependencies listed in requirements.txt.
Installation and Setup
Before using the scripts programmatically, install the required dependencies:
pip install -r requirements.txt
The repository requires only requests, beautifulsoup4, and lxml, making it lightweight for production environments.
Fetching Page Data with fetch_page.py
The scripts/fetch_page.py module serves as the foundation for page analysis, performing HTTP GET requests with browser-like headers and parsing HTML via BeautifulSoup.
Command-Line Usage
Fetch a single page and output structured JSON:
python scripts/fetch_page.py https://example.com
This outputs a JSON object containing title, meta_tags, heading_structure, word_count, internal_links, external_links, and an has_ssr_content boolean flag indicating whether the page relies on client-side rendering.
Programmatic Integration
Import the fetch_page function directly into your Python applications:
from scripts.fetch_page import fetch_page
data = fetch_page("https://example.com")
print(data["title"])
print(data["heading_structure"])
print(data["has_ssr_content"])
Managing AI Crawler Guidance with llmstxt_generator.py
The scripts/llmstxt_generator.py module handles the emerging standard for AI crawler guidance files, supporting both validation of existing files and generation of new ones.
Validating Existing llms.txt Files
Verify that an existing llms.txt meets formatting requirements:
python scripts/llmstxt_generator.py https://example.com validate
This returns a JSON response with format_valid status and specific issues or suggestions for improvement.
Generating New llms.txt and llms-full.txt
Create fresh AI guidance files by crawling and categorizing site content:
python scripts/llmstxt_generator.py https://example.com generate
For programmatic use, import the generate_llmstxt function:
from scripts.llmstxt_generator import generate_llmstxt
result = generate_llmstxt("https://example.com")
with open("llms.txt", "w") as f:
f.write(result["generated_llmstxt"])
with open("llms-full.txt", "w") as f:
f.write(result["generated_llmstxt_full"])
The function categorizes discovered pages into Products, Resources, Company, Support, and Main sections, producing both a concise llms.txt and a detailed llms-full.txt with optional page descriptions.
Auditing Brand Presence with brand_scanner.py
The scripts/brand_scanner.py module queries platforms that most influence AI citation graphs, including YouTube, Reddit, Wikipedia/Wikidata, and LinkedIn.
Running Brand Scans from the Terminal
Execute a comprehensive brand audit:
python scripts/brand_scanner.py "Acme Corp" acmecorp.com
This produces a JSON report containing platform-specific presence flags, direct search URLs, and prioritized overall_recommendations for improving visibility.
Importing as a Python Module
Integrate brand monitoring into automated workflows:
from scripts.brand_scanner import generate_brand_report
report = generate_brand_report("Acme Corp", "acmecorp.com")
print(report["overall_recommendations"][0])
Scoring and Exporting Results
After gathering raw data, use the remaining scripts to calculate metrics and format outputs for business intelligence tools.
Calculating Citability Scores
The scripts/citability_scorer.py module combines metrics from fetch_page.py (SSR flags, structured data presence, word count) with brand_scanner.py outputs to generate a numeric citability score and explanatory rationale.
Exporting to CSV for Reporting
Convert JSON outputs from any script into CRM-ready CSV format:
python scripts/crm_dashboard.py fetch_page_output.json > page_report.csv
This reads the JSON payloads, flattens key fields, and writes a CSV file suitable for import into spreadsheets or customer relationship management systems.
Typical Workflow for Automated GEO-SEO
To implement a complete programmatic pipeline using the GEO-SEO Claude scripts:
- Extract page signals – Run
fetch_page.pyto capture technical SEO data and SSR status. - Verify AI guidance – Use
llmstxt_generator.pyin validate mode to check existing llms.txt files, or generate mode to create new ones. - Audit brand presence – Execute
brand_scanner.pyto identify visibility gaps on AI-impact platforms. - Calculate citability – Process outputs through
citability_scorer.pyto prioritize optimization efforts. - Export and report – Use
crm_dashboard.pyto flatten results for stakeholder dashboards.
These scripts expose their core functions (fetch_page, validate_llmstxt, generate_llmstxt, generate_brand_report) for seamless integration into larger Python applications, supporting both one-off analyses and fully automated monitoring pipelines.
Summary
- The
zubair-trabzada/geo-seo-clauderepository provides five specialized scripts for programmatic GEO-SEO analysis. fetch_page.pyextracts comprehensive SEO data including SSR detection and structured data parsing.llmstxt_generator.pyvalidates and generates AI crawler guidance files in both concise and detailed formats.brand_scanner.pyaudits brand presence across YouTube, Reddit, Wikipedia, and LinkedIn.- All scripts support both CLI usage and Python module imports via functions like
fetch_page()andgenerate_brand_report(). - The
DEFAULT_HEADERSconstant ensures requests mimic legitimate browser traffic to avoid anti-bot blocking.
Frequently Asked Questions
What dependencies are required to run these scripts?
The scripts require only requests, beautifulsoup4, and lxml as specified in requirements.txt. These standard libraries handle HTTP requests and HTML parsing without additional heavy frameworks.
Can I use these scripts in a production pipeline?
Yes. The scripts are designed for both command-line usage and programmatic import. Core functions like fetch_page(), generate_llmstxt(), and generate_brand_report() return structured dictionaries that can be consumed by scheduling systems, CI/CD pipelines, or monitoring dashboards.
What is the difference between llms.txt and llms-full.txt?
According to the llmstxt_generator.py implementation, llms.txt provides a concise categorized list of URLs (Products, Resources, Company, Support, Main), while llms-full.txt includes the same structure with detailed page descriptions and additional metadata for AI crawlers requiring comprehensive context.
How does the citability scorer determine ratings?
The citability_scorer.py algorithm combines technical signals from fetch_page.py—including the has_ssr_content flag, JSON-LD presence, heading hierarchy quality, and word count—with brand authority metrics from brand_scanner.py to produce a composite score indicating likelihood of AI citation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →