What Are the Sub-Factors for Citability Scoring in GEO?

The citability scoring system in zubair-trabzada/geo-seo-claude evaluates content using five deterministic sub-factors—Answer Quality, Self-Containment, Structure, Statistical Density, and Uniqueness—which are aggregated via the score_passage() function in scripts/citability_scorer.py to generate a 0–100 page-level score.

Generative Engine Optimization (GEO) requires content that Large Language Models (LLMs) can easily extract and cite. The zubair-trabzada/geo-seo-claude repository implements a deterministic citability rubric defined in skills/geo-citability/SKILL.md that scores individual content blocks before aggregating them into page-level metrics. Understanding the sub-factors for citability scoring is essential for optimizing content structures that perform well in AI-driven search engines like ChatGPT, Claude, and Perplexity.

The Five Sub-Factors for Citability Scoring

According to the source documentation in skills/geo-citability/SKILL.md (lines 3, 18, and 105), the citability scorer evaluates every substantive content block against five distinct dimensions. These sub-factors determine how readily an LLM can extract and cite the passage without requiring additional context:

  • Answer Quality: Measures how directly and completely the passage answers a potential user query. High-scoring content provides immediate, unambiguous answers rather than requiring inference or external knowledge.

  • Self-Containment: Evaluates whether the passage stands alone without requiring external context. Self-contained blocks include necessary definitions, background, and scope within the text itself.

  • Structure: Assesses logical organization through semantic HTML elements, hierarchical headings, lists, and tables. Well-structured content enables AI systems to parse relationships between concepts efficiently.

  • Statistical Density: Quantifies the presence of concrete data, figures, percentages, and verifiable facts. Passages rich in statistics provide citable evidence that LLMs prefer when generating factual responses.

  • Uniqueness: Measures the originality of phrasing and ideas compared to existing indexed content. Unique wording reduces semantic overlap with competing sources and increases the probability of LLM citation.

How the Scoring Algorithm Works

The citability algorithm processes content through the score_passage() function implemented in scripts/citability_scorer.py. As documented in docs/scoring-methodology.md (lines 78–88), the scoring workflow follows these steps:

  1. Individual content blocks are parsed and evaluated against the five sub-factors
  2. Each block receives a 0–100 numeric score based on the dimension rubric
  3. The page-level citability score is calculated as the average of the top-five highest-scoring blocks (or all blocks if fewer than five exist)

The methodology assigns citability a weight of 0.25 (25%) within the overall GEO scoring matrix, making it a primary pillar of the optimization framework alongside technical accessibility and brand presence.

Implementing Citability Scoring

You can calculate citability scores programmatically using the Python API or via the command-line interface.

Python Implementation

from scripts.citability_scorer import score_passage

text = """
According to a 2024 Gartner report, AI-driven search adoption increased by 30%,
with enterprise implementations showing 45% faster information retrieval compared
to traditional keyword-based systems.
"""

# Returns dict with total_score and dimension breakdown

result = score_passage(text)

print(result)

# {

#   "total_score": 87,

#   "dimensions": {

#     "answer_quality": 85,

#     "self_containment": 90,

#     "structure": 80,

#     "statistical_density": 95,

#     "uniqueness": 88

#   }

# }

CLI Usage


# Analyze a remote URL

geo citability https://example.com/blog-post

# Generates GEO-CITABILITY-SCORE.md with:

# • Overall citability score (0-100)

# • Weighted breakdown of the five sub-factors

# • Weakest block identification and rewrite suggestions

Key Implementation Files

The citability scoring system spans several critical components:

Summary

  • The citability scoring system uses five deterministic sub-factors: Answer Quality, Self-Containment, Structure, Statistical Density, and Uniqueness.
  • Each content block receives a 0–100 score via the score_passage() function in scripts/citability_scorer.py.
  • Page-level scores aggregate the top-five block scores to represent the strongest citable content on the page.
  • According to docs/scoring-methodology.md, citability carries a 25% weight in the overall Generative Engine Optimization score.

Frequently Asked Questions

How is the overall citability score calculated from the sub-factors?

The score_passage() function in scripts/citability_scorer.py evaluates individual content blocks against the five dimensions, returning a numeric score for each. The page-level citability score is then computed as the average of the top-five highest-scoring blocks, as defined in docs/scoring-methodology.md (lines 78–88). If fewer than five blocks exist, the system averages all available blocks.

What weight does citability carry in the total GEO score?

According to docs/scoring-methodology.md, the citability sub-score is weighted at 0.25 (25%) of the total Generative Engine Optimization score. This makes it one of the four primary evaluation pillars in the framework, alongside technical factors, brand presence, and schema markup compliance.

Can I improve individual sub-factors without rewriting entire articles?

Yes. The deterministic nature of the score_passage() algorithm allows targeted optimization. You can improve Statistical Density by adding concrete percentages and figures, enhance Structure by implementing semantic HTML headings, or boost Self-Containment by including contextual definitions within individual passages rather than relying on surrounding content.

Is the citability scoring algorithm deterministic?

Yes. As implemented in scripts/citability_scorer.py, the scoring algorithm produces consistent, reproducible results for identical inputs. This determinism enables content teams to verify optimization changes through A/B testing and ensures that the geo citability CLI command returns the same scores across different environments when analyzing identical content.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →