How Community Detection with the Leiden Algorithm Powers Code‑Review‑Graph's Architecture Analysis

The Leiden algorithm in code‑review‑graph clusters related code units into communities by maximizing modularity on a weighted dependency graph, with deterministic output and automatic fallback to directory‑based grouping when igraph is unavailable.

The tirth8205/code-review‑graph project builds a code‑knowledge graph from repository structure and then applies community detection to surface architectural patterns. The Leiden algorithm serves as its primary clustering engine, translating dependency relationships into meaningful, named communities that represent tightly‑coupled functional units.

What the Leiden Algorithm Does in Code‑Review‑Graph

Community detection transforms a complex web of code dependencies into human‑readable clusters. In code_review_graph/communities.py, the implementation treats each function, class, or module as a graph node and each CALLS, IMPORTS_FROM, or INHERITS_FROM relationship as a weighted, undirected edge. The Leiden algorithm then optimizes modularity—a measure of how densely connected nodes are within groups versus between them—to find the natural boundaries of cohesive code.

The algorithm is particularly well‑suited for this use case because it produces higher‑quality partitions than the older Louvain method, with guarantees that communities are well‑connected and that the process converges efficiently even on large codebases.

Step‑by‑Step Implementation of Community Detection

The detection pipeline in detect_communities(store, min_size=2) follows a rigorous sequence:

Optional igraph Import and Fallback

The system first attempts to import igraph at line 33 of code_review_graph/communities.py. If the library is absent, it seamlessly falls back to _detect_file_based, which groups nodes by directory depth instead:

try:
    import igraph
    HAS_IGRAPH = True
except ImportError:
    HAS_IGRAPH = False

Building Node and Edge Mappings

Before invoking the algorithm, the code constructs translation dictionaries to interface with igraph's integer‑based indexing. At line 68, it builds:

  • qn_to_idx: maps qualified node names (e.g., mymodule.MyClass.method) to vertex indices
  • idx_to_node: reverse mapping back to GraphNode objects

Weighted Edge Assembly

Starting at line 84, the system iterates through all GraphEdge instances, converts source/target qualified names to indices, deduplicates undirected edges, and assigns per‑kind weights—so a CALLS edge may carry different influence than an IMPORTS_FROM edge in the clustering calculation.

Resolution Scaling for Repository Size

At line 100, the resolution parameter (controlling community granularity) is computed as inversely proportional to the logarithm of node count:

resolution = 1.0 / math.log(max(len(nodes), 10))

This yields coarser clusters for massive repositories and finer‑grained communities for smaller codebases.

Deterministic Seeding

Reproducibility is enforced at line 110 via the CRG_LEIDEN_SEED environment variable (default: 42). The seed is injected into igraph's random number generator so identical graphs always produce identical community IDs.

Running Leiden with Bounded Iterations

The core call at line 116 uses fixed parameters optimized for code graphs:

partition = g.community_leiden(
    objective_function="modularity",
    weights="weight",
    resolution=resolution,
    n_iterations=2
)

The hard‑coded 2 iterations prevent exponential runtimes on dense dependency graphs while still achieving stable, high‑quality partitions.

Post‑Processing: Test Nodes, Cohesion, and Hierarchy

Reassigning Test Nodes

After raw detection, the _reassign_test_nodes function (line 61) migrates test‑only nodes to the community containing the most unique subjects they test. This preserves deterministic output while placing tests near their relevant production code.

Cohesion Scoring via Batch Computation

Community quality is assessed using _compute_cohesion_batch at line 87. This single‑pass O(edges) routine tallies internal versus external edges for all communities simultaneously—avoiding the N² penalty of per‑community scans. The resulting cohesion score (internal edges ÷ total edges) identifies well‑encapsulated modules versus tangled architectural hotspots.

Recursive Splitting of Oversized Communities

Communities exceeding 25% of total node count trigger _split_oversized (line 66), which recursively re‑runs Leiden on the sub‑graph. This produces a hierarchy of sub‑communities for monolithic codebases without manual intervention.

Name Generation and Deduplication

Each community receives a readable name based on dominant language and shared path prefixes. The _dedupe_community_names function (line 332) resolves collisions by appending distinctive keywords or numeric suffixes.

Detecting Communities in Practice

The public API requires minimal boilerplate:

from code_review_graph.graph import GraphStore
from code_review_graph.communities import detect_communities, get_architecture_overview

# Load your graph from a previously analyzed repository

store = GraphStore(path="my_repo_graph.db")

# Run detection—Leiden activates automatically if igraph is installed

communities = detect_communities(store, min_size=5)

# Inspect results

for c in communities:
    print(f"{c['name']}: {c['size']} nodes, cohesion={c['cohesion']:.2f}")

# Generate cross‑community coupling warnings

overview = get_architecture_overview(store)
for warning in overview["warnings"]:
    print(f"Architectural concern: {warning}")

If igraph is missing, the identical call transparently uses the directory‑based detector—no code changes required.

Key Configuration and Environment Variables

Variable Purpose Default
CRG_LEIDEN_SEED RNG seed for deterministic community IDs 42
min_size parameter Minimum nodes for a community to be returned 2

The resolution scaling and 2‑iteration bound are currently hard‑coded but tuned for typical software repository sizes (10³–10⁶ nodes).

Summary

  • Leiden algorithm maximizes modularity on weighted dependency graphs to identify natural code boundaries
  • Automatic fallback to directory‑based grouping when igraph is unavailable
  • Deterministic output via CRG_LEIDEN_SEED and bounded 2‑iteration convergence
  • Test‑aware post‑processing reassigns test files to their most‑tested production communities
  • Batch cohesion scoring computes cluster quality in O(edges) time
  • Recursive splitting handles oversized communities without manual tuning
  • Entry point detect_communities() in code_review_graph/communities.py orchestrates the pipeline

Frequently Asked Questions

What happens if igraph is not installed?

The system falls back to _detect_file_based, which groups nodes by directory depth and applies the same cohesion and naming logic. The public detect_communities() API remains unchanged—callers need no conditional code.

Why is the iteration count fixed at 2?

Two iterations provide stable, high‑quality partitions for code‑dependency graphs while preventing exponential runtime growth on dense edges. The developers empirically validated this trade‑off against repository sizes ranging from small libraries to monorepos exceeding 100,000 nodes.

How does cohesion relate to architectural quality?

Cohesion measures the fraction of a community's edges that stay internal versus connect outward. Values near 1.0 indicate well‑encapsulated modules; low values flag architectural tangles that may benefit from refactoring. The batch computation makes this metric practical even for large graphs.

Can I customize the resolution parameter for finer or coarser clusters?

Currently, resolution is auto‑scaled via 1.0 / log(node_count). Direct override is not exposed in the public API, though the logarithmic scaling naturally adapts to repository size. Users requiring specific granularity can pre‑filter nodes or post‑process the returned communities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →