What Community Detection Algorithm Does Code-Review-Graph Use to Identify Code Modules?

Code-review-graph uses the Leiden community detection algorithm from the igraph library to identify code modules and boundaries, with a deterministic file-based fallback when igraph is unavailable.

Code-review-graph is an open-source tool that analyzes codebases to discover logical modules and architectural boundaries. At the heart of its analysis pipeline lies a sophisticated community detection algorithm that transforms raw dependency graphs into meaningful code communities. This article explains exactly how the system detects communities, what algorithm it employs, and how you can leverage its detection capabilities in your own workflows.

The Leiden Algorithm as Primary Community Detector

The core community detection algorithm in code-review-graph is the Leiden algorithm, implemented through igraph.Graph.community_leiden. The project authors chose this algorithm for its speed, quality guarantees, and ability to handle large-scale graphs efficiently.

In code_review_graph/communities.py, the detection logic configures Leiden with two specific constraints for reproducibility and performance:


# From code_review_graph/communities.py lines 53-60

# Pseudocode representation of the actual implementation

import igraph

def _run_leiden_communities(graph):
    # Seed RNG for deterministic results across runs

    random.seed(42)
    # Cap at 2 iterations for performance

    communities = graph.community_leiden(
        weights='weight',
        n_iterations=2
    )
    return communities

The two-iteration limit balances detection quality with execution speed, while the fixed random seed ensures identical community structures on every run—critical for consistent code reviews and CI/CD integrations.

The Complete Community Detection Flow

Code-review-graph applies a multi-stage pipeline after invoking the Leiden algorithm. Understanding these stages helps explain why the detected communities accurately reflect real code modules.

Stage 1: Test Node Reassignment

Raw Leiden clusters often separate test files from the code they exercise. The system corrects this by reassigning test nodes to communities where they cover the most unique production subjects:


# Conceptual flow from code_review_graph/communities.py lines 61-70

def reassign_test_nodes(communities, graph):
    for test_node in graph.test_nodes:
        # Find which production nodes this test exercises

        exercised = graph.get_exercised_subjects(test_node)
        # Move to community with maximum unique coverage

        best_community = max(
            communities,
            key=lambda c: len(exercised & c.production_nodes)
        )
        best_community.add(test_node)

This step ensures that test files live alongside the code they test, making communities more semantically meaningful.

Stage 2: Cohesion Scoring

Each community receives a cohesion score representing its internal edge ratio—essentially how tightly connected the nodes are versus their external connections. This calculation runs in O(edges) time:


# From code_review_graph/communities.py lines 87-94

def calculate_cohesion(community, graph):
    internal_edges = sum(
        1 for edge in subgraph_edges(community)
        if edge.target in community
    )
    total_edges = len(subgraph_edges(community))
    return internal_edges / total_edges if total_edges else 0.0

High cohesion indicates a well-defined module; low cohesion suggests potential architectural refactoring opportunities.

Stage 3: Oversized Community Splitting

Communities exceeding 25% of the total graph size trigger recursive splitting. The system extracts the subgraph and re-runs Leiden on that subset:


# From code_review_graph/communities.py lines 62-70

def split_oversized_communities(communities, graph):
    threshold = len(graph.nodes) * 0.25
    for comm in communities:
        if len(comm) > threshold:
            subgraph = graph.induced_subgraph(comm)
            sub_communities = _run_leiden_communities(subgraph)
            # Recursively apply same pipeline

            yield from process_communities(sub_communities, subgraph)
        else:
            yield comm

This prevents "mega-communities" that obscure meaningful internal structure—particularly common in monolithic codebases.

The Fallback: File-Based Deterministic Grouping

When igraph is not installed, code-review-graph degrades gracefully to a directory-prefix heuristic. This fallback groups nodes by common file path prefixes:


# From code_review_graph/communities.py lines 78-85

def file_based_fallback(nodes):
    from collections import defaultdict
    groups = defaultdict(list)
    for node in nodes:
        # Group by top-level directory

        prefix = node.filepath.split('/')[0]
        groups[prefix].append(node)
    return list(groups.values())

This ensures 100% reproducible results in restricted environments, though with lower semantic accuracy than Leiden-based detection.

Public API for Community Detection

The detection functionality exposes two main entry points in code_review_graph/tools/community_tools.py:

Function Purpose Typical Use Case
detect_communities Full detection run on entire codebase Initial project analysis, CI pipelines
incremental_detect_communities Detection limited to changed files Pull request reviews, incremental updates

Both functions delegate to the same lower-level helpers described above, ensuring consistent behavior.

Practical Usage Examples

Query detected communities through the tool interface:

from code_review_graph.tools.community_tools import list_communities_func, get_architecture_overview_func

# List all detected communities (default sorting by size)

communities = list_communities_func()
print("Found:", communities["summary"])

# Get a concise architecture overview (minimal detail)

overview = get_architecture_overview_func(detail_level="minimal")
print("Architecture:", overview["summary"])
for comm in overview["communities"]:
    print(f"- {comm['name']} ({comm['size']} nodes, cohesion={comm['cohesion']})")

For lower-level access when you manage your own GraphStore:

from code_review_graph.communities import detect_communities

# Direct call with custom minimum community size

store = ...               # GraphStore created elsewhere

all_communities = detect_communities(store, min_size=3)
for c in all_communities:
    print(c["name"], c["size"], c["cohesion"])

Key Implementation Files

Understanding the codebase structure helps when extending or debugging community detection:

Summary

  • Primary algorithm: Leiden community detection via igraph.Graph.community_leiden
  • Key constraints: 2 iterations max, seeded RNG for reproducibility
  • Post-processing: Test node reassignment, cohesion scoring, and recursive splitting of oversized communities
  • Fallback: Directory-prefix grouping when igraph unavailable
  • Entry points: detect_communities and incremental_detect_communities in community_tools.py

Frequently Asked Questions

Why does code-review-graph use Leiden instead of Louvain?

The Leiden algorithm improves upon Louvain by guaranteeing well-connected communities and avoiding arbitrarily badly connected clusters—critical when code modules must represent functional units rather than arbitrary groupings. The implementation in igraph also offers better performance characteristics on the sparse dependency graphs typical of software codebases.

How can I ensure deterministic community detection results?

Install igraph and rely on the default Leiden path. The code explicitly seeds the random number generator with a fixed value (line 58 in communities.py), ensuring identical partitions across runs. The file-based fallback is also deterministic by design.

What happens if my codebase has no clear modular boundaries?

The cohesion scores exposed in community metadata quantify boundary clarity. Low cohesion values suggest highly coupled code that resists clean modularization. The oversized-community splitting logic (25% threshold) also attempts to force structure even in monolithic codebases, though manual architectural review may be warranted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →