# What Community Detection Algorithm Does Code-Review-Graph Use to Identify Code Modules?

> Discover how code-review-graph identifies code modules using the Leiden community detection algorithm. Learn about its deterministic fallback for module boundary analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-15

---

**Code-review-graph uses the Leiden community detection algorithm from the igraph library** to identify code modules and boundaries, with a deterministic file-based fallback when igraph is unavailable.

Code-review-graph is an open-source tool that analyzes codebases to discover logical modules and architectural boundaries. At the heart of its analysis pipeline lies a sophisticated **community detection algorithm** that transforms raw dependency graphs into meaningful code communities. This article explains exactly how the system detects communities, what algorithm it employs, and how you can leverage its detection capabilities in your own workflows.

## The Leiden Algorithm as Primary Community Detector

The core community detection algorithm in code-review-graph is the **Leiden algorithm**, implemented through `igraph.Graph.community_leiden`. The project authors chose this algorithm for its speed, quality guarantees, and ability to handle large-scale graphs efficiently.

In [`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py), the detection logic configures Leiden with two specific constraints for reproducibility and performance:

```python

# From code_review_graph/communities.py lines 53-60

# Pseudocode representation of the actual implementation

import igraph

def _run_leiden_communities(graph):
    # Seed RNG for deterministic results across runs

    random.seed(42)
    # Cap at 2 iterations for performance

    communities = graph.community_leiden(
        weights='weight',
        n_iterations=2
    )
    return communities

```

The **two-iteration limit** balances detection quality with execution speed, while the **fixed random seed** ensures identical community structures on every run—critical for consistent code reviews and CI/CD integrations.

## The Complete Community Detection Flow

Code-review-graph applies a multi-stage pipeline after invoking the Leiden algorithm. Understanding these stages helps explain why the detected communities accurately reflect real code modules.

### Stage 1: Test Node Reassignment

Raw Leiden clusters often separate test files from the code they exercise. The system corrects this by reassigning test nodes to communities where they cover the most unique production subjects:

```python

# Conceptual flow from code_review_graph/communities.py lines 61-70

def reassign_test_nodes(communities, graph):
    for test_node in graph.test_nodes:
        # Find which production nodes this test exercises

        exercised = graph.get_exercised_subjects(test_node)
        # Move to community with maximum unique coverage

        best_community = max(
            communities,
            key=lambda c: len(exercised & c.production_nodes)
        )
        best_community.add(test_node)

```

This step ensures that **test files live alongside the code they test**, making communities more semantically meaningful.

### Stage 2: Cohesion Scoring

Each community receives a **cohesion score** representing its internal edge ratio—essentially how tightly connected the nodes are versus their external connections. This calculation runs in **O(edges)** time:

```python

# From code_review_graph/communities.py lines 87-94

def calculate_cohesion(community, graph):
    internal_edges = sum(
        1 for edge in subgraph_edges(community)
        if edge.target in community
    )
    total_edges = len(subgraph_edges(community))
    return internal_edges / total_edges if total_edges else 0.0

```

High cohesion indicates a well-defined module; low cohesion suggests potential architectural refactoring opportunities.

### Stage 3: Oversized Community Splitting

Communities exceeding **25% of the total graph size** trigger recursive splitting. The system extracts the subgraph and re-runs Leiden on that subset:

```python

# From code_review_graph/communities.py lines 62-70

def split_oversized_communities(communities, graph):
    threshold = len(graph.nodes) * 0.25
    for comm in communities:
        if len(comm) > threshold:
            subgraph = graph.induced_subgraph(comm)
            sub_communities = _run_leiden_communities(subgraph)
            # Recursively apply same pipeline

            yield from process_communities(sub_communities, subgraph)
        else:
            yield comm

```

This prevents "mega-communities" that obscure meaningful internal structure—particularly common in monolithic codebases.

## The Fallback: File-Based Deterministic Grouping

When **igraph is not installed**, code-review-graph degrades gracefully to a **directory-prefix heuristic**. This fallback groups nodes by common file path prefixes:

```python

# From code_review_graph/communities.py lines 78-85

def file_based_fallback(nodes):
    from collections import defaultdict
    groups = defaultdict(list)
    for node in nodes:
        # Group by top-level directory

        prefix = node.filepath.split('/')[0]
        groups[prefix].append(node)
    return list(groups.values())

```

This ensures **100% reproducible results** in restricted environments, though with lower semantic accuracy than Leiden-based detection.

## Public API for Community Detection

The detection functionality exposes two main entry points in [`code_review_graph/tools/community_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/community_tools.py):

| Function | Purpose | Typical Use Case |
|----------|---------|----------------|
| `detect_communities` | Full detection run on entire codebase | Initial project analysis, CI pipelines |
| `incremental_detect_communities` | Detection limited to changed files | Pull request reviews, incremental updates |

Both functions delegate to the same lower-level helpers described above, ensuring consistent behavior.

### Practical Usage Examples

Query detected communities through the tool interface:

```python
from code_review_graph.tools.community_tools import list_communities_func, get_architecture_overview_func

# List all detected communities (default sorting by size)

communities = list_communities_func()
print("Found:", communities["summary"])

# Get a concise architecture overview (minimal detail)

overview = get_architecture_overview_func(detail_level="minimal")
print("Architecture:", overview["summary"])
for comm in overview["communities"]:
    print(f"- {comm['name']} ({comm['size']} nodes, cohesion={comm['cohesion']})")

```

For lower-level access when you manage your own `GraphStore`:

```python
from code_review_graph.communities import detect_communities

# Direct call with custom minimum community size

store = ...               # GraphStore created elsewhere

all_communities = detect_communities(store, min_size=3)
for c in all_communities:
    print(c["name"], c["size"], c["cohesion"])

```

## Key Implementation Files

Understanding the codebase structure helps when extending or debugging community detection:

- **[`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py)** — Core Leiden implementation, test reassignment, cohesion scoring, splitting logic, and fallback method
- **[`code_review_graph/tools/community_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/community_tools.py)** — User-facing tools: `list_communities_func`, `get_community_func`, `get_architecture_overview_func`
- **[`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py)** — `GraphNode`, `GraphEdge`, and `GraphStore` definitions that feed the detector
- **[`code_review_graph/postprocessing.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py)** — Integration point where `detect_communities` runs during full analysis pipelines

## Summary

- **Primary algorithm**: Leiden community detection via `igraph.Graph.community_leiden`
- **Key constraints**: 2 iterations max, seeded RNG for reproducibility
- **Post-processing**: Test node reassignment, cohesion scoring, and recursive splitting of oversized communities
- **Fallback**: Directory-prefix grouping when igraph unavailable
- **Entry points**: `detect_communities` and `incremental_detect_communities` in [`community_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/community_tools.py)

## Frequently Asked Questions

### Why does code-review-graph use Leiden instead of Louvain?

The **Leiden algorithm** improves upon Louvain by guaranteeing well-connected communities and avoiding arbitrarily badly connected clusters—critical when code modules must represent functional units rather than arbitrary groupings. The implementation in igraph also offers better performance characteristics on the sparse dependency graphs typical of software codebases.

### How can I ensure deterministic community detection results?

Install **igraph** and rely on the default Leiden path. The code explicitly seeds the random number generator with a fixed value (line 58 in [`communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/communities.py)), ensuring identical partitions across runs. The file-based fallback is also deterministic by design.

### What happens if my codebase has no clear modular boundaries?

The **cohesion scores** exposed in community metadata quantify boundary clarity. Low cohesion values suggest highly coupled code that resists clean modularization. The **oversized-community splitting** logic (25% threshold) also attempts to force structure even in monolithic codebases, though manual architectural review may be warranted.