# How Community Detection with the Leiden Algorithm Powers Code‑Review‑Graph's Architecture Analysis

> Discover how community detection with the Leiden algorithm structures code-review-graph. It clusters related code units, enhancing architecture analysis with modularity maximization and reliable fallbacks.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: architecture
- Published: 2026-08-11

---

**The Leiden algorithm in code‑review‑graph clusters related code units into communities by maximizing modularity on a weighted dependency graph, with deterministic output and automatic fallback to directory‑based grouping when igraph is unavailable.**

The `tirth8205/code-review‑graph` project builds a code‑knowledge graph from repository structure and then applies **community detection** to surface architectural patterns. The Leiden algorithm serves as its primary clustering engine, translating dependency relationships into meaningful, named communities that represent tightly‑coupled functional units.

## What the Leiden Algorithm Does in Code‑Review‑Graph

Community detection transforms a complex web of code dependencies into human‑readable clusters. In [`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py), the implementation treats each function, class, or module as a graph node and each `CALLS`, `IMPORTS_FROM`, or `INHERITS_FROM` relationship as a weighted, undirected edge. The Leiden algorithm then optimizes **modularity**—a measure of how densely connected nodes are within groups versus between them—to find the natural boundaries of cohesive code.

The algorithm is particularly well‑suited for this use case because it produces **higher‑quality partitions** than the older Louvain method, with guarantees that communities are well‑connected and that the process converges efficiently even on large codebases.

## Step‑by‑Step Implementation of Community Detection

The detection pipeline in `detect_communities(store, min_size=2)` follows a rigorous sequence:

### Optional igraph Import and Fallback

The system first attempts to import `igraph` at line 33 of [`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py). If the library is absent, it seamlessly falls back to `_detect_file_based`, which groups nodes by directory depth instead:

```python
try:
    import igraph
    HAS_IGRAPH = True
except ImportError:
    HAS_IGRAPH = False

```

### Building Node and Edge Mappings

Before invoking the algorithm, the code constructs translation dictionaries to interface with igraph's integer‑based indexing. At line 68, it builds:

- `qn_to_idx`: maps qualified node names (e.g., `mymodule.MyClass.method`) to vertex indices
- `idx_to_node`: reverse mapping back to `GraphNode` objects

### Weighted Edge Assembly

Starting at line 84, the system iterates through all `GraphEdge` instances, converts source/target qualified names to indices, deduplicates undirected edges, and assigns **per‑kind weights**—so a `CALLS` edge may carry different influence than an `IMPORTS_FROM` edge in the clustering calculation.

### Resolution Scaling for Repository Size

At line 100, the resolution parameter (controlling community granularity) is computed as inversely proportional to the logarithm of node count:

```python
resolution = 1.0 / math.log(max(len(nodes), 10))

```

This yields **coarser clusters for massive repositories** and finer‑grained communities for smaller codebases.

### Deterministic Seeding

Reproducibility is enforced at line 110 via the `CRG_LEIDEN_SEED` environment variable (default: `42`). The seed is injected into igraph's random number generator so identical graphs always produce identical community IDs.

### Running Leiden with Bounded Iterations

The core call at line 116 uses fixed parameters optimized for code graphs:

```python
partition = g.community_leiden(
    objective_function="modularity",
    weights="weight",
    resolution=resolution,
    n_iterations=2
)

```

The **hard‑coded 2 iterations** prevent exponential runtimes on dense dependency graphs while still achieving stable, high‑quality partitions.

## Post‑Processing: Test Nodes, Cohesion, and Hierarchy

### Reassigning Test Nodes

After raw detection, the `_reassign_test_nodes` function (line 61) migrates test‑only nodes to the community containing the most **unique subjects they test**. This preserves deterministic output while placing tests near their relevant production code.

### Cohesion Scoring via Batch Computation

Community quality is assessed using `_compute_cohesion_batch` at line 87. This single‑pass O(edges) routine tallies internal versus external edges for all communities simultaneously—avoiding the N² penalty of per‑community scans. The resulting **cohesion score** (internal edges ÷ total edges) identifies well‑encapsulated modules versus tangled architectural hotspots.

### Recursive Splitting of Oversized Communities

Communities exceeding 25% of total node count trigger `_split_oversized` (line 66), which recursively re‑runs Leiden on the sub‑graph. This produces a **hierarchy of sub‑communities** for monolithic codebases without manual intervention.

### Name Generation and Deduplication

Each community receives a readable name based on dominant language and shared path prefixes. The `_dedupe_community_names` function (line 332) resolves collisions by appending distinctive keywords or numeric suffixes.

## Detecting Communities in Practice

The public API requires minimal boilerplate:

```python
from code_review_graph.graph import GraphStore
from code_review_graph.communities import detect_communities, get_architecture_overview

# Load your graph from a previously analyzed repository

store = GraphStore(path="my_repo_graph.db")

# Run detection—Leiden activates automatically if igraph is installed

communities = detect_communities(store, min_size=5)

# Inspect results

for c in communities:
    print(f"{c['name']}: {c['size']} nodes, cohesion={c['cohesion']:.2f}")

# Generate cross‑community coupling warnings

overview = get_architecture_overview(store)
for warning in overview["warnings"]:
    print(f"Architectural concern: {warning}")

```

If `igraph` is missing, the identical call transparently uses the directory‑based detector—no code changes required.

## Key Configuration and Environment Variables

| Variable | Purpose | Default |
|----------|---------|---------|
| `CRG_LEIDEN_SEED` | RNG seed for deterministic community IDs | `42` |
| `min_size` parameter | Minimum nodes for a community to be returned | `2` |

The resolution scaling and 2‑iteration bound are currently hard‑coded but tuned for typical software repository sizes (10³–10⁶ nodes).

## Summary

- **Leiden algorithm** maximizes modularity on weighted dependency graphs to identify natural code boundaries
- **Automatic fallback** to directory‑based grouping when igraph is unavailable
- **Deterministic output** via `CRG_LEIDEN_SEED` and bounded 2‑iteration convergence
- **Test‑aware post‑processing** reassigns test files to their most‑tested production communities
- **Batch cohesion scoring** computes cluster quality in O(edges) time
- **Recursive splitting** handles oversized communities without manual tuning
- **Entry point** `detect_communities()` in [`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py) orchestrates the pipeline

## Frequently Asked Questions

### What happens if igraph is not installed?

The system falls back to `_detect_file_based`, which groups nodes by directory depth and applies the same cohesion and naming logic. The public `detect_communities()` API remains unchanged—callers need no conditional code.

### Why is the iteration count fixed at 2?

Two iterations provide stable, high‑quality partitions for code‑dependency graphs while preventing exponential runtime growth on dense edges. The developers empirically validated this trade‑off against repository sizes ranging from small libraries to monorepos exceeding 100,000 nodes.

### How does cohesion relate to architectural quality?

Cohesion measures the fraction of a community's edges that stay internal versus connect outward. Values near 1.0 indicate well‑encapsulated modules; low values flag architectural tangles that may benefit from refactoring. The batch computation makes this metric practical even for large graphs.

### Can I customize the resolution parameter for finer or coarser clusters?

Currently, resolution is auto‑scaled via `1.0 / log(node_count)`. Direct override is not exposed in the public API, though the logarithmic scaling naturally adapts to repository size. Users requiring specific granularity can pre‑filter nodes or post‑process the returned communities.