# How the Leiden Community Detection Algorithm Handles Oversized Graphs in the Code‑Review‑Graph Repository

> Learn how the Leiden community detection algorithm tackles oversized graphs. Discover its recursive splitting technique for efficient community discovery in large datasets.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: internals
- Published: 2026-08-16

---

**The Leiden algorithm handles oversized graphs by first running on the full graph, then recursively splitting any community that exceeds 25% of total nodes into smaller sub-communities until all groups meet size thresholds.**

The **code-review-graph** repository by tirth8205 implements hierarchical community detection on code-knowledge graphs using the Leiden algorithm. Rather than applying a single pass and accepting potentially dominant mega-communities, the system enforces balanced partitions through recursive sub-graph analysis. This approach ensures that code review recommendations remain granular and actionable.

## Initial Community Detection Pass

The detection workflow begins with a standard Leiden run across the entire graph. In [`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py), the `detect_communities` function orchestrates this initial pass using the optional **igraph** library.

When igraph is unavailable, the routine returns communities unchanged and falls back to directory-based grouping elsewhere in the codebase. This graceful degradation preserves functionality without hard dependencies.

## Detecting and Splitting Oversized Communities

After the initial detection, each community is evaluated against two criteria:

- **`threshold_pct`** (default 0.25 or 25%): Maximum allowed proportion of total nodes
- **`min_split_size`** (default 10): Minimum community size worth splitting

Communities exceeding these limits trigger the `_split_oversized` function.

### The Recursive Splitting Process

The `_split_oversized` function operates through six distinct steps:

1. **Sub-graph construction** – Extract only the nodes belonging to the oversized community
2. **Edge weight assignment** – Apply static weights from `EDGE_WEIGHTS` (CALLS = 1.0, IMPORTS_FROM = 0.5, etc.)
3. **Deterministic Leiden rerun** – Execute with fixed RNG seed `CRG_LEIDEN_SEED` or fallback `42`
4. **Test node reassignment** – Call `_reassign_test_nodes` to keep test code cohesive
5. **Sub-community naming** – Generate human-readable names via `_generate_community_name`
6. **Cohesion recomputation** – Calculate new internal metrics for the refined groups

The original oversized community is replaced by its sub-communities, and the process repeats recursively until all communities satisfy size constraints.

## Implementation Example

```python
from code_review_graph.graph import GraphStore
from code_review_graph.communities import detect_communities

# Build graph from source files

store = GraphStore.from_path("my_project")

# Detect with automatic oversized handling

communities = detect_communities(store)

for comm in communities:
    print(f"Community {comm['id']}: {comm['name']} ({comm['size']} nodes)")

```

For manual control over splitting parameters:

```python
from code_review_graph.communities import _split_oversized

refined = _split_oversized(
    communities=initial_communities,
    nodes=store.nodes,
    edges=store.edges,
    threshold_pct=0.20,   # Stricter: 20% threshold

    min_split_size=15,    # Larger minimum group size

)

```

## Key Source Files and Their Roles

| File | Purpose |
|------|---------|
| [`code_review_graph/communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py) | Leiden detection, `_split_oversized`, naming, and cohesion logic |
| [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | `GraphNode`, `GraphEdge`, and `GraphStore` data structures |
| [`tests/test_communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_communities.py) | Unit tests for Leiden handling, fallback behavior, recursive splitting |

## Logging and Failure Handling

The implementation includes comprehensive logging for observability. Split actions generate messages like "Split oversized community … into …" for debugging and auditing.

If sub-graph creation or Leiden execution fails during splitting, the system safely falls back to preserving the original community intact rather than propagating errors.

## Summary

- **Hierarchical application**: Leiden runs globally first, then locally on oversized communities
- **Configurable thresholds**: `threshold_pct` and `min_split_size` control splitting behavior
- **Deterministic results**: Fixed RNG seed ensures reproducible community assignments
- **Graceful degradation**: igraph absence triggers fallback to directory-based grouping
- **Test-aware processing**: `_reassign_test_nodes` maintains logical code cohesion

## Frequently Asked Questions

### What triggers community splitting in the Leiden implementation?

A community splits when its node count exceeds `threshold_pct` (default 25%) of the total graph nodes **and** surpasses `min_split_size` (default 10). Both conditions must be met to avoid fragmenting small but proportionally significant groups.

### Why does the algorithm use a fixed random seed?

The `CRG_LEIDEN_SEED` environment variable (falling back to `42`) ensures **reproducible community detection**. This determinism is critical for code review workflows where consistent recommendations across runs build user trust.

### What happens if igraph is not installed?

The detection routine returns communities unchanged and relies on **directory-based grouping** elsewhere in the codebase. This optional dependency design keeps the core package lightweight while enabling advanced analysis when igraph is available.

### How does the algorithm handle test files during splitting?

The `_reassign_test_nodes` function post-processes Leiden partitions to **keep test code together** within communities. This prevents test utilities from scattering across unrelated production code groups, preserving logical project structure.