How the Leiden Community Detection Algorithm Handles Oversized Graphs in the Code‑Review‑Graph Repository
The Leiden algorithm handles oversized graphs by first running on the full graph, then recursively splitting any community that exceeds 25% of total nodes into smaller sub-communities until all groups meet size thresholds.
The code-review-graph repository by tirth8205 implements hierarchical community detection on code-knowledge graphs using the Leiden algorithm. Rather than applying a single pass and accepting potentially dominant mega-communities, the system enforces balanced partitions through recursive sub-graph analysis. This approach ensures that code review recommendations remain granular and actionable.
Initial Community Detection Pass
The detection workflow begins with a standard Leiden run across the entire graph. In code_review_graph/communities.py, the detect_communities function orchestrates this initial pass using the optional igraph library.
When igraph is unavailable, the routine returns communities unchanged and falls back to directory-based grouping elsewhere in the codebase. This graceful degradation preserves functionality without hard dependencies.
Detecting and Splitting Oversized Communities
After the initial detection, each community is evaluated against two criteria:
threshold_pct(default 0.25 or 25%): Maximum allowed proportion of total nodesmin_split_size(default 10): Minimum community size worth splitting
Communities exceeding these limits trigger the _split_oversized function.
The Recursive Splitting Process
The _split_oversized function operates through six distinct steps:
- Sub-graph construction – Extract only the nodes belonging to the oversized community
- Edge weight assignment – Apply static weights from
EDGE_WEIGHTS(CALLS = 1.0, IMPORTS_FROM = 0.5, etc.) - Deterministic Leiden rerun – Execute with fixed RNG seed
CRG_LEIDEN_SEEDor fallback42 - Test node reassignment – Call
_reassign_test_nodesto keep test code cohesive - Sub-community naming – Generate human-readable names via
_generate_community_name - Cohesion recomputation – Calculate new internal metrics for the refined groups
The original oversized community is replaced by its sub-communities, and the process repeats recursively until all communities satisfy size constraints.
Implementation Example
from code_review_graph.graph import GraphStore
from code_review_graph.communities import detect_communities
# Build graph from source files
store = GraphStore.from_path("my_project")
# Detect with automatic oversized handling
communities = detect_communities(store)
for comm in communities:
print(f"Community {comm['id']}: {comm['name']} ({comm['size']} nodes)")
For manual control over splitting parameters:
from code_review_graph.communities import _split_oversized
refined = _split_oversized(
communities=initial_communities,
nodes=store.nodes,
edges=store.edges,
threshold_pct=0.20, # Stricter: 20% threshold
min_split_size=15, # Larger minimum group size
)
Key Source Files and Their Roles
| File | Purpose |
|---|---|
code_review_graph/communities.py |
Leiden detection, _split_oversized, naming, and cohesion logic |
code_review_graph/graph.py |
GraphNode, GraphEdge, and GraphStore data structures |
tests/test_communities.py |
Unit tests for Leiden handling, fallback behavior, recursive splitting |
Logging and Failure Handling
The implementation includes comprehensive logging for observability. Split actions generate messages like "Split oversized community … into …" for debugging and auditing.
If sub-graph creation or Leiden execution fails during splitting, the system safely falls back to preserving the original community intact rather than propagating errors.
Summary
- Hierarchical application: Leiden runs globally first, then locally on oversized communities
- Configurable thresholds:
threshold_pctandmin_split_sizecontrol splitting behavior - Deterministic results: Fixed RNG seed ensures reproducible community assignments
- Graceful degradation: igraph absence triggers fallback to directory-based grouping
- Test-aware processing:
_reassign_test_nodesmaintains logical code cohesion
Frequently Asked Questions
What triggers community splitting in the Leiden implementation?
A community splits when its node count exceeds threshold_pct (default 25%) of the total graph nodes and surpasses min_split_size (default 10). Both conditions must be met to avoid fragmenting small but proportionally significant groups.
Why does the algorithm use a fixed random seed?
The CRG_LEIDEN_SEED environment variable (falling back to 42) ensures reproducible community detection. This determinism is critical for code review workflows where consistent recommendations across runs build user trust.
What happens if igraph is not installed?
The detection routine returns communities unchanged and relies on directory-based grouping elsewhere in the codebase. This optional dependency design keeps the core package lightweight while enabling advanced analysis when igraph is available.
How does the algorithm handle test files during splitting?
The _reassign_test_nodes function post-processes Leiden partitions to keep test code together within communities. This prevents test utilities from scattering across unrelated production code groups, preserving logical project structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →