How the Leiden Community Detection Algorithm Handles Oversized Graphs in the Code‑Review‑Graph Repository

The Leiden algorithm handles oversized graphs by first running on the full graph, then recursively splitting any community that exceeds 25% of total nodes into smaller sub-communities until all groups meet size thresholds.

The code-review-graph repository by tirth8205 implements hierarchical community detection on code-knowledge graphs using the Leiden algorithm. Rather than applying a single pass and accepting potentially dominant mega-communities, the system enforces balanced partitions through recursive sub-graph analysis. This approach ensures that code review recommendations remain granular and actionable.

Initial Community Detection Pass

The detection workflow begins with a standard Leiden run across the entire graph. In code_review_graph/communities.py, the detect_communities function orchestrates this initial pass using the optional igraph library.

When igraph is unavailable, the routine returns communities unchanged and falls back to directory-based grouping elsewhere in the codebase. This graceful degradation preserves functionality without hard dependencies.

Detecting and Splitting Oversized Communities

After the initial detection, each community is evaluated against two criteria:

  • threshold_pct (default 0.25 or 25%): Maximum allowed proportion of total nodes
  • min_split_size (default 10): Minimum community size worth splitting

Communities exceeding these limits trigger the _split_oversized function.

The Recursive Splitting Process

The _split_oversized function operates through six distinct steps:

  1. Sub-graph construction – Extract only the nodes belonging to the oversized community
  2. Edge weight assignment – Apply static weights from EDGE_WEIGHTS (CALLS = 1.0, IMPORTS_FROM = 0.5, etc.)
  3. Deterministic Leiden rerun – Execute with fixed RNG seed CRG_LEIDEN_SEED or fallback 42
  4. Test node reassignment – Call _reassign_test_nodes to keep test code cohesive
  5. Sub-community naming – Generate human-readable names via _generate_community_name
  6. Cohesion recomputation – Calculate new internal metrics for the refined groups

The original oversized community is replaced by its sub-communities, and the process repeats recursively until all communities satisfy size constraints.

Implementation Example

from code_review_graph.graph import GraphStore
from code_review_graph.communities import detect_communities

# Build graph from source files

store = GraphStore.from_path("my_project")

# Detect with automatic oversized handling

communities = detect_communities(store)

for comm in communities:
    print(f"Community {comm['id']}: {comm['name']} ({comm['size']} nodes)")

For manual control over splitting parameters:

from code_review_graph.communities import _split_oversized

refined = _split_oversized(
    communities=initial_communities,
    nodes=store.nodes,
    edges=store.edges,
    threshold_pct=0.20,   # Stricter: 20% threshold

    min_split_size=15,    # Larger minimum group size

)

Key Source Files and Their Roles

File Purpose
code_review_graph/communities.py Leiden detection, _split_oversized, naming, and cohesion logic
code_review_graph/graph.py GraphNode, GraphEdge, and GraphStore data structures
tests/test_communities.py Unit tests for Leiden handling, fallback behavior, recursive splitting

Logging and Failure Handling

The implementation includes comprehensive logging for observability. Split actions generate messages like "Split oversized community … into …" for debugging and auditing.

If sub-graph creation or Leiden execution fails during splitting, the system safely falls back to preserving the original community intact rather than propagating errors.

Summary

  • Hierarchical application: Leiden runs globally first, then locally on oversized communities
  • Configurable thresholds: threshold_pct and min_split_size control splitting behavior
  • Deterministic results: Fixed RNG seed ensures reproducible community assignments
  • Graceful degradation: igraph absence triggers fallback to directory-based grouping
  • Test-aware processing: _reassign_test_nodes maintains logical code cohesion

Frequently Asked Questions

What triggers community splitting in the Leiden implementation?

A community splits when its node count exceeds threshold_pct (default 25%) of the total graph nodes and surpasses min_split_size (default 10). Both conditions must be met to avoid fragmenting small but proportionally significant groups.

Why does the algorithm use a fixed random seed?

The CRG_LEIDEN_SEED environment variable (falling back to 42) ensures reproducible community detection. This determinism is critical for code review workflows where consistent recommendations across runs build user trust.

What happens if igraph is not installed?

The detection routine returns communities unchanged and relies on directory-based grouping elsewhere in the codebase. This optional dependency design keeps the core package lightweight while enabling advanced analysis when igraph is available.

How does the algorithm handle test files during splitting?

The _reassign_test_nodes function post-processes Leiden partitions to keep test code together within communities. This prevents test utilities from scattering across unrelated production code groups, preserving logical project structure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →