# How Knowledge Gap Detection Identifies Structural Weaknesses in Code Using code‑review‑graph

> Discover how knowledge gap detection in code-review-graph reveals hidden structural weaknesses in your codebase by analyzing its knowledge graph. Find architectural soft spots.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-15

---

**Knowledge gap detection in `code-review-graph` builds a persistent SQLite-backed knowledge graph of your codebase and analyzes node connectivity, community structure, and test coverage to flag architectural "soft spots" that are invisible in raw source code.**

The `code-review-graph` tool transforms your repository into a queryable knowledge graph where every source file, class, function, type, and test becomes a node, and relationships like imports, calls, and test coverage become edges. The **knowledge gap detector** (`find_knowledge_gaps`) traverses this graph to surface four distinct categories of structural weakness that threaten code maintainability.

## How the Knowledge Graph Enables Gap Detection

At the heart of knowledge gap detection lies the `GraphStore` class defined in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py). This SQLite-backed store persists:

- **Nodes**: Code entities (files, classes, functions, types, tests)
- **Edges**: Relationships (imports, calls, inheritance, `TESTED_BY` links)
- **Communities**: Pre-computed cluster assignments for modularity analysis

The detector consumes this structured data without re-parsing source files, making repeated analyses fast and consistent.

## Four Structural Weaknesses the Detector Uncovers

The `find_knowledge_gaps` function in [`code_review_graph/analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/analysis.py) implements four detection strategies. Each targets a specific architectural anti-pattern.

### Isolated Nodes: Code That Lives in Isolation

The detector calculates the **total degree** (inbound plus outbound connections) for every non-File node by iterating through all edges:

```python

# From analysis.py lines 28-42

degree: Dict[str, int] = {}
for e in edges:
    degree[e.source_qualified] = degree.get(e.source_qualified, 0) + 1
    degree[e.target_qualified] = degree.get(e.target_qualified, 0) + 1

isolated = [
    node for node in nodes
    if not node.is_file and degree.get(node.qualified_name, 0) <= 1
]

```

Nodes with **degree ≤ 1** are flagged as isolated. These represent:

- Dead code that nothing calls and that calls nothing
- Hidden entry points accidentally excluded from the main flow
- Poorly integrated utility functions that should be more discoverable

### Thin Communities: Over-Fragmented Modules

The detector leverages community detection already computed in the graph store:

```python

# From analysis.py lines 59-71

comm_sizes: Dict[str, int] = {}
for node_id, comm_id in store.get_all_community_ids():
    comm_sizes[comm_id] = comm_sizes.get(comm_id, 0) + 1

thin_communities = [
    comm_id for comm_id, size in comm_sizes.items()
    if size < 3
]

```

Communities with **fewer than three members** indicate modules that may be overly granular. These fragments could be consolidated for better cohesion or may signal extraction boundaries drawn too aggressively.

### Untested Hotspots: High-Risk, High-Connectivity Code

The detector identifies dangerous gaps in test coverage by cross-referencing node connectivity with `TESTED_BY` edges:

```python

# From analysis.py lines 72-80

tested_nodes = {e.source_qualified for e in edges if e.rel_type == "TESTED_BY"}

untested_hotspots = [
    node for node in nodes
    if (not node.is_test
        and degree.get(node.qualified_name, 0) >= 5
        and node.qualified_name not in tested_nodes)
]

```

Nodes with **degree ≥ 5** that lack test coverage represent architectural hotspots where regressions would have cascading effects. These demand prioritization in test planning.

### Single-File Communities: God Objects in Disguise

For each community, the detector tracks how many distinct files host its members:

```python

# From analysis.py lines 91-100

comm_files: Dict[str, Set[str]] = defaultdict(set)
for node in nodes:
    if node.community_id:
        comm_files[node.community_id].add(node.file_path)

single_file_communities = [
    comm_id for comm_id in comm_files
    if len(comm_files[comm_id]) == 1 and comm_sizes[comm_id] >= 3
]

```

Communities of **three or more nodes confined to a single file** often signal a "God object" or file with too many responsibilities. This structural clustering hurts readability and complicates future extension.

## The Three-Stage Detection Pipeline

Knowledge gap detection operates in discrete phases that build analytical momentum:

1. **Graph traversal** ([`analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis.py) lines 24-34): Fetches all edges via `store.get_all_edges()` and constructs the degree map while simultaneously collecting `TESTED_BY` relationships for coverage analysis.

2. **Community analysis** ([`analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis.py) lines 50-60): Retrieves community assignments with `store.get_all_community_ids()`, then aggregates member counts and file distributions to identify both thin communities and single-file clusters.

3. **Gap synthesis** ([`analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis.py) lines 102-110): Assembles findings into a structured dictionary consumable by CLI renderers, PR comment generators, or downstream automation.

## Running Knowledge Gap Detection

### Python API

Execute detection programmatically after building your graph:

```python
from code_review_graph.main import build_graph
from code_review_graph.analysis import find_knowledge_gaps

# Build or update the knowledge graph

graph_store = build_graph(paths=["."])

# Run the detector

gaps = find_knowledge_gaps(graph_store)

# Inspect results

print("Isolated nodes:", gaps["isolated_nodes"][:5])
print("Thin communities:", gaps["thin_communities"])
print("Untested hotspots:", gaps["untested_hotspots"])
print("Single-file communities:", gaps["single_file_communities"])

```

### CLI Interface

The same functionality exposes through a command-line tool:

```bash
code-review-graph get_knowledge_gaps

```

This emits a JSON summary suitable for piping to other tools or CI systems.

### PR Comment Integration

For automated code review workflows:

```python
from code_review_graph.main import get_knowledge_gaps_tool

comment = get_knowledge_gaps_tool()
print(comment)  # Markdown-formatted for GitHub comments

```

The `get_knowledge_gaps_tool` wrapper (registered in [`main.py`](https://github.com/tirth8205/code-review-graph/blob/main/main.py) and implemented via [`tools/analysis_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/tools/analysis_tools.py)) formats findings as actionable review commentary.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`code_review_graph/analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/analysis.py) | Core `find_knowledge_gaps` implementation with all four detection algorithms |
| [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | `GraphStore` class for SQLite persistence and community ID handling |
| [`code_review_graph/main.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/main.py) | Orchestration layer and CLI command registration |
| [`code_review_graph/tools/analysis_tools.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/analysis_tools.py) | Tool-based API wrapper for MCP integration |

These components work together to transform raw graph topology into maintainability intelligence that would require hours of manual code review to replicate.

## Summary

- **Knowledge gap detection** operates on a persistent SQLite-backed knowledge graph of code entities and relationships
- **Four weakness categories**: isolated nodes (degree ≤ 1), thin communities (< 3 members), untested hotspots (degree ≥ 5, no tests), and single-file communities (≥ 3 nodes, one file)
- **Three-stage pipeline**: graph traversal → community analysis → gap synthesis
- **Multiple interfaces**: Python API, CLI command, and PR comment generation
- **Source locations**: All detection logic lives in [`analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis.py) with storage in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) and orchestration in [`main.py`](https://github.com/tirth8205/code-review-graph/blob/main/main.py)

## Frequently Asked Questions

### What makes knowledge gap detection different from traditional static analysis?

Traditional static analyzers flag syntactic issues or style violations. Knowledge gap detection analyzes **architectural topology**—how code entities connect and cluster—revealing structural problems like hidden dependencies, coverage gaps in critical paths, and modularity failures that linters cannot see.

### How does the detector handle large codebases efficiently?

The detector leverages the pre-built knowledge graph stored in SQLite. Rather than re-parsing source files, it executes targeted SQL queries via `GraphStore` methods like `get_all_edges()` and `get_all_community_ids()`. This makes repeated analyses fast even for repositories with thousands of files.

### Can I customize the thresholds for what constitutes a "gap"?

The current implementation uses fixed thresholds (degree ≤ 1 for isolation, degree ≥ 5 for hotspots, community size < 3, single-file communities ≥ 3 nodes). These are hardcoded in [`analysis.py`](https://github.com/tirth8205/code-review-graph/blob/main/analysis.py). For custom analysis, you can fork the `find_knowledge_gaps` function and adjust the comparison operators while preserving the same graph traversal patterns.