How Knowledge Gap Detection Identifies Structural Weaknesses in Code Using code‑review‑graph

Knowledge gap detection in code-review-graph builds a persistent SQLite-backed knowledge graph of your codebase and analyzes node connectivity, community structure, and test coverage to flag architectural "soft spots" that are invisible in raw source code.

The code-review-graph tool transforms your repository into a queryable knowledge graph where every source file, class, function, type, and test becomes a node, and relationships like imports, calls, and test coverage become edges. The knowledge gap detector (find_knowledge_gaps) traverses this graph to surface four distinct categories of structural weakness that threaten code maintainability.

How the Knowledge Graph Enables Gap Detection

At the heart of knowledge gap detection lies the GraphStore class defined in code_review_graph/graph.py. This SQLite-backed store persists:

  • Nodes: Code entities (files, classes, functions, types, tests)
  • Edges: Relationships (imports, calls, inheritance, TESTED_BY links)
  • Communities: Pre-computed cluster assignments for modularity analysis

The detector consumes this structured data without re-parsing source files, making repeated analyses fast and consistent.

Four Structural Weaknesses the Detector Uncovers

The find_knowledge_gaps function in code_review_graph/analysis.py implements four detection strategies. Each targets a specific architectural anti-pattern.

Isolated Nodes: Code That Lives in Isolation

The detector calculates the total degree (inbound plus outbound connections) for every non-File node by iterating through all edges:


# From analysis.py lines 28-42

degree: Dict[str, int] = {}
for e in edges:
    degree[e.source_qualified] = degree.get(e.source_qualified, 0) + 1
    degree[e.target_qualified] = degree.get(e.target_qualified, 0) + 1

isolated = [
    node for node in nodes
    if not node.is_file and degree.get(node.qualified_name, 0) <= 1
]

Nodes with degree ≤ 1 are flagged as isolated. These represent:

  • Dead code that nothing calls and that calls nothing
  • Hidden entry points accidentally excluded from the main flow
  • Poorly integrated utility functions that should be more discoverable

Thin Communities: Over-Fragmented Modules

The detector leverages community detection already computed in the graph store:


# From analysis.py lines 59-71

comm_sizes: Dict[str, int] = {}
for node_id, comm_id in store.get_all_community_ids():
    comm_sizes[comm_id] = comm_sizes.get(comm_id, 0) + 1

thin_communities = [
    comm_id for comm_id, size in comm_sizes.items()
    if size < 3
]

Communities with fewer than three members indicate modules that may be overly granular. These fragments could be consolidated for better cohesion or may signal extraction boundaries drawn too aggressively.

Untested Hotspots: High-Risk, High-Connectivity Code

The detector identifies dangerous gaps in test coverage by cross-referencing node connectivity with TESTED_BY edges:


# From analysis.py lines 72-80

tested_nodes = {e.source_qualified for e in edges if e.rel_type == "TESTED_BY"}

untested_hotspots = [
    node for node in nodes
    if (not node.is_test
        and degree.get(node.qualified_name, 0) >= 5
        and node.qualified_name not in tested_nodes)
]

Nodes with degree ≥ 5 that lack test coverage represent architectural hotspots where regressions would have cascading effects. These demand prioritization in test planning.

Single-File Communities: God Objects in Disguise

For each community, the detector tracks how many distinct files host its members:


# From analysis.py lines 91-100

comm_files: Dict[str, Set[str]] = defaultdict(set)
for node in nodes:
    if node.community_id:
        comm_files[node.community_id].add(node.file_path)

single_file_communities = [
    comm_id for comm_id in comm_files
    if len(comm_files[comm_id]) == 1 and comm_sizes[comm_id] >= 3
]

Communities of three or more nodes confined to a single file often signal a "God object" or file with too many responsibilities. This structural clustering hurts readability and complicates future extension.

The Three-Stage Detection Pipeline

Knowledge gap detection operates in discrete phases that build analytical momentum:

  1. Graph traversal (analysis.py lines 24-34): Fetches all edges via store.get_all_edges() and constructs the degree map while simultaneously collecting TESTED_BY relationships for coverage analysis.

  2. Community analysis (analysis.py lines 50-60): Retrieves community assignments with store.get_all_community_ids(), then aggregates member counts and file distributions to identify both thin communities and single-file clusters.

  3. Gap synthesis (analysis.py lines 102-110): Assembles findings into a structured dictionary consumable by CLI renderers, PR comment generators, or downstream automation.

Running Knowledge Gap Detection

Python API

Execute detection programmatically after building your graph:

from code_review_graph.main import build_graph
from code_review_graph.analysis import find_knowledge_gaps

# Build or update the knowledge graph

graph_store = build_graph(paths=["."])

# Run the detector

gaps = find_knowledge_gaps(graph_store)

# Inspect results

print("Isolated nodes:", gaps["isolated_nodes"][:5])
print("Thin communities:", gaps["thin_communities"])
print("Untested hotspots:", gaps["untested_hotspots"])
print("Single-file communities:", gaps["single_file_communities"])

CLI Interface

The same functionality exposes through a command-line tool:

code-review-graph get_knowledge_gaps

This emits a JSON summary suitable for piping to other tools or CI systems.

PR Comment Integration

For automated code review workflows:

from code_review_graph.main import get_knowledge_gaps_tool

comment = get_knowledge_gaps_tool()
print(comment)  # Markdown-formatted for GitHub comments

The get_knowledge_gaps_tool wrapper (registered in main.py and implemented via tools/analysis_tools.py) formats findings as actionable review commentary.

Key Implementation Files

File Purpose
code_review_graph/analysis.py Core find_knowledge_gaps implementation with all four detection algorithms
code_review_graph/graph.py GraphStore class for SQLite persistence and community ID handling
code_review_graph/main.py Orchestration layer and CLI command registration
code_review_graph/tools/analysis_tools.py Tool-based API wrapper for MCP integration

These components work together to transform raw graph topology into maintainability intelligence that would require hours of manual code review to replicate.

Summary

  • Knowledge gap detection operates on a persistent SQLite-backed knowledge graph of code entities and relationships
  • Four weakness categories: isolated nodes (degree ≤ 1), thin communities (< 3 members), untested hotspots (degree ≥ 5, no tests), and single-file communities (≥ 3 nodes, one file)
  • Three-stage pipeline: graph traversal → community analysis → gap synthesis
  • Multiple interfaces: Python API, CLI command, and PR comment generation
  • Source locations: All detection logic lives in analysis.py with storage in graph.py and orchestration in main.py

Frequently Asked Questions

What makes knowledge gap detection different from traditional static analysis?

Traditional static analyzers flag syntactic issues or style violations. Knowledge gap detection analyzes architectural topology—how code entities connect and cluster—revealing structural problems like hidden dependencies, coverage gaps in critical paths, and modularity failures that linters cannot see.

How does the detector handle large codebases efficiently?

The detector leverages the pre-built knowledge graph stored in SQLite. Rather than re-parsing source files, it executes targeted SQL queries via GraphStore methods like get_all_edges() and get_all_community_ids(). This makes repeated analyses fast even for repositories with thousands of files.

Can I customize the thresholds for what constitutes a "gap"?

The current implementation uses fixed thresholds (degree ≤ 1 for isolation, degree ≥ 5 for hotspots, community size < 3, single-file communities ≥ 3 nodes). These are hardcoded in analysis.py. For custom analysis, you can fork the find_knowledge_gaps function and adjust the comparison operators while preserving the same graph traversal patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →