Git Diff Impact Mapping and Risk Classification in Codebase Memory
Codebase Memory analyzes Git diffs by traversing a stored dependency graph to identify impacted symbols and classifies each by risk level based on hop distance from modified code.
The DeusData/codebase-memory-mcp repository provides intelligent code analysis through git diff impact mapping, a feature that transforms raw file changes into actionable risk assessments. By parsing the Git working tree and walking the codebase's internal graph, the system identifies which functions, types, and constants are affected by modifications. This enables developers to prioritize review efforts based on precise risk classification rather than file counts.
How Git Diff Impact Mapping Works
The impact mapping system operates through a graph traversal algorithm that connects changed files to their downstream dependencies.
Detecting Changes in the Working Tree
The detect_changes command initiates the process by scanning the Git index and working tree for modified files. As implemented in src/mcp/mcp.c (lines 2775-2777), this function parses each changed file to extract symbols that have been added, removed, or altered.
Traversing the Dependency Graph
Once changed symbols are identified, the system executes a breadth-first search (BFS) through the stored call-graph and definition graph. For every visited node, the algorithm records the fully-qualified symbol name and the hop distance—the number of edges between the original change and the impacted symbol. This data populates the impacted_symbols array in the resulting JSON output.
Risk Classification Logic
The hop distance determines the severity level through a deterministic mapping function.
Hop-to-Risk Conversion
In src/store/store.c at line 2948, the cbm_hop_to_risk function translates hop counts into four distinct risk enums:
| Hop Distance | Risk Level | Enum Value |
|---|---|---|
| 1 | CRITICAL | CBM_RISK_CRITICAL |
| 2 | HIGH | CBM_RISK_HIGH |
| 3 | MEDIUM | CBM_RISK_MEDIUM |
| ≥ 4 or ≤ 0 | LOW | CBM_RISK_LOW |
Label Generation
The cbm_risk_label function at line 2961 in src/store/store.c converts these enums into human-readable strings. When the risk_labels flag is enabled, each entry in the impacted_symbols array includes a risk field populated with values like "CRITICAL", "HIGH", "MEDIUM", or "LOW".
Implementation Details
The core logic resides across three primary source files. The detect_changes implementation in src/mcp/mcp.c handles the BFS traversal and JSON construction. Risk level utilities live in src/store/store.c, while the command-line interface exposing the risk_labels option appears in src/cli/cli.c at line 460.
/* Example: enabling risk classification programmatically */
int main(int argc, char **argv) {
/* ... argument parsing ... */
bool risk_labels = true; // request risk classification
detect_changes(store, risk_labels); // traverses graph and assigns risk levels
return 0;
}
CLI Usage and Output
Developers enable risk classification through the CLI flag documented in the source:
cbm detect_changes --risk_labels=true
The resulting JSON structure follows this format:
{
"impacted_symbols": [
{
"symbol": "my_project::utils::parse_input",
"hop": 1,
"risk": "CRITICAL"
},
{
"symbol": "my_project::core::process",
"hop": 3,
"risk": "MEDIUM"
}
]
}
Summary
- Git diff impact mapping converts file changes into a graph of affected symbols using BFS traversal through the stored dependency graph.
- Risk classification assigns CRITICAL, HIGH, MEDIUM, or LOW labels based on hop distance from modified code, as defined in
cbm_hop_to_risk. - The
cbm_hop_to_riskandcbm_risk_labelfunctions insrc/store/store.cimplement the severity mapping at lines 2948 and 2961. - Enable the feature via
--risk_labels=trueto receive risk-aware JSON output for impact-aware code reviews.
Frequently Asked Questions
What is git diff impact mapping in Codebase Memory?
Git diff impact mapping is the process of analyzing changes in the Git working tree to identify which symbols (functions, types, constants) are affected by those modifications. The system traverses the codebase's stored dependency graph to calculate the "hop distance" between changed files and dependent symbols, producing a machine-readable impact report.
How does Codebase Memory classify risk levels?
Risk levels are determined by the cbm_hop_to_risk function based on how many edges separate a symbol from the original change. A hop distance of 1 yields CRITICAL risk, 2 yields HIGH, 3 yields MEDIUM, and distances of 4 or greater (or zero/negative) yield LOW risk. The cbm_risk_label function converts these internal enums to printable strings like "CRITICAL" or "HIGH".
Where is the risk classification implemented in the source code?
The risk classification logic resides in src/store/store.c, specifically at line 2948 for cbm_hop_to_risk and line 2961 for cbm_risk_label. The BFS traversal that generates the impact data is implemented in src/mcp/mcp.c around lines 2775-2777, while the CLI interface providing the --risk_labels flag is located in src/cli/cli.c at line 460.
How do I enable risk labels in the output?
Pass the --risk_labels=true flag when running the detect_changes command. This instructs the system to include the risk field in each entry of the impacted_symbols JSON array, allowing you to filter or prioritize symbols based on their calculated risk level.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →