What Causes the Circular Recall Issue in Impact Accuracy Benchmarks?
The circular recall issue occurs because the impact-accuracy benchmark uses the same dependency graph to define the ground-truth files and to generate predictions, forcing recall to always equal 1.0.
The code-review-graph repository provides tools for blast-radius analysis that predict which files are affected by code changes. To evaluate these predictions, the repository includes an impact-accuracy benchmark with a notorious quirk: when run in "graph-derived" mode, recall is locked at perfect scores. Understanding why this happens—and why it matters—helps interpret benchmark results correctly.
The Two Modes of Ground Truth
The impact accuracy benchmark operates in two distinct modes with fundamentally different ground-truth sources:
-
Graph-derived (circular) mode: Builds the ground-truth set from files that changed plus every file with a
CALLSorIMPORTS_FROMedge pointing to those changed files. This set is constructed from the same graph that the predictor traverses. -
Co-change mode: Uses files that actually changed together in the same Git commit, providing an independent, honest ground truth.
The graph-derived mode exists as a deliberate upper-bound reference, not as a realistic performance measurement.
Why Recall Becomes Trivial in Graph-Derived Mode
The circular recall issue stems from a self-referential design. In code_review_graph/eval/benchmarks/impact_accuracy.py, the _graph_neighbor_files function (lines 67-78) walks the exact same CALLS and IMPORTS_FROM edges that the analyze_changes predictor later uses.
Because the ground truth is defined as "all files reachable via these edges," and the predictor explores those same edges, the predictor can never miss a ground-truth file—it can only add extraneous predictions. This guarantees recall = 1.0 for every successful run.
As documented in lines 5-9 of impact_accuracy.py, this is explicitly acknowledged as a circular measurement. The README reinforces this interpretation at lines 47-52, stating that "recall 1.0 is a circular upper bound, not '100% recall'."
Examining the Ground-Truth Construction
The _graph_neighbor_files function illustrates the problem directly:
# From code_review_graph/eval/benchmarks/impact_accuracy.py (lines 67-78)
def _graph_neighbor_files(store, changed_files):
"""
Build ground truth from the same graph edges used for prediction.
This creates the circular recall condition.
"""
ground_truth = set(changed_files)
for f in changed_files:
# These CALLS/IMPORTS_FROM edges are identical to what analyze_changes uses
for incoming in store.get_edges(target=f, types=["CALLS", "IMPORTS_FROM"]):
ground_truth.add(incoming.source)
return ground_truth
When analyze_changes later traverses these same edges to generate predictions, it necessarily finds every file in this ground-truth set.
Interpreting Benchmark Results Correctly
Since recall is non-informative in graph-derived mode, precision becomes the meaningful metric. The circular recall simply establishes how many predictions were necessary to achieve perfect coverage of the graph-defined neighborhood.
For realistic evaluation, use co-change mode:
from pathlib import Path
from code_review_graph.eval.benchmarks.impact_accuracy import run, aggregate
# Configure for CO-CHANGE mode (honest ground truth)
repo_path = Path("/path/to/repo")
store = ... # populated CodeReviewGraph store
config = {
"name": "my-project",
"test_commits": [{"sha": "a1b2c3d"}],
"ground_truth": "co-change", # KEY: use independent Git history
}
results = run(repo_path, store, config)
summary = aggregate(results)
# Now recall reflects actual predictive capability
print(summary["co_change"]["recall"]) # Realistic value (typically < 1.0)
print(summary["co_change"]["precision"]) # Still meaningful metric
Summary
- The circular recall issue in impact accuracy benchmarks is intentional and documented, not a bug.
- It occurs because the same dependency graph defines both ground truth and predictions.
- Graph-derived mode provides an upper bound; co-change mode provides honest evaluation.
- Key implementation resides in
code_review_graph/eval/benchmarks/impact_accuracy.pyat lines 67-78. - Always verify which ground-truth mode produced reported results before interpreting recall scores.
Frequently Asked Questions
Why would a benchmark intentionally include a circular measurement?
The circular upper bound serves as a sanity check and theoretical ceiling. It answers: "What's the best possible recall if we had perfect knowledge of all dependencies?" This isolates the impact of graph coverage quality from prediction algorithm quality. Teams can compare their actual precision against this ceiling to understand how much the underlying dependency data limits their performance.
How do I switch from graph-derived to co-change mode?
Pass "ground_truth": "co-change" in your benchmark configuration. The run() function in impact_accuracy.py branches on this parameter, using Git history-derived co-change data instead of graph neighbors. This requires properly configured test commits with sufficient commit history in the target repository.
Does the circular recall issue affect other benchmarks in the repository?
No. The circular recall is specific to impact accuracy and its self-referential dependency graph design. Other benchmarks in the repository use independent ground-truth sources. Always consult the documentation in README.md and docs/REPRODUCING.md to understand each benchmark's assumptions and limitations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →