# What Causes the Circular Recall Issue in Impact Accuracy Benchmarks?

> Understand the circular recall issue in impact accuracy benchmarks. Discover why using the same dependency graph for ground-truth and predictions forces recall to 1.0.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-16

---

**The circular recall issue occurs because the impact-accuracy benchmark uses the same dependency graph to define the ground-truth files and to generate predictions, forcing recall to always equal 1.0.**

The `code-review-graph` repository provides tools for **blast-radius analysis** that predict which files are affected by code changes. To evaluate these predictions, the repository includes an impact-accuracy benchmark with a notorious quirk: when run in "graph-derived" mode, recall is locked at perfect scores. Understanding why this happens—and why it matters—helps interpret benchmark results correctly.

## The Two Modes of Ground Truth

The impact accuracy benchmark operates in two distinct modes with fundamentally different ground-truth sources:

- **Graph-derived (circular) mode:** Builds the ground-truth set from files that changed plus every file with a `CALLS` or `IMPORTS_FROM` edge pointing to those changed files. This set is constructed **from the same graph** that the predictor traverses.

- **Co-change mode:** Uses files that actually changed together in the same Git commit, providing an independent, honest ground truth.

The graph-derived mode exists as a deliberate **upper-bound reference**, not as a realistic performance measurement.

## Why Recall Becomes Trivial in Graph-Derived Mode

The circular recall issue stems from a self-referential design. In [`code_review_graph/eval/benchmarks/impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/impact_accuracy.py), the `_graph_neighbor_files` function (lines 67-78) walks the exact same `CALLS` and `IMPORTS_FROM` edges that the `analyze_changes` predictor later uses.

Because the ground truth is defined as "all files reachable via these edges," and the predictor explores those same edges, the predictor can never miss a ground-truth file—it can only add extraneous predictions. This guarantees **recall = 1.0** for every successful run.

As documented in lines 5-9 of [`impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/impact_accuracy.py), this is explicitly acknowledged as a circular measurement. The README reinforces this interpretation at lines 47-52, stating that "recall 1.0 is a circular upper bound, not '100% recall'."

## Examining the Ground-Truth Construction

The `_graph_neighbor_files` function illustrates the problem directly:

```python

# From code_review_graph/eval/benchmarks/impact_accuracy.py (lines 67-78)

def _graph_neighbor_files(store, changed_files):
    """
    Build ground truth from the same graph edges used for prediction.
    This creates the circular recall condition.
    """
    ground_truth = set(changed_files)
    for f in changed_files:
        # These CALLS/IMPORTS_FROM edges are identical to what analyze_changes uses

        for incoming in store.get_edges(target=f, types=["CALLS", "IMPORTS_FROM"]):
            ground_truth.add(incoming.source)
    return ground_truth

```

When `analyze_changes` later traverses these same edges to generate predictions, it necessarily finds every file in this ground-truth set.

## Interpreting Benchmark Results Correctly

Since recall is non-informative in graph-derived mode, **precision becomes the meaningful metric**. The circular recall simply establishes how many predictions were necessary to achieve perfect coverage of the graph-defined neighborhood.

For realistic evaluation, use co-change mode:

```python
from pathlib import Path
from code_review_graph.eval.benchmarks.impact_accuracy import run, aggregate

# Configure for CO-CHANGE mode (honest ground truth)

repo_path = Path("/path/to/repo")
store = ...  # populated CodeReviewGraph store

config = {
    "name": "my-project",
    "test_commits": [{"sha": "a1b2c3d"}],
    "ground_truth": "co-change",  # KEY: use independent Git history

}
results = run(repo_path, store, config)
summary = aggregate(results)

# Now recall reflects actual predictive capability

print(summary["co_change"]["recall"])   # Realistic value (typically < 1.0)

print(summary["co_change"]["precision"]) # Still meaningful metric

```

## Summary

- The circular recall issue in impact accuracy benchmarks is **intentional and documented**, not a bug.
- It occurs because the **same dependency graph** defines both ground truth and predictions.
- **Graph-derived mode** provides an upper bound; **co-change mode** provides honest evaluation.
- Key implementation resides in [`code_review_graph/eval/benchmarks/impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/eval/benchmarks/impact_accuracy.py) at lines 67-78.
- Always verify which ground-truth mode produced reported results before interpreting recall scores.

## Frequently Asked Questions

### Why would a benchmark intentionally include a circular measurement?

The circular upper bound serves as a **sanity check and theoretical ceiling**. It answers: "What's the best possible recall if we had perfect knowledge of all dependencies?" This isolates the impact of graph coverage quality from prediction algorithm quality. Teams can compare their actual precision against this ceiling to understand how much the underlying dependency data limits their performance.

### How do I switch from graph-derived to co-change mode?

Pass `"ground_truth": "co-change"` in your benchmark configuration. The `run()` function in [`impact_accuracy.py`](https://github.com/tirth8205/code-review-graph/blob/main/impact_accuracy.py) branches on this parameter, using Git history-derived co-change data instead of graph neighbors. This requires properly configured test commits with sufficient commit history in the target repository.

### Does the circular recall issue affect other benchmarks in the repository?

No. The circular recall is **specific to impact accuracy** and its self-referential dependency graph design. Other benchmarks in the repository use independent ground-truth sources. Always consult the documentation in [`README.md`](https://github.com/tirth8205/code-review-graph/blob/main/README.md) and [`docs/REPRODUCING.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/REPRODUCING.md) to understand each benchmark's assumptions and limitations.