# How Edge Confidence Scoring Categorizes Edge Types in Code-Review-Graph

> Discover how edge confidence scoring categorizes edge types in code-review-graph into EXTRACTED and INFERRED tiers. Learn about the numeric confidence scores.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-18

---

**The edge confidence scoring system in code-review-graph categorizes edges into two confidence tiers: `EXTRACTED` for edges directly parsed from source code, and `INFERRED` for edges generated by resolver analysis, with each tier paired with a numeric `confidence` score from 0 to 1.**

Understanding how edge confidence scoring categorizes edge types is essential when working with the **code-review-graph** repository, a tool that builds dependency graphs from static code analysis. The system uses a dual-layer confidence model that separates *how* an edge originated from *how certain* the system is about that relationship. This design lets downstream tools filter or weight edges differently based on their provenance.

---

## The Two-Field Confidence Model

Every edge in the graph stores **two distinct confidence values** in the SQLite database, as defined in [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) and documented in [`docs/schema.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/schema.md):

| Field | Type | Default | Purpose |
|-------|------|---------|---------|
| `confidence` | `float` | `1.0` | Numeric certainty (0–1) for fine-grained weighting |
| `confidence_tier` | `str` | `"EXTRACTED"` | Categorical origin classification |

These fields were added to the `edges` table in **migration v9** ([`code_review_graph/migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/migrations.py)), enabling the graph to distinguish between ground-truth extractions and analytical inferences.

---

## The Two Confidence Tiers

The `confidence_tier` field implements exactly **two edge type categories**:

### EXTRACTED Tier

Edges with `confidence_tier: "EXTRACTED"` are taken **verbatim from the source code** by the static parser. These represent direct observations—function calls, imports, or references that appear literally in the analyzed code.

- Set automatically as the default in `graph.add_edge()`
- Considered fully reliable (`confidence` typically `1.0`)
- No resolver mediation involved

### INFERRED Tier

Edges with `confidence_tier: "INFERRED"` are **generated by resolver passes** that rewrite or deduce relationships based on higher-level analysis. These edges originate from computational inference rather than direct parsing.

- Explicitly set by resolver code in [`code_review_graph/scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/scoped_resolver.py) and [`code_review_graph/event_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/event_resolver.py)
- Typically paired with reduced numeric confidence (e.g., `0.95`)
- May represent scoped resolutions or event-based relationships not visible in raw code

---

## How Resolvers Assign the INFERRED Tier

The categorization logic lives in two resolver implementations:

### Scoped Resolver ([`code_review_graph/scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/scoped_resolver.py))

When the scoped resolver rewrites edges—for example, resolving a call to its target across module boundaries—it explicitly tags them as inferred. The source code comment states: *"Rewritten edges are tagged `confidence_tier = INFERRED`"*. The accompanying SQL update statements execute:

```python

# From scoped_resolver.py - edge rewrite operation

UPDATE edges 
SET confidence_tier = 'INFERRED', 
    confidence = :confidence 
WHERE edge_id = :edge_id

```

### Event Resolver ([`code_review_graph/event_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/event_resolver.py))

For event-based analysis, inferred edges are created directly with both reduced confidence and the inferred tier:

```python

# Creating an inferred event edge in event_resolver.py

graph.add_edge(
    source_qualified="event_handler.on_click",
    target_qualified="event_handler.process",
    line=15,
    extra={"event_type": "click"},
    confidence=0.95,           # Slightly reduced certainty

    confidence_tier="INFERRED" # Marked as analytically derived

)

```

---

## Practical Code Examples

### Default Extracted Edge

```python
from code_review_graph.graph import CodeGraph

graph = CodeGraph()

# Standard extraction from parsed source

graph.add_edge(
    source_qualified="auth.login",
    target_qualified="db.query_user",
    line=42,
    extra={},                          # Optional metadata

    confidence=1.0,                    # Full confidence

    confidence_tier="EXTRACTED"        # Default: directly observed

)

```

### Inferred Edge After Resolution

```python

# After scoped resolution finds a cross-module target

graph.add_edge(
    source_qualified="api.router",
    target_qualified="internal.handler_v2",  # Rewritten from aliased reference

    line=10,
    extra={"resolution_pass": "scoped"},
    confidence=0.95,                   # Reduced confidence

    confidence_tier="INFERRED"         # Marked as derived

)

```

---

## Verification in Test Suite

The distinction between tiers is **enforced by tests** across multiple language frontends:

- [`tests/test_rust_scoped_calls.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_rust_scoped_calls.py): Asserts that resolver-rewritten Rust call edges have `confidence_tier == "INFERRED"`
- [`tests/test_php_scoped_calls.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_php_scoped_calls.py): Verifies PHP edges directly from the parser retain `confidence_tier == "EXTRACTED"`

These tests ensure that the edge confidence scoring correctly categorizes edge types regardless of the source language being analyzed.

---

## Key Files for Implementation Details

| File | Role in Confidence Scoring |
|------|---------------------------|
| [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | Defines `Edge` dataclass with `confidence` and `confidence_tier` fields; handles database serialization |
| [`code_review_graph/scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/scoped_resolver.py) | Marks rewritten edges with `confidence_tier = 'INFERRED'` |
| [`code_review_graph/event_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/event_resolver.py) | Creates inferred event edges with reduced confidence scores |
| [`code_review_graph/migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/migrations.py) | Database migration v9 adding confidence columns to `edges` table |
| [`docs/schema.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/schema.md) | Schema documentation for `confidence_tier` column |
| [`tests/test_rust_scoped_calls.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_rust_scoped_calls.py) | Validates tier assignment for Rust scoped resolution |

---

## Summary

- Edge confidence scoring uses **two fields**: numeric `confidence` (0–1) and categorical `confidence_tier` (`EXTRACTED` | `INFERRED`)
- **EXTRACTED** edges come directly from static parsing and carry full reliability
- **INFERRED** edges are resolver-generated and marked with reduced confidence
- The **[`scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/scoped_resolver.py)** and **[`event_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/event_resolver.py)** modules explicitly set `confidence_tier = "INFERRED"` for rewritten or deduced edges
- The design enables downstream consumers to filter, weight, or visualize edges based on **both origin and certainty**

---

## Frequently Asked Questions

### What is the default confidence_tier for new edges?

**EXTRACTED**. When calling `graph.add_edge()` without specifying `confidence_tier`, the system defaults to `"EXTRACTED"` with `confidence=1.0`. This reflects the common case of edges directly observed by the static parser.

### When should I use confidence_tier=INFERRED?

Use `INFERRED` when your code **generates edges through analysis rather than direct parsing**—for example, when resolving scoped references, inferring event handlers, or deducing implicit dependencies. The code-review-graph resolvers automatically apply this tier during such operations.

### How does the numeric confidence field interact with confidence_tier?

The `confidence` field provides **granular weighting** within a tier, while `confidence_tier` indicates **provenance**. An inferred edge might have `confidence=0.95` (high certainty inference) or `confidence=0.6` (speculative deduction). Both fields can be queried together: `WHERE confidence_tier = 'INFERRED' AND confidence > 0.9`.

### Where are confidence values stored and queried?

Confidence data persists in the **SQLite `edges` table**, with columns added by migration v9 in [`code_review_graph/migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/migrations.py). The `CodeGraph` class in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) handles serialization, enabling SQL queries like `SELECT * FROM edges WHERE confidence_tier = 'INFERRED'` for downstream filtering.