# Edge Confidence Levels in Code-Review-Graph: EXTRACTED, INFERRED, and AMBIGUOUS Explained

> Understand edge confidence levels EXTRACTED, INFERRED, and AMBIGUOUS in Code-Review-Graph. Learn how these tiers and numeric scores reveal relationship certainty in code reviews.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: deep-dive
- Published: 2026-08-16

---

**Code-Review-Graph uses three symbolic confidence tiers—EXTRACTED, INFERRED, and AMBIGUOUS—to categorize how certain each relationship edge is, combined with a numeric 0–1 score for fine-grained certainty.**

Code-Review-Graph models relationships between symbols (functions, classes, modules) as edges in an SQLite database. Every edge carries both a numeric `confidence` score and a `confidence_tier` label that tells downstream tools how the relationship was derived. Understanding these tiers helps you interpret query results and build reliable refactoring or impact analysis tools on top of the graph.

## How Edge Confidence Works

Each edge stores two related confidence fields:

| Field | Type | Purpose | Default |
|-------|------|---------|---------|
| `confidence` | float (0–1) | Numeric certainty score | 1.0 |
| `confidence_tier` | text | Symbolic source-of-derivation category | `EXTRACTED` |

The `confidence_tier` groups edges by their origin, while the `confidence` float provides granular scoring within each tier.

## The Three Confidence Tiers

### EXTRACTED: Directly Parsed from Source

**`EXTRACTED`** edges represent concrete syntax relationships found directly in source files. These have the highest certainty because they correspond to explicit code constructs—actual function calls, imports, or inheritance declarations visible in the text.

In [`code_review_graph/migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/migrations.py), the `edges` table schema sets `confidence_tier` with a default value:

```python

# migrations.py lines 229-238

cursor.execute("""
    CREATE TABLE edges (
        id INTEGER PRIMARY KEY AUTOINCREMENT,
        source_qualified TEXT NOT NULL,
        target_qualified TEXT NOT NULL,
        line INTEGER NOT NULL,
        extra TEXT NOT NULL DEFAULT '',
        confidence REAL NOT NULL DEFAULT 1.0,
        confidence_tier TEXT NOT NULL DEFAULT 'EXTRACTED',  -- highest certainty default
        created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
    )
""")

```

When [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) inserts edges without specifying a tier, SQLite automatically assigns `EXTRACTED`:

```python

# graph.py lines 102-106

def add_edge(
    self,
    source_qualified: str,
    target_qualified: str,
    line: int,
    extra: str = "",
    confidence: float = 1.0,
    confidence_tier: str = "EXTRACTED"  # explicit default matching schema

) -> int:

```

### INFERRED: Derived Through Static Analysis

**`INFERRED`** edges come from analysis rather than direct parsing. The system generates these through **Language Server Protocol (LSP) resolution**, **type inference**, or **scoped resolution** that discovers relationships not explicitly written in source.

In [`code_review_graph/scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/scoped_resolver.py), rewritten edges during type-aware resolution are tagged as `INFERRED`:

```python

# scoped_resolver.py lines 44-49

def resolve_scoped_call(
    self,
    call_node: Call,
    scope: Scope
) -> Optional[Edge]:
    # ... resolution logic ...

    return Edge(
        source_qualified=scope.qualified_name,
        target_qualified=resolved_target,
        line=call_node.lineno,
        confidence=0.95,  # high but not certain

        confidence_tier="INFERRED"  # analysis-derived, not surface syntax

    )

```

Common sources of `INFERRED` edges include:

- Method override relationships discovered via LSP symbol resolution
- Dynamic dispatch targets resolved through type inference
- Call targets rewritten after scoped name resolution

These edges remain reliable but carry slightly lower nominal certainty because they depend on static analysis heuristics and may change if analysis rules improve.

### AMBIGUOUS: Multiple Possible Targets

**`AMBIGUOUS`** edges indicate the system cannot pinpoint a single target. This occurs when:

- Multiple definitions match a reference (e.g., ambiguous imports or shadowed names)
- Overload resolution cannot determine the precise callee
- Type information is insufficient to disambiguate polymorphic calls

The `AMBIGUOUS` tier signals to downstream consumers that the relationship should be treated cautiously. As noted in [`docs/FAQ.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/FAQ.md) (line 48), this tier explicitly marks edges where "the analysis cannot decide between multiple valid targets."

## Practical Code Examples

### Inserting an EXTRACTED Edge (Default)

```python
from code_review_graph.graph import GraphStore

store = GraphStore("reviews.db")

# Direct function call found in source parsing

store.add_edge(
    source_qualified="billing.invoice.process",
    target_qualified="payment.gateway.charge",
    line=42,
    extra="",
    confidence=1.0,
    confidence_tier="EXTRACTED"
)

```

### Inserting an INFERRED Edge from LSP Analysis

```python

# Edge discovered via Language Server Protocol resolution

store.add_edge(
    source_qualified="reports.generator.Renderer.render",
    target_qualified="reports.backend.PDFBackend.write",
    line=156,
    extra="",
    confidence=0.92,
    confidence_tier="INFERRED"
)

```

### Inserting an AMBIGUOUS Edge

```python

# Multiple possible targets for overloaded or ambiguous call

store.add_edge(
    source_qualified="utils.format.display",
    target_qualified="utils.format.HTMLFormatter|utils.format.TextFormatter",
    line=89,
    extra="unresolved_overload",
    confidence=0.65,
    confidence_tier="AMBIGUOUS"
)

```

## Querying Edges by Confidence Tier

Filter your graph queries to exclude uncertain edges for critical path analysis:

```python

# High-certainty edges only

high_confidence_edges = store.execute("""
    SELECT * FROM edges 
    WHERE confidence_tier = 'EXTRACTED' 
       OR (confidence_tier = 'INFERRED' AND confidence >= 0.90)
""")

# Review ambiguous relationships manually

ambiguous = store.execute("""
    SELECT * FROM edges 
    WHERE confidence_tier = 'AMBIGUOUS'
    ORDER BY confidence ASC
""")

```

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`code_review_graph/graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | `Edge` dataclass, default confidence values, CRUD operations |
| [`code_review_graph/migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/migrations.py) | Schema migration adding `confidence` and `confidence_tier` columns (v9) |
| [`code_review_graph/scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/scoped_resolver.py) | Tags rewritten edges as `INFERRED` during resolution |
| [`docs/FAQ.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/FAQ.md) | User-facing documentation of the three-tier system |

## Summary

- **EXTRACTED** edges come from direct source parsing with maximum certainty (default tier)
- **INFERRED** edges derive from static analysis and LSP resolution, slightly lower certainty
- **AMBIGUOUS** edges mark unresolved multiple targets requiring manual review
- The numeric `confidence` score (0–1) provides granularity within each tier
- Schema defaults in [`migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/migrations.py) and insertion logic in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) ensure `EXTRACTED` is the safe baseline

## Frequently Asked Questions

### How do I change the confidence tier for an existing edge?

Update the `edges` table directly with SQL. Use parameterized queries to avoid injection: `UPDATE edges SET confidence_tier = 'INFERRED', confidence = 0.85 WHERE id = ?`. The `GraphStore` class in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) does not expose a dedicated `update_edge_tier()` method, so direct SQL is currently required.

### Can I disable inferred edges entirely?

Yes. When building queries, filter with `WHERE confidence_tier = 'EXTRACTED'`. This excludes all analysis-derived relationships and uses only surface syntax edges. Note that this may miss legitimate call relationships in dynamic or polymorphic code.

### What confidence score should I assign to ambiguous edges?

The codebase uses 0.70 as a representative value, but any score below 0.80 signals notable uncertainty. Consider the severity: unresolved imports might warrant 0.60, while minor overload ambiguity could use 0.75. Document your scoring rationale for downstream consumers.

### Where is the confidence tier system documented for end users?

The three-tier system is explained in [`docs/FAQ.md`](https://github.com/tirth8205/code-review-graph/blob/main/docs/FAQ.md) at line 48, with additional context in the project [`README.md`](https://github.com/tirth8205/code-review-graph/blob/main/README.md). Developers should also consult [`migrations.py`](https://github.com/tirth8205/code-review-graph/blob/main/migrations.py) for schema defaults and [`scoped_resolver.py`](https://github.com/tirth8205/code-review-graph/blob/main/scoped_resolver.py) to understand when `INFERRED` edges are generated.