How Edge Confidence Scoring Categorizes Edge Types in Code-Review-Graph
The edge confidence scoring system in code-review-graph categorizes edges into two confidence tiers: EXTRACTED for edges directly parsed from source code, and INFERRED for edges generated by resolver analysis, with each tier paired with a numeric confidence score from 0 to 1.
Understanding how edge confidence scoring categorizes edge types is essential when working with the code-review-graph repository, a tool that builds dependency graphs from static code analysis. The system uses a dual-layer confidence model that separates how an edge originated from how certain the system is about that relationship. This design lets downstream tools filter or weight edges differently based on their provenance.
The Two-Field Confidence Model
Every edge in the graph stores two distinct confidence values in the SQLite database, as defined in code_review_graph/graph.py and documented in docs/schema.md:
| Field | Type | Default | Purpose |
|---|---|---|---|
confidence |
float |
1.0 |
Numeric certainty (0–1) for fine-grained weighting |
confidence_tier |
str |
"EXTRACTED" |
Categorical origin classification |
These fields were added to the edges table in migration v9 (code_review_graph/migrations.py), enabling the graph to distinguish between ground-truth extractions and analytical inferences.
The Two Confidence Tiers
The confidence_tier field implements exactly two edge type categories:
EXTRACTED Tier
Edges with confidence_tier: "EXTRACTED" are taken verbatim from the source code by the static parser. These represent direct observations—function calls, imports, or references that appear literally in the analyzed code.
- Set automatically as the default in
graph.add_edge() - Considered fully reliable (
confidencetypically1.0) - No resolver mediation involved
INFERRED Tier
Edges with confidence_tier: "INFERRED" are generated by resolver passes that rewrite or deduce relationships based on higher-level analysis. These edges originate from computational inference rather than direct parsing.
- Explicitly set by resolver code in
code_review_graph/scoped_resolver.pyandcode_review_graph/event_resolver.py - Typically paired with reduced numeric confidence (e.g.,
0.95) - May represent scoped resolutions or event-based relationships not visible in raw code
How Resolvers Assign the INFERRED Tier
The categorization logic lives in two resolver implementations:
Scoped Resolver (code_review_graph/scoped_resolver.py)
When the scoped resolver rewrites edges—for example, resolving a call to its target across module boundaries—it explicitly tags them as inferred. The source code comment states: "Rewritten edges are tagged confidence_tier = INFERRED". The accompanying SQL update statements execute:
# From scoped_resolver.py - edge rewrite operation
UPDATE edges
SET confidence_tier = 'INFERRED',
confidence = :confidence
WHERE edge_id = :edge_id
Event Resolver (code_review_graph/event_resolver.py)
For event-based analysis, inferred edges are created directly with both reduced confidence and the inferred tier:
# Creating an inferred event edge in event_resolver.py
graph.add_edge(
source_qualified="event_handler.on_click",
target_qualified="event_handler.process",
line=15,
extra={"event_type": "click"},
confidence=0.95, # Slightly reduced certainty
confidence_tier="INFERRED" # Marked as analytically derived
)
Practical Code Examples
Default Extracted Edge
from code_review_graph.graph import CodeGraph
graph = CodeGraph()
# Standard extraction from parsed source
graph.add_edge(
source_qualified="auth.login",
target_qualified="db.query_user",
line=42,
extra={}, # Optional metadata
confidence=1.0, # Full confidence
confidence_tier="EXTRACTED" # Default: directly observed
)
Inferred Edge After Resolution
# After scoped resolution finds a cross-module target
graph.add_edge(
source_qualified="api.router",
target_qualified="internal.handler_v2", # Rewritten from aliased reference
line=10,
extra={"resolution_pass": "scoped"},
confidence=0.95, # Reduced confidence
confidence_tier="INFERRED" # Marked as derived
)
Verification in Test Suite
The distinction between tiers is enforced by tests across multiple language frontends:
tests/test_rust_scoped_calls.py: Asserts that resolver-rewritten Rust call edges haveconfidence_tier == "INFERRED"tests/test_php_scoped_calls.py: Verifies PHP edges directly from the parser retainconfidence_tier == "EXTRACTED"
These tests ensure that the edge confidence scoring correctly categorizes edge types regardless of the source language being analyzed.
Key Files for Implementation Details
| File | Role in Confidence Scoring |
|---|---|
code_review_graph/graph.py |
Defines Edge dataclass with confidence and confidence_tier fields; handles database serialization |
code_review_graph/scoped_resolver.py |
Marks rewritten edges with confidence_tier = 'INFERRED' |
code_review_graph/event_resolver.py |
Creates inferred event edges with reduced confidence scores |
code_review_graph/migrations.py |
Database migration v9 adding confidence columns to edges table |
docs/schema.md |
Schema documentation for confidence_tier column |
tests/test_rust_scoped_calls.py |
Validates tier assignment for Rust scoped resolution |
Summary
- Edge confidence scoring uses two fields: numeric
confidence(0–1) and categoricalconfidence_tier(EXTRACTED|INFERRED) - EXTRACTED edges come directly from static parsing and carry full reliability
- INFERRED edges are resolver-generated and marked with reduced confidence
- The
scoped_resolver.pyandevent_resolver.pymodules explicitly setconfidence_tier = "INFERRED"for rewritten or deduced edges - The design enables downstream consumers to filter, weight, or visualize edges based on both origin and certainty
Frequently Asked Questions
What is the default confidence_tier for new edges?
EXTRACTED. When calling graph.add_edge() without specifying confidence_tier, the system defaults to "EXTRACTED" with confidence=1.0. This reflects the common case of edges directly observed by the static parser.
When should I use confidence_tier=INFERRED?
Use INFERRED when your code generates edges through analysis rather than direct parsing—for example, when resolving scoped references, inferring event handlers, or deducing implicit dependencies. The code-review-graph resolvers automatically apply this tier during such operations.
How does the numeric confidence field interact with confidence_tier?
The confidence field provides granular weighting within a tier, while confidence_tier indicates provenance. An inferred edge might have confidence=0.95 (high certainty inference) or confidence=0.6 (speculative deduction). Both fields can be queried together: WHERE confidence_tier = 'INFERRED' AND confidence > 0.9.
Where are confidence values stored and queried?
Confidence data persists in the SQLite edges table, with columns added by migration v9 in code_review_graph/migrations.py. The CodeGraph class in graph.py handles serialization, enabling SQL queries like SELECT * FROM edges WHERE confidence_tier = 'INFERRED' for downstream filtering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →