Edge Confidence Levels in Code-Review-Graph: EXTRACTED, INFERRED, and AMBIGUOUS Explained
Code-Review-Graph uses three symbolic confidence tiers—EXTRACTED, INFERRED, and AMBIGUOUS—to categorize how certain each relationship edge is, combined with a numeric 0–1 score for fine-grained certainty.
Code-Review-Graph models relationships between symbols (functions, classes, modules) as edges in an SQLite database. Every edge carries both a numeric confidence score and a confidence_tier label that tells downstream tools how the relationship was derived. Understanding these tiers helps you interpret query results and build reliable refactoring or impact analysis tools on top of the graph.
How Edge Confidence Works
Each edge stores two related confidence fields:
| Field | Type | Purpose | Default |
|---|---|---|---|
confidence |
float (0–1) | Numeric certainty score | 1.0 |
confidence_tier |
text | Symbolic source-of-derivation category | EXTRACTED |
The confidence_tier groups edges by their origin, while the confidence float provides granular scoring within each tier.
The Three Confidence Tiers
EXTRACTED: Directly Parsed from Source
EXTRACTED edges represent concrete syntax relationships found directly in source files. These have the highest certainty because they correspond to explicit code constructs—actual function calls, imports, or inheritance declarations visible in the text.
In code_review_graph/migrations.py, the edges table schema sets confidence_tier with a default value:
# migrations.py lines 229-238
cursor.execute("""
CREATE TABLE edges (
id INTEGER PRIMARY KEY AUTOINCREMENT,
source_qualified TEXT NOT NULL,
target_qualified TEXT NOT NULL,
line INTEGER NOT NULL,
extra TEXT NOT NULL DEFAULT '',
confidence REAL NOT NULL DEFAULT 1.0,
confidence_tier TEXT NOT NULL DEFAULT 'EXTRACTED', -- highest certainty default
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
)
""")
When code_review_graph/graph.py inserts edges without specifying a tier, SQLite automatically assigns EXTRACTED:
# graph.py lines 102-106
def add_edge(
self,
source_qualified: str,
target_qualified: str,
line: int,
extra: str = "",
confidence: float = 1.0,
confidence_tier: str = "EXTRACTED" # explicit default matching schema
) -> int:
INFERRED: Derived Through Static Analysis
INFERRED edges come from analysis rather than direct parsing. The system generates these through Language Server Protocol (LSP) resolution, type inference, or scoped resolution that discovers relationships not explicitly written in source.
In code_review_graph/scoped_resolver.py, rewritten edges during type-aware resolution are tagged as INFERRED:
# scoped_resolver.py lines 44-49
def resolve_scoped_call(
self,
call_node: Call,
scope: Scope
) -> Optional[Edge]:
# ... resolution logic ...
return Edge(
source_qualified=scope.qualified_name,
target_qualified=resolved_target,
line=call_node.lineno,
confidence=0.95, # high but not certain
confidence_tier="INFERRED" # analysis-derived, not surface syntax
)
Common sources of INFERRED edges include:
- Method override relationships discovered via LSP symbol resolution
- Dynamic dispatch targets resolved through type inference
- Call targets rewritten after scoped name resolution
These edges remain reliable but carry slightly lower nominal certainty because they depend on static analysis heuristics and may change if analysis rules improve.
AMBIGUOUS: Multiple Possible Targets
AMBIGUOUS edges indicate the system cannot pinpoint a single target. This occurs when:
- Multiple definitions match a reference (e.g., ambiguous imports or shadowed names)
- Overload resolution cannot determine the precise callee
- Type information is insufficient to disambiguate polymorphic calls
The AMBIGUOUS tier signals to downstream consumers that the relationship should be treated cautiously. As noted in docs/FAQ.md (line 48), this tier explicitly marks edges where "the analysis cannot decide between multiple valid targets."
Practical Code Examples
Inserting an EXTRACTED Edge (Default)
from code_review_graph.graph import GraphStore
store = GraphStore("reviews.db")
# Direct function call found in source parsing
store.add_edge(
source_qualified="billing.invoice.process",
target_qualified="payment.gateway.charge",
line=42,
extra="",
confidence=1.0,
confidence_tier="EXTRACTED"
)
Inserting an INFERRED Edge from LSP Analysis
# Edge discovered via Language Server Protocol resolution
store.add_edge(
source_qualified="reports.generator.Renderer.render",
target_qualified="reports.backend.PDFBackend.write",
line=156,
extra="",
confidence=0.92,
confidence_tier="INFERRED"
)
Inserting an AMBIGUOUS Edge
# Multiple possible targets for overloaded or ambiguous call
store.add_edge(
source_qualified="utils.format.display",
target_qualified="utils.format.HTMLFormatter|utils.format.TextFormatter",
line=89,
extra="unresolved_overload",
confidence=0.65,
confidence_tier="AMBIGUOUS"
)
Querying Edges by Confidence Tier
Filter your graph queries to exclude uncertain edges for critical path analysis:
# High-certainty edges only
high_confidence_edges = store.execute("""
SELECT * FROM edges
WHERE confidence_tier = 'EXTRACTED'
OR (confidence_tier = 'INFERRED' AND confidence >= 0.90)
""")
# Review ambiguous relationships manually
ambiguous = store.execute("""
SELECT * FROM edges
WHERE confidence_tier = 'AMBIGUOUS'
ORDER BY confidence ASC
""")
Key Implementation Files
| File | Purpose |
|---|---|
code_review_graph/graph.py |
Edge dataclass, default confidence values, CRUD operations |
code_review_graph/migrations.py |
Schema migration adding confidence and confidence_tier columns (v9) |
code_review_graph/scoped_resolver.py |
Tags rewritten edges as INFERRED during resolution |
docs/FAQ.md |
User-facing documentation of the three-tier system |
Summary
- EXTRACTED edges come from direct source parsing with maximum certainty (default tier)
- INFERRED edges derive from static analysis and LSP resolution, slightly lower certainty
- AMBIGUOUS edges mark unresolved multiple targets requiring manual review
- The numeric
confidencescore (0–1) provides granularity within each tier - Schema defaults in
migrations.pyand insertion logic ingraph.pyensureEXTRACTEDis the safe baseline
Frequently Asked Questions
How do I change the confidence tier for an existing edge?
Update the edges table directly with SQL. Use parameterized queries to avoid injection: UPDATE edges SET confidence_tier = 'INFERRED', confidence = 0.85 WHERE id = ?. The GraphStore class in graph.py does not expose a dedicated update_edge_tier() method, so direct SQL is currently required.
Can I disable inferred edges entirely?
Yes. When building queries, filter with WHERE confidence_tier = 'EXTRACTED'. This excludes all analysis-derived relationships and uses only surface syntax edges. Note that this may miss legitimate call relationships in dynamic or polymorphic code.
What confidence score should I assign to ambiguous edges?
The codebase uses 0.70 as a representative value, but any score below 0.80 signals notable uncertainty. Consider the severity: unresolved imports might warrant 0.60, while minor overload ambiguity could use 0.75. Document your scoring rationale for downstream consumers.
Where is the confidence tier system documented for end users?
The three-tier system is explained in docs/FAQ.md at line 48, with additional context in the project README.md. Developers should also consult migrations.py for schema defaults and scoped_resolver.py to understand when INFERRED edges are generated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →