What Do Edge Confidence Scores (EXTRACTED/INFERRED/AMBIGUOUS) Mean in code-review-graph?
Edge confidence scores in code-review-graph indicate how a relationship between two code symbols was derived: EXTRACTED means directly parsed from source, INFERRED means heuristically resolved by the scoped resolver, and AMBIGUOUS means the target could not be uniquely determined.
The code-review-graph project stores code relationships—calls, imports, inheritance, and tests—as rows in an SQLite database. Each edge carries a confidence_tier column that tells downstream tools exactly how much trust they should place in that relationship. Understanding these tiers is essential for building accurate static analysis pipelines, risk-scored CI reviews, and developer debugging workflows.
The Three Confidence Tiers Explained
EXTRACTED: Directly Parsed from Source
EXTRACTED edges represent relationships that were directly extracted from the tree-sitter AST during the initial parse with no post-processing required.
The database schema in code_review_graph/graph.py defines this as the default:
# From code_review_graph/graph.py (lines 103-104)
edges.confidence_tier TEXT DEFAULT 'EXTRACTED'
When GraphStore.upsert_edge writes a new edge, it uses this default unless explicitly overridden. These edges carry the highest trust because callers, imports, and inheritance relationships are known with full source information.
INFERRED: Heuristically Resolved
INFERRED edges were rewritten after parsing through the scoped resolver. This typically happens when raw AST nodes contain partially qualified names that must be mapped to canonical fully-qualified symbols.
In code_review_graph/scoped_resolver.py, a comment explains the mechanism:
"Rewritten edges are tagged
confidence_tier = INFERRED…"
After successfully resolving a scoped call like PHP's Mailer::send to the canonical src/mail.php::Mailer.send, the resolver updates the edge's confidence_tier to 'INFERRED'.
These edges remain usable but signal to downstream tools that heuristic resolution occurred. The test suite validates this behavior explicitly:
# From tests/test_rust_scoped_calls.py (line 89)
assert resolved[0]["confidence_tier"] == "INFERRED"
# From tests/test_php_scoped_calls.py (line 113)
assert resolved[0]["confidence_tier"] == "INFERRED"
AMBIGUOUS: Unresolvable Target
AMBIGUOUS edges indicate that multiple candidate definitions exist or import information was insufficient to determine a unique target. The original raw target is preserved, but the edge is flagged for special handling.
The core library does not yet write this tier automatically during normal operation. Instead, tests/test_forget_parity.py sets up an _AMBIGUOUS_IMPORT_FILES scenario (line 36) to verify that the system handles ambiguous edges gracefully. Consumers such as blast-radius analysis tools may choose to:
- Omit ambiguous edges from precise impact calculations
- Flag them for manual developer review
- Apply confidence-weighted scoring that reduces their influence
How Confidence Tiers Are Used in Practice
Risk Scoring in CI Integration
According to the documentation in docs/README.md, confidence tiers directly affect risk-scored PR reviews. Higher-confidence edges contribute more significantly to risk calculations, while INFERRED and AMBIGUOUS edges may be weighted down or filtered out.
Incremental Graph Updates
The resolver's incremental update strategy leverages these tiers:
- Only edges eligible for inference rewriting are processed
AMBIGUOUSedges remain untouched to avoid propagating incorrect rewritesEXTRACTEDedges serve as ground truth anchors
Querying and Filtering Edges
Use these Python snippets to work with confidence tiers directly:
# Query all edges with their confidence tiers
from code_review_graph.graph import GraphStore
store = GraphStore("myrepo/.code-review-graph/graph.db")
with store:
rows = store._conn.execute(
"SELECT source_qualified, target_qualified, confidence_tier FROM edges"
).fetchall()
for r in rows:
print(f"{r['source_qualified']} → {r['target_qualified']} [{r['confidence_tier']}]")
Example output:
src/main.py::main → src/util.py::helper [EXTRACTED]
src/mail.php::Mailer::send → src/mail.php::Mailer.send [INFERRED]
src/ambiguous.rs::foo → src/ambiguous.rs::Bar [AMBIGUOUS]
# Filter for high-confidence edges only
high_conf = store._conn.execute(
"SELECT * FROM edges WHERE confidence_tier = 'EXTRACTED'"
).fetchall()
# Update an edge after successful inference
store._conn.execute(
"""
UPDATE edges
SET target_qualified = ?, confidence_tier = 'INFERRED'
WHERE id = ?
""",
("src/mail.php::Mailer.send", edge_id),
)
store._invalidate_cache()
Key Source Files for Edge Confidence
| File | Purpose |
|---|---|
code_review_graph/graph.py |
Defines edges table schema with default confidence_tier = 'EXTRACTED' and the GraphEdge dataclass |
code_review_graph/migrations.py |
Adds the confidence_tier column during database migrations (lines 234-236) |
code_review_graph/scoped_resolver.py |
Explains INFERRED tier assignment for rewritten edges (line 44) |
docs/schema.md |
Documents the database schema including column defaults (lines 207-208) |
tests/test_rust_scoped_calls.py |
Asserts INFERRED confidence for resolved Rust scoped calls (line 89) |
tests/test_php_scoped_calls.py |
Asserts INFERRED confidence for resolved PHP scoped calls (line 113) |
tests/test_forget_parity.py |
Sets up ambiguous import test scenario (line 36) |
Summary
- EXTRACTED edges come straight from tree-sitter parsing with maximum reliability
- INFERRED edges were heuristically resolved by the scoped resolver and remain trustworthy but flagged
- AMBIGUOUS edges represent unresolved targets that consumers should handle carefully
- The
confidence_tiercolumn defaults to'EXTRACTED'incode_review_graph/graph.pyand propagates throughGraphStore.upsert_edge - Downstream tools use these tiers for risk scoring, filtering, and incremental updates
Frequently Asked Questions
How do I filter out low-confidence edges from my analysis?
Query with WHERE confidence_tier = 'EXTRACTED' to include only directly parsed relationships. For analyses that can tolerate heuristic resolution, include both 'EXTRACTED' and 'INFERRED'. Omit or specially handle 'AMBIGUOUS' depending on your precision requirements.
Can I manually change an edge's confidence tier?
Yes. Use standard SQL UPDATE statements against the edges table, then call store._invalidate_cache() to ensure the GraphStore picks up changes. The code example in the querying section demonstrates this pattern.
Why does INFERRED exist instead of just updating the edge in place?
The INFERRED tier preserves provenance. It allows debugging tools to trace whether a relationship came directly from source parsing or required heuristic resolution. This audit trail is valuable for debugging resolver behavior and understanding why particular edges appear in query results.
Does the system automatically create AMBIGUOUS edges during normal operation?
Not currently. The AMBIGUOUS tier is primarily used in test scenarios (tests/test_forget_parity.py) to verify graceful handling. Future versions may automatically mark edges as ambiguous when the resolver encounters multiple valid targets, but this requires explicit implementation in the resolution pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →