Understanding Confidence Tags in Graphify: EXTRACTED vs INFERRED Explained
Graphify annotates every edge in its knowledge graph with a confidence tag—either EXTRACTED for relationships parsed directly from source code or INFERRED for relationships derived by the resolution engine—allowing developers to distinguish explicit facts from heuristic suggestions.
The Graphify-Labs/graphify repository transforms codebases into queryable knowledge graphs, tagging each edge with provenance metadata that reveals how the relationship was discovered. These confidence tags appear in the generated JSON output and CLI results, giving you precise control over whether you trust only explicit code links or want to explore speculative connections. Understanding these tags is critical for tuning the precision of automated code analysis and refactoring workflows.
What Are the Confidence Tags in Graphify?
Graphify defines exactly two confidence tags that label every edge in the graph:
EXTRACTED– The edge represents a relationship that was explicitly parsed from source code or other artifacts. This includes direct function calls, import statements, class inheritance, and other syntactic structures that the parser can read verbatim.INFERRED– The edge was derived by the resolution engine using heuristics, name-based similarity, or cross-language matching. These connections are not explicit in the source but represent plausible links that Graphify suggests based on its analysis algorithms.
According to the repository documentation in README.md, this distinction ensures that users can "tell what was read directly from what was inferred" when querying the graph. This provenance tracking prevents false positives from being treated as ground truth when executing automated refactorings or security audits.
Where Confidence Tags Appear in the Codebase
The confidence tag system is implemented across several key files in the repository:
README.md
The primary documentation defines the semantic meaning of both tags, explaining that EXTRACTED edges carry the highest certainty because they mirror actual code structures, while INFERRED edges carry the resolution engine's confidence scores derived from pattern matching.
worked/rsl-siege-manager/graph.json
This sample output demonstrates the concrete JSON structure where confidence tags appear. Each edge object includes a "confidence" field set to "EXTRACTED" and often a "confidence_score" floating-point value (e.g., 1.0 for fully certain extracted edges).
graphify/semantics/relationships.py
This module defines the internal edge data structures that store confidence metadata. When the parser traverses the codebase, it instantiates relationship objects with the EXTRACTED tag, while the resolution engine later augments the graph with INFERRED edges based on symbol matching heuristics.
graphify/cli.py
The command-line interface reads these tags during query execution and exposes them in formatted output, allowing users to see confidence labels when investigating code relationships.
How to Query and Filter by Confidence Tags
You can inspect confidence tags programmatically or via the CLI to build trust-level filters into your analysis pipeline.
Using the Graphify CLI
Query the knowledge graph and pipe the results to filter for only extracted relationships:
# Show all edges with their confidence tags
graphify query "How does a Discord member get synced?" --output json | \
jq '.edges[] | {source: .source, target: .target, confidence: .confidence}'
This returns JSON objects containing the "confidence" field, which you can filter further to exclude INFERRED edges when requiring high certainty.
Programmatic Access with Node.js
If consuming Graphify output in a TypeScript or JavaScript application, you can inspect the confidence property of edge objects:
import { Graphify } from "graphify";
const g = new Graphify();
const subgraph = await g.query("What is the confidence of a Discord sync?");
// Find a specific edge and check its confidence tag
const edge = subgraph.edges.find(e => e.id === "components_discordsyncmodal_confidence_label");
if (edge && edge.confidence === "EXTRACTED") {
console.log("High certainty relationship found");
} else {
console.log("Heuristic relationship - verify manually");
}
This pattern allows you to gate automated refactoring tools behind confidence checks, ensuring that only EXTRACTED edges trigger breaking changes while INFERRED edges queue for manual review.
Why Confidence Tags Matter for Code Analysis
The dual-tag system solves the "black box" problem in automated code intelligence:
Trust Calibration
When migrating legacy code, you can restrict automated transformations to EXTRACTED edges only, guaranteeing that the tool modifies only relationships that actually exist in the source. This prevents the resolution engine's speculative matches from introducing errors into critical refactoring operations.
False Positive Filtering
Cross-language graphs often generate INFERRED edges through name-based heuristics that may produce false positives. By surfacing these as distinct from EXTRACTED facts, Graphify lets you set thresholds or manually review low-confidence connections before they influence architectural decisions.
Audit Trail
Security and compliance audits require provenance tracking. The confidence tags in worked/rsl-siege-manager/graph.json and similar outputs provide an immutable record of whether a reported dependency was parsed from an import statement or suggested by heuristic analysis.
Summary
- Confidence tags in Graphify label every graph edge as either
EXTRACTED(parsed directly from code) orINFERRED(computed by the resolution engine). - The distinction is documented in
README.mdand implemented ingraphify/semantics/relationships.py, with concrete examples visible inworked/rsl-siege-manager/graph.json. - Use the
graphifyCLI with JSON output filters to separate high-certainty (EXTRACTED) relationships from heuristic suggestions. - These tags prevent false positives in automated refactoring by ensuring that only explicit code relationships trigger changes without manual review.
Frequently Asked Questions
What is the difference between EXTRACTED and INFERRED confidence tags in Graphify?
The EXTRACTED tag indicates that Graphify found the relationship by parsing source code directly, such as reading an import statement or function call. The INFERRED tag means the resolution engine generated the relationship using heuristics like name matching or cross-language similarity, without explicit syntactic evidence in the source files.
How can I filter query results to show only EXTRACTED edges?
Pipe the JSON output from graphify query to a filter like jq and select edges where the "confidence" field equals "EXTRACTED". Alternatively, when building applications with the Graphify client library, check the edge.confidence property programmatically before processing relationships in your business logic.
Where are confidence tags stored in the Graphify output format?
Confidence tags appear as string values in the "confidence" field of edge objects within the generated JSON graph files, as demonstrated in worked/rsl-siege-manager/graph.json. Each edge may also include a numeric "confidence_score" field that provides additional granularity for INFERRED relationships.
Can I adjust the thresholds for INFERRED edges in Graphify?
While the resolution engine produces INFERRED edges based on internal heuristics defined in modules like graphify/semantics/relationships.py, you can filter these edges post-generation by examining the confidence_score values in the JSON output or implementing custom logic in your consuming application to ignore edges below a specific confidence threshold.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →