# Understanding Confidence Tags in Graphify: EXTRACTED vs INFERRED Explained

> Learn about Graphify confidence tags: EXTRACTED vs INFERRED. Understand how Graphify distinguishes explicit facts from heuristic suggestions in knowledge graphs.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: deep-dive
- Published: 2026-07-16

---

**Graphify annotates every edge in its knowledge graph with a confidence tag—either `EXTRACTED` for relationships parsed directly from source code or `INFERRED` for relationships derived by the resolution engine—allowing developers to distinguish explicit facts from heuristic suggestions.**

The **Graphify-Labs/graphify** repository transforms codebases into queryable knowledge graphs, tagging each edge with provenance metadata that reveals how the relationship was discovered. These **confidence tags** appear in the generated JSON output and CLI results, giving you precise control over whether you trust only explicit code links or want to explore speculative connections. Understanding these tags is critical for tuning the precision of automated code analysis and refactoring workflows.

## What Are the Confidence Tags in Graphify?

Graphify defines exactly two confidence tags that label every edge in the graph:

- **`EXTRACTED`** – The edge represents a relationship that was **explicitly parsed** from source code or other artifacts. This includes direct function calls, import statements, class inheritance, and other syntactic structures that the parser can read verbatim.
- **`INFERRED`** – The edge was **derived by the resolution engine** using heuristics, name-based similarity, or cross-language matching. These connections are not explicit in the source but represent plausible links that Graphify suggests based on its analysis algorithms.

According to the repository documentation in [`README.md`](https://github.com/Graphify-Labs/graphify/blob/main/README.md), this distinction ensures that users can "tell what was read directly from what was inferred" when querying the graph. This provenance tracking prevents false positives from being treated as ground truth when executing automated refactorings or security audits.

## Where Confidence Tags Appear in the Codebase

The confidence tag system is implemented across several key files in the repository:

**[`README.md`](https://github.com/Graphify-Labs/graphify/blob/main/README.md)**  
The primary documentation defines the semantic meaning of both tags, explaining that `EXTRACTED` edges carry the highest certainty because they mirror actual code structures, while `INFERRED` edges carry the resolution engine's confidence scores derived from pattern matching.

**[`worked/rsl-siege-manager/graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/worked/rsl-siege-manager/graph.json)**  
This sample output demonstrates the concrete JSON structure where confidence tags appear. Each edge object includes a `"confidence"` field set to `"EXTRACTED"` and often a `"confidence_score"` floating-point value (e.g., `1.0` for fully certain extracted edges).

**[`graphify/semantics/relationships.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/semantics/relationships.py)**  
This module defines the internal edge data structures that store confidence metadata. When the parser traverses the codebase, it instantiates relationship objects with the `EXTRACTED` tag, while the resolution engine later augments the graph with `INFERRED` edges based on symbol matching heuristics.

**[`graphify/cli.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/cli.py)**  
The command-line interface reads these tags during query execution and exposes them in formatted output, allowing users to see confidence labels when investigating code relationships.

## How to Query and Filter by Confidence Tags

You can inspect confidence tags programmatically or via the CLI to build trust-level filters into your analysis pipeline.

### Using the Graphify CLI

Query the knowledge graph and pipe the results to filter for only extracted relationships:

```bash

# Show all edges with their confidence tags

graphify query "How does a Discord member get synced?" --output json | \
  jq '.edges[] | {source: .source, target: .target, confidence: .confidence}'

```

This returns JSON objects containing the `"confidence"` field, which you can filter further to exclude `INFERRED` edges when requiring high certainty.

### Programmatic Access with Node.js

If consuming Graphify output in a TypeScript or JavaScript application, you can inspect the confidence property of edge objects:

```typescript
import { Graphify } from "graphify";

const g = new Graphify();
const subgraph = await g.query("What is the confidence of a Discord sync?");

// Find a specific edge and check its confidence tag
const edge = subgraph.edges.find(e => e.id === "components_discordsyncmodal_confidence_label");
if (edge && edge.confidence === "EXTRACTED") {
  console.log("High certainty relationship found");
} else {
  console.log("Heuristic relationship - verify manually");
}

```

This pattern allows you to gate automated refactoring tools behind confidence checks, ensuring that only `EXTRACTED` edges trigger breaking changes while `INFERRED` edges queue for manual review.

## Why Confidence Tags Matter for Code Analysis

The dual-tag system solves the "black box" problem in automated code intelligence:

**Trust Calibration**  
When migrating legacy code, you can restrict automated transformations to `EXTRACTED` edges only, guaranteeing that the tool modifies only relationships that actually exist in the source. This prevents the resolution engine's speculative matches from introducing errors into critical refactoring operations.

**False Positive Filtering**  
Cross-language graphs often generate `INFERRED` edges through name-based heuristics that may produce false positives. By surfacing these as distinct from `EXTRACTED` facts, Graphify lets you set thresholds or manually review low-confidence connections before they influence architectural decisions.

**Audit Trail**  
Security and compliance audits require provenance tracking. The confidence tags in [`worked/rsl-siege-manager/graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/worked/rsl-siege-manager/graph.json) and similar outputs provide an immutable record of whether a reported dependency was parsed from an import statement or suggested by heuristic analysis.

## Summary

- **Confidence tags** in Graphify label every graph edge as either `EXTRACTED` (parsed directly from code) or `INFERRED` (computed by the resolution engine).
- The distinction is documented in [`README.md`](https://github.com/Graphify-Labs/graphify/blob/main/README.md) and implemented in [`graphify/semantics/relationships.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/semantics/relationships.py), with concrete examples visible in [`worked/rsl-siege-manager/graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/worked/rsl-siege-manager/graph.json).
- Use the `graphify` CLI with JSON output filters to separate high-certainty (`EXTRACTED`) relationships from heuristic suggestions.
- These tags prevent false positives in automated refactoring by ensuring that only explicit code relationships trigger changes without manual review.

## Frequently Asked Questions

### What is the difference between EXTRACTED and INFERRED confidence tags in Graphify?

The `EXTRACTED` tag indicates that Graphify found the relationship by parsing source code directly, such as reading an import statement or function call. The `INFERRED` tag means the resolution engine generated the relationship using heuristics like name matching or cross-language similarity, without explicit syntactic evidence in the source files.

### How can I filter query results to show only EXTRACTED edges?

Pipe the JSON output from `graphify query` to a filter like `jq` and select edges where the `"confidence"` field equals `"EXTRACTED"`. Alternatively, when building applications with the Graphify client library, check the `edge.confidence` property programmatically before processing relationships in your business logic.

### Where are confidence tags stored in the Graphify output format?

Confidence tags appear as string values in the `"confidence"` field of edge objects within the generated JSON graph files, as demonstrated in [`worked/rsl-siege-manager/graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/worked/rsl-siege-manager/graph.json). Each edge may also include a numeric `"confidence_score"` field that provides additional granularity for `INFERRED` relationships.

### Can I adjust the thresholds for INFERRED edges in Graphify?

While the resolution engine produces `INFERRED` edges based on internal heuristics defined in modules like [`graphify/semantics/relationships.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/semantics/relationships.py), you can filter these edges post-generation by examining the `confidence_score` values in the JSON output or implementing custom logic in your consuming application to ignore edges below a specific confidence threshold.