Graphify Confidence Scoring: EXTRACTED, INFERRED, and AMBIGUOUS Explained

Graphify assigns every relationship in a knowledge graph one of three confidence labels—EXTRACTED (1.0), INFERRED (0.55–0.95), or AMBIGUOUS (0.1–0.3)—to indicate whether the connection is explicitly confirmed, reasonably inferred, or uncertain and requires manual review.

Graphify, an open-source knowledge graph extraction tool developed by Graphify-Labs, classifies every edge it discovers using a rigorous confidence scoring system. This system helps users distinguish between explicitly verified code relationships and speculative connections that require human validation, ensuring transparency in automated knowledge graph generation.

The Three Confidence Levels

Graphify's extraction sub-agent follows a strict rubric defined in tools/skillgen/fragments/references/shared/extraction-spec.md to categorize every edge.

EXTRACTED: Explicitly Verified Relationships

EXTRACTED edges represent relationships that are explicitly present in the source material. These receive a perfect confidence score of 1.0 because the connection is directly confirmed by syntactic evidence in the AST or literal references in documents.

According to the extraction specification at lines 47-49, this label applies to direct imports, function calls, citations, and explicit cross-references like "see §3.2". Since the evidence is unambiguous, the system never assigns a value other than 1.0 to EXTRACTED relationships.

INFERRED: Contextually Derived Connections

INFERRED edges represent relationships that are reasonably inferred from surrounding context but not spelled out verbatim. Unlike EXTRACTED edges, INFERRED connections use discrete confidence values chosen from a fixed set: 0.95, 0.85, 0.75, 0.65, or 0.55.

As documented in extraction-spec.md lines 50-58, the sub-agent selects values based on evidence strength:

  • 0.95: Direct structural evidence (e.g., shared data structures)
  • 0.85: Strong inference with clear functional alignment but no direct symbol link
  • 0.75: Reasonable inference based on shared problem domain and similar shape
  • 0.65: Weak inference indicating thematic relation without shape evidence
  • 0.55: Speculative but plausible connections based only on surface-level co-occurrence

The specification enforces that the sub-agent never falls back to a generic 0.5 value. If the evidence does not fit one of these discrete tiers, the edge must be marked as AMBIGUOUS instead.

AMBIGUOUS: Uncertain Connections Requiring Review

AMBIGUOUS edges act as a safety net for uncertain relationships. These edges receive scores in the 0.1–0.3 range, signaling that the extractor cannot decide whether a connection truly exists and that human inspection is required.

As noted in lines 59-60 of the extraction specification, this label prevents the system from presenting speculative connections as factual, preserving the integrity of the knowledge graph.

How Confidence Scores Drive the Pipeline

Graphify utilizes confidence scoring throughout its processing pipeline to prioritize review and weight algorithmic decisions.

Sorting and Prioritization

During graph generation, edges are sorted in the order AMBIGUOUS → INFERRED → EXTRACTED to ensure uncertain connections surface first for human review. This sorting logic is visible in worked/rsl-siege-manager/graph.json at lines 5840-5845, where the system explicitly orders edges by confidence level to facilitate quality assurance workflows.

Community Detection

The community detection algorithm considers numeric confidence scores when weighting edges, influencing how nodes cluster into logical communities. Higher-confidence edges exert stronger influence on clustering decisions, while AMBIGUOUS edges contribute minimal weight until verified.

Audit Trail Transparency

Every edge in the final knowledge graph carries its confidence label and score, creating what the documentation in tools/skillgen/fragments/core/devin.md (lines 53-55) calls an "honest audit trail". This transparency allows users to distinguish between what the system found versus what it invented, critical for downstream tasks like automated refactoring or code navigation.

JSON Schema Examples

Below are minimal JSON snippets illustrating how each confidence type appears in Graphify's output, matching the schema defined in the extraction specification:

{
  "edges": [
    {
      "source": "src_auth_login_validatecredentials",
      "target": "src_auth_login_checkpassword",
      "relation": "calls",
      "confidence": "EXTRACTED",
      "confidence_score": 1.0
    },
    {
      "source": "src_auth_login_validatecredentials",
      "target": "src_auth_security_hashpassword",
      "relation": "calls",
      "confidence": "INFERRED",
      "confidence_score": 0.85
    },
    {
      "source": "src_auth_login_validatecredentials",
      "target": "src_auth_ui_loginpage",
      "relation": "conceptually_related_to",
      "confidence": "AMBIGUOUS",
      "confidence_score": 0.2
    }
  ]
}

In this example:

  • The first edge shows a direct function call with EXTRACTED confidence at 1.0
  • The second edge infers a security dependency without direct evidence, marked INFERRED at 0.85
  • The third edge suggests a vague UI connection that cannot be verified, marked AMBIGUOUS at 0.2

Real-world usage statistics appear in GRAPH_REPORT.md files throughout the repository, typically showing distributions like "90% EXTRACTED / 10% INFERRED" that illustrate how the system prioritizes explicit evidence over inference.

Summary

  • EXTRACTED edges (1.0) represent explicitly verified relationships from direct syntactic evidence in source code or documents
  • INFERRED edges use discrete values (0.95, 0.85, 0.75, 0.65, 0.55) based on evidence strength, with strict rules preventing generic 0.5 assignments
  • AMBIGUOUS edges (0.1–0.3) flag uncertain connections for manual review, acting as a safety net against false positives
  • The pipeline uses these scores for sorting (AMBIGUOUS first), community detection weighting, and maintaining transparent audit trails
  • All confidence rules are strictly defined in tools/skillgen/fragments/references/shared/extraction-spec.md

Frequently Asked Questions

What is the difference between INFERRED and AMBIGUOUS in Graphify?

INFERRED relationships have sufficient contextual evidence to justify a specific discrete confidence score (0.55–0.95), while AMBIGUOUS relationships lack enough evidence to qualify for even the lowest INFERRED tier (0.55). If the sub-agent cannot confidently select one of the five discrete INFERRED values, it must label the edge as AMBIGUOUS and assign a score between 0.1 and 0.3, signaling that human review is necessary.

Why does EXTRACTED always use 1.0 instead of a range?

EXTRACTED edges represent relationships with direct syntactic evidence from the AST or explicit literal references in documents. Because these connections are explicitly confirmed by the source material rather than deduced, Graphify treats them as binary facts. The extraction specification at lines 47-49 of extraction-spec.md mandates that EXTRACTED edges always receive exactly 1.0, reflecting absolute certainty based on direct evidence.

How does Graphify use confidence scores when building knowledge graphs?

Graphify uses confidence scores in three primary ways: sorting edges so AMBIGUOUS connections appear first for review, weighting edges during community detection to influence clustering decisions, and maintaining audit trails that distinguish discovered facts from inferred connections. As implemented in worked/rsl-siege-manager/graph.json, the system explicitly orders edges by confidence level to facilitate quality assurance workflows.

Where can I find the official extraction specification for confidence scoring?

The canonical extraction specification resides at tools/skillgen/fragments/references/shared/extraction-spec.md in the Graphify-Labs/graphify repository. This document defines the confidence categories, scoring rubric, and JSON schema for all edges. Platform-specific variations exist in graphify/skills/*/references/extraction-spec.md for different target environments like VS Code or Windows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →