How Wigolo's Tool Scoring and Evidence System Ensure Explainable Results

Wigolo emits parallel scores—a legacy flat relevance metric and a structured evidence object containing full component breakdowns—to provide complete transparency into every ranking decision.

Modern AI agents require more than ranked search results; they need to understand the reasoning behind each ranking. The Wigolo search tool in KnockOutEZ/wigolo implements a sophisticated tool scoring and evidence system that decomposes every result into inspectable components, enabling LLM-driven agents to explain why specific sources were selected.

Dual-Score Architecture for Backward Compatibility

Wigolo returns two parallel scoring fields for every result to satisfy both legacy integrations and modern explainability requirements.

The Legacy relevance_score

This field contains a flat aggregate numeric value in the range [0, 1]. Early versions of Wigolo relied solely on this score for simple ranking, and it remains available to maintain backward compatibility with existing callers.

The Explainable evidence_score

Newer implementations use this object, defined in src/types.ts (lines 839-866), which contains the same aggregate value via the final property plus a complete component breakdown that explains precisely why the result received that score. This separation allows the system to support traditional ranking while delivering the transparency required by LLM-driven agents.

Evidence Score Components Breakdown

The evidence_score.components map includes granular signals that contributed to the final ranking. According to the source code in src/types.ts, these components include:

  • base_rrf – The raw Reciprocal Rank Fusion score before any boost multipliers.
  • context_cosine – Similarity between the result and query embeddings when context-ranking is active.
  • domain_quality – A multiplier reflecting the authority and quality of the source host.
  • lexical_alignment – The degree to which query tokens match the title and snippet text.
  • recency_boost – Additional weight applied to fresh content when queries indicate temporal intent.
  • engine_consensus – Count of distinct search engines that returned the same URL, indicating broader agreement.
  • cross_encoder – Optional signals from cross-encoder reranking models.
  • rare_terms – Optional bonuses for matching uncommon query terms.

The Three-Stage Scoring Pipeline

The pipeline that builds these scores operates through three distinct phases, each implemented in specific source files.

Stage 1: Raw Search Results

Individual search engines return plain lists contained in RawSearchResult structures, each carrying only the flat relevance_score. At this stage, no explainability metadata exists.

Stage 2: Orchestrator Core Processing

The src/search/core/core-provider.ts file (lines 660-698) handles result merging and scoring. During this phase, the system:

  1. Merges results from multiple engines.
  2. Runs Reciprocal Rank Fusion (RRF) to create base rankings.
  3. Invokes applyEvidenceDefault from src/search/evidence.ts (lines 193-279).

The applyEvidenceDefault function performs several critical operations:

  • Extracts high-quality passages using extractHighlights.
  • Filters excerpts through isUsefulEvidenceExcerpt (lines 31-38), discarding snippets shorter than 40 characters or those dominated by markdown link markup.
  • Respects the user-supplied max_tokens_out budget, truncating evidence to stay within limits.
  • Generates deterministic citation IDs via stableCitationId for reliable excerpt referencing.

Stage 3: Post-Ranking Boost Modules

After establishing base RRF scores, Wigolo applies specialized boost modules that update specific component fields:

Each module modifies the relevant component fields before the final evidence_score object attaches to the result.

Budget-Aware Evidence Handling

To ensure returned excerpts contain meaningful prose rather than boilerplate, Wigolo implements strict quality filters. The isUsefulEvidenceExcerpt function in src/search/evidence.ts (lines 31-38) automatically excludes passages that are too short or contain excessive link markup.

The system also enforces token budget management through the max_tokens_out parameter. When evidence would exceed the specified limit, the pipeline truncates or discards lower-priority excerpts, ensuring the LLM receives only high-value context within cost constraints.

Consuming Explainable Results in Practice

Developers can inspect the full scoring breakdown using Wigolo's public API. The following example demonstrates how to access both legacy and explainable scores:

// Example: run a search with the default explainable output
import { wigolo } from '@wigolo/sdk';

async function demo() {
  const out = await wigolo.search({
    query: 'latest TypeScript 5.5 features',
    max_results: 5,               // limit number of results
    max_tokens_out: 3000,          // overall response budget
    include_full_markdown: false, // keep only evidence, no full page body
  });

  // Walk the results and print the explainable score breakdown
  for (const r of out.results) {
    console.log('▶️', r.title);
    console.log('   URL:', r.url);
    console.log('   Ranking (flat):', r.relevance_score);
    console.log('   Explainable score:', r.evidence_score?.final);
    console.log('   Breakdown:', r.evidence_score?.components);
    console.log('   Why? →', r.evidence_score?.explanation);
    console.log('   Evidence excerpt:', out.evidence?.find(e => e.citation_id === r.evidence?.[0]?.citation_id)?.excerpt);
    console.log('---');
  }
}
demo();

To verify the stable citation IDs that tie evidence passages to specific results:

// Example: inspect the citation IDs that tie an evidence passage to a result
import { wigolo } from '@wigolo/sdk';

const out = await wigolo.search({ query: 'Node.js 20 release notes', max_results: 3 });
out.evidence?.forEach(e => {
  console.log(`Citation ${e.citation_id} → ${e.title} (${e.url})`);
});

Both examples rely on the default API behavior that propagates evidence_score and evidence automatically without requiring additional flags.

Summary

  • Wigolo's tool scoring and evidence system uses dual parallel scores to maintain backward compatibility while providing full explainability.
  • The evidence_score object contains a final aggregate value and a detailed components breakdown including signals like base_rrf, domain_quality, and lexical_alignment.
  • Source files src/search/evidence.ts and src/search/core/core-provider.ts implement a three-stage pipeline that extracts, filters, and scores evidence while respecting token budgets.
  • Deterministic stableCitationId values and the isUsefulEvidenceExcerpt filter ensure that only high-quality, referenceable prose reaches downstream LLM agents.
  • Post-ranking boost modules in score-floor.ts and rerank-fold.ts transparently update specific scoring components before final result delivery.

Frequently Asked Questions

What is the difference between relevance_score and evidence_score in Wigolo?

The relevance_score is a legacy flat numeric value in the range [0, 1] used for simple ranking and backward compatibility. The evidence_score is a structured object containing the same final aggregate plus a complete component breakdown—including factors like base_rrf, domain_quality, and recency_boost—that explains precisely why each result ranks where it does.

How does Wigolo ensure evidence excerpts are high quality?

Wigolo filters all potential evidence through the isUsefulEvidenceExcerpt function in src/search/evidence.ts (lines 31-38), which discards snippets shorter than 40 characters or those dominated by markdown link markup. This ensures that only substantive, meaningful prose reaches the LLM, avoiding boilerplate or navigation-heavy text.

How does the token budget affect evidence selection?

When users specify max_tokens_out, Wigolo's pipeline in applyEvidenceDefault monitors the cumulative token count of extracted evidence. If the budget would be exceeded, the system truncates or discards lower-priority excerpts while preserving the most relevant passages, guaranteeing predictable response sizes without sacrificing the highest-quality evidence.

Can downstream systems reconstruct the ranking logic using the evidence components?

Yes. Because the full component map travels end-to-end from src/search/core/core-provider.ts through to the final result, any consumer can visualize numeric contributions (e.g., "Domain quality = 0.12, Recency = 0.08") or re-run the pipeline with different boost configurations to get deterministic, comparable scores. The evidence_score.explanation field also provides a human-readable summary of the scoring factors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →