# How Wigolo's Tool Scoring and Evidence System Ensure Explainable Results

> Discover how Wigolo's tool scoring and evidence system ensures explainable results with parallel scores and structured evidence objects for complete ranking transparency.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: deep-dive
- Published: 2026-07-29

---

**Wigolo emits parallel scores—a legacy flat relevance metric and a structured evidence object containing full component breakdowns—to provide complete transparency into every ranking decision.**

Modern AI agents require more than ranked search results; they need to understand the reasoning behind each ranking. The Wigolo search tool in KnockOutEZ/wigolo implements a sophisticated tool scoring and evidence system that decomposes every result into inspectable components, enabling LLM-driven agents to explain why specific sources were selected.

## Dual-Score Architecture for Backward Compatibility

Wigolo returns two parallel scoring fields for every result to satisfy both legacy integrations and modern explainability requirements.

### The Legacy relevance_score

This field contains a flat aggregate numeric value in the range **[0, 1]**. Early versions of Wigolo relied solely on this score for simple ranking, and it remains available to maintain backward compatibility with existing callers.

### The Explainable evidence_score

Newer implementations use this object, defined in [`src/types.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/types.ts) (lines 839-866), which contains the same aggregate value via the `final` property plus a complete **component breakdown** that explains precisely why the result received that score. This separation allows the system to support traditional ranking while delivering the transparency required by LLM-driven agents.

## Evidence Score Components Breakdown

The `evidence_score.components` map includes granular signals that contributed to the final ranking. According to the source code in [`src/types.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/types.ts), these components include:

- **base_rrf** – The raw Reciprocal Rank Fusion score before any boost multipliers.
- **context_cosine** – Similarity between the result and query embeddings when context-ranking is active.
- **domain_quality** – A multiplier reflecting the authority and quality of the source host.
- **lexical_alignment** – The degree to which query tokens match the title and snippet text.
- **recency_boost** – Additional weight applied to fresh content when queries indicate temporal intent.
- **engine_consensus** – Count of distinct search engines that returned the same URL, indicating broader agreement.
- **cross_encoder** – Optional signals from cross-encoder reranking models.
- **rare_terms** – Optional bonuses for matching uncommon query terms.

## The Three-Stage Scoring Pipeline

The pipeline that builds these scores operates through three distinct phases, each implemented in specific source files.

### Stage 1: Raw Search Results

Individual search engines return plain lists contained in `RawSearchResult` structures, each carrying only the flat `relevance_score`. At this stage, no explainability metadata exists.

### Stage 2: Orchestrator Core Processing

The [`src/search/core/core-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/core-provider.ts) file (lines 660-698) handles result merging and scoring. During this phase, the system:

1. Merges results from multiple engines.
2. Runs **Reciprocal Rank Fusion (RRF)** to create base rankings.
3. Invokes `applyEvidenceDefault` from [`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts) (lines 193-279).

The `applyEvidenceDefault` function performs several critical operations:

- Extracts high-quality passages using `extractHighlights`.
- Filters excerpts through `isUsefulEvidenceExcerpt` (lines 31-38), discarding snippets shorter than 40 characters or those dominated by markdown link markup.
- Respects the user-supplied `max_tokens_out` budget, truncating evidence to stay within limits.
- Generates deterministic **citation IDs** via `stableCitationId` for reliable excerpt referencing.

### Stage 3: Post-Ranking Boost Modules

After establishing base RRF scores, Wigolo applies specialized boost modules that update specific component fields:

- **[`src/search/core/score-floor.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/score-floor.ts)** (lines 68-78) applies **lexical_alignment** and other quality floors.
- **[`src/search/core/rerank-fold.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/rerank-fold.ts)** (lines 185-215) incorporates **cross_encoder** signals and other reranker-derived metrics.

Each module modifies the relevant component fields before the final `evidence_score` object attaches to the result.

## Budget-Aware Evidence Handling

To ensure returned excerpts contain meaningful prose rather than boilerplate, Wigolo implements strict quality filters. The `isUsefulEvidenceExcerpt` function in [`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts) (lines 31-38) automatically excludes passages that are too short or contain excessive link markup.

The system also enforces **token budget management** through the `max_tokens_out` parameter. When evidence would exceed the specified limit, the pipeline truncates or discards lower-priority excerpts, ensuring the LLM receives only high-value context within cost constraints.

## Consuming Explainable Results in Practice

Developers can inspect the full scoring breakdown using Wigolo's public API. The following example demonstrates how to access both legacy and explainable scores:

```typescript
// Example: run a search with the default explainable output
import { wigolo } from '@wigolo/sdk';

async function demo() {
  const out = await wigolo.search({
    query: 'latest TypeScript 5.5 features',
    max_results: 5,               // limit number of results
    max_tokens_out: 3000,          // overall response budget
    include_full_markdown: false, // keep only evidence, no full page body
  });

  // Walk the results and print the explainable score breakdown
  for (const r of out.results) {
    console.log('▶️', r.title);
    console.log('   URL:', r.url);
    console.log('   Ranking (flat):', r.relevance_score);
    console.log('   Explainable score:', r.evidence_score?.final);
    console.log('   Breakdown:', r.evidence_score?.components);
    console.log('   Why? →', r.evidence_score?.explanation);
    console.log('   Evidence excerpt:', out.evidence?.find(e => e.citation_id === r.evidence?.[0]?.citation_id)?.excerpt);
    console.log('---');
  }
}
demo();

```

To verify the stable citation IDs that tie evidence passages to specific results:

```typescript
// Example: inspect the citation IDs that tie an evidence passage to a result
import { wigolo } from '@wigolo/sdk';

const out = await wigolo.search({ query: 'Node.js 20 release notes', max_results: 3 });
out.evidence?.forEach(e => {
  console.log(`Citation ${e.citation_id} → ${e.title} (${e.url})`);
});

```

Both examples rely on the default API behavior that propagates `evidence_score` and `evidence` automatically without requiring additional flags.

## Summary

- **Wigolo's tool scoring and evidence system** uses dual parallel scores to maintain backward compatibility while providing full explainability.
- The `evidence_score` object contains a `final` aggregate value and a detailed `components` breakdown including signals like `base_rrf`, `domain_quality`, and `lexical_alignment`.
- Source files [`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts) and [`src/search/core/core-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/core-provider.ts) implement a three-stage pipeline that extracts, filters, and scores evidence while respecting token budgets.
- Deterministic `stableCitationId` values and the `isUsefulEvidenceExcerpt` filter ensure that only high-quality, referenceable prose reaches downstream LLM agents.
- Post-ranking boost modules in [`score-floor.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/score-floor.ts) and [`rerank-fold.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/rerank-fold.ts) transparently update specific scoring components before final result delivery.

## Frequently Asked Questions

### What is the difference between relevance_score and evidence_score in Wigolo?

The `relevance_score` is a legacy flat numeric value in the range **[0, 1]** used for simple ranking and backward compatibility. The `evidence_score` is a structured object containing the same final aggregate plus a complete component breakdown—including factors like `base_rrf`, `domain_quality`, and `recency_boost`—that explains precisely why each result ranks where it does.

### How does Wigolo ensure evidence excerpts are high quality?

Wigolo filters all potential evidence through the `isUsefulEvidenceExcerpt` function in [`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts) (lines 31-38), which discards snippets shorter than 40 characters or those dominated by markdown link markup. This ensures that only substantive, meaningful prose reaches the LLM, avoiding boilerplate or navigation-heavy text.

### How does the token budget affect evidence selection?

When users specify `max_tokens_out`, Wigolo's pipeline in `applyEvidenceDefault` monitors the cumulative token count of extracted evidence. If the budget would be exceeded, the system truncates or discards lower-priority excerpts while preserving the most relevant passages, guaranteeing predictable response sizes without sacrificing the highest-quality evidence.

### Can downstream systems reconstruct the ranking logic using the evidence components?

Yes. Because the **full component map travels end-to-end** from [`src/search/core/core-provider.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/core/core-provider.ts) through to the final result, any consumer can visualize numeric contributions (e.g., "Domain quality = 0.12, Recency = 0.08") or re-run the pipeline with different boost configurations to get deterministic, comparable scores. The `evidence_score.explanation` field also provides a human-readable summary of the scoring factors.