How Wigolo's Tool Scoring and Evidence System Ensure Explainable Results
Wigolo emits parallel scores—a legacy flat relevance metric and a structured evidence object containing full component breakdowns—to provide complete transparency into every ranking decision.
Modern AI agents require more than ranked search results; they need to understand the reasoning behind each ranking. The Wigolo search tool in KnockOutEZ/wigolo implements a sophisticated tool scoring and evidence system that decomposes every result into inspectable components, enabling LLM-driven agents to explain why specific sources were selected.
Dual-Score Architecture for Backward Compatibility
Wigolo returns two parallel scoring fields for every result to satisfy both legacy integrations and modern explainability requirements.
The Legacy relevance_score
This field contains a flat aggregate numeric value in the range [0, 1]. Early versions of Wigolo relied solely on this score for simple ranking, and it remains available to maintain backward compatibility with existing callers.
The Explainable evidence_score
Newer implementations use this object, defined in src/types.ts (lines 839-866), which contains the same aggregate value via the final property plus a complete component breakdown that explains precisely why the result received that score. This separation allows the system to support traditional ranking while delivering the transparency required by LLM-driven agents.
Evidence Score Components Breakdown
The evidence_score.components map includes granular signals that contributed to the final ranking. According to the source code in src/types.ts, these components include:
- base_rrf – The raw Reciprocal Rank Fusion score before any boost multipliers.
- context_cosine – Similarity between the result and query embeddings when context-ranking is active.
- domain_quality – A multiplier reflecting the authority and quality of the source host.
- lexical_alignment – The degree to which query tokens match the title and snippet text.
- recency_boost – Additional weight applied to fresh content when queries indicate temporal intent.
- engine_consensus – Count of distinct search engines that returned the same URL, indicating broader agreement.
- cross_encoder – Optional signals from cross-encoder reranking models.
- rare_terms – Optional bonuses for matching uncommon query terms.
The Three-Stage Scoring Pipeline
The pipeline that builds these scores operates through three distinct phases, each implemented in specific source files.
Stage 1: Raw Search Results
Individual search engines return plain lists contained in RawSearchResult structures, each carrying only the flat relevance_score. At this stage, no explainability metadata exists.
Stage 2: Orchestrator Core Processing
The src/search/core/core-provider.ts file (lines 660-698) handles result merging and scoring. During this phase, the system:
- Merges results from multiple engines.
- Runs Reciprocal Rank Fusion (RRF) to create base rankings.
- Invokes
applyEvidenceDefaultfromsrc/search/evidence.ts(lines 193-279).
The applyEvidenceDefault function performs several critical operations:
- Extracts high-quality passages using
extractHighlights. - Filters excerpts through
isUsefulEvidenceExcerpt(lines 31-38), discarding snippets shorter than 40 characters or those dominated by markdown link markup. - Respects the user-supplied
max_tokens_outbudget, truncating evidence to stay within limits. - Generates deterministic citation IDs via
stableCitationIdfor reliable excerpt referencing.
Stage 3: Post-Ranking Boost Modules
After establishing base RRF scores, Wigolo applies specialized boost modules that update specific component fields:
src/search/core/score-floor.ts(lines 68-78) applies lexical_alignment and other quality floors.src/search/core/rerank-fold.ts(lines 185-215) incorporates cross_encoder signals and other reranker-derived metrics.
Each module modifies the relevant component fields before the final evidence_score object attaches to the result.
Budget-Aware Evidence Handling
To ensure returned excerpts contain meaningful prose rather than boilerplate, Wigolo implements strict quality filters. The isUsefulEvidenceExcerpt function in src/search/evidence.ts (lines 31-38) automatically excludes passages that are too short or contain excessive link markup.
The system also enforces token budget management through the max_tokens_out parameter. When evidence would exceed the specified limit, the pipeline truncates or discards lower-priority excerpts, ensuring the LLM receives only high-value context within cost constraints.
Consuming Explainable Results in Practice
Developers can inspect the full scoring breakdown using Wigolo's public API. The following example demonstrates how to access both legacy and explainable scores:
// Example: run a search with the default explainable output
import { wigolo } from '@wigolo/sdk';
async function demo() {
const out = await wigolo.search({
query: 'latest TypeScript 5.5 features',
max_results: 5, // limit number of results
max_tokens_out: 3000, // overall response budget
include_full_markdown: false, // keep only evidence, no full page body
});
// Walk the results and print the explainable score breakdown
for (const r of out.results) {
console.log('▶️', r.title);
console.log(' URL:', r.url);
console.log(' Ranking (flat):', r.relevance_score);
console.log(' Explainable score:', r.evidence_score?.final);
console.log(' Breakdown:', r.evidence_score?.components);
console.log(' Why? →', r.evidence_score?.explanation);
console.log(' Evidence excerpt:', out.evidence?.find(e => e.citation_id === r.evidence?.[0]?.citation_id)?.excerpt);
console.log('---');
}
}
demo();
To verify the stable citation IDs that tie evidence passages to specific results:
// Example: inspect the citation IDs that tie an evidence passage to a result
import { wigolo } from '@wigolo/sdk';
const out = await wigolo.search({ query: 'Node.js 20 release notes', max_results: 3 });
out.evidence?.forEach(e => {
console.log(`Citation ${e.citation_id} → ${e.title} (${e.url})`);
});
Both examples rely on the default API behavior that propagates evidence_score and evidence automatically without requiring additional flags.
Summary
- Wigolo's tool scoring and evidence system uses dual parallel scores to maintain backward compatibility while providing full explainability.
- The
evidence_scoreobject contains afinalaggregate value and a detailedcomponentsbreakdown including signals likebase_rrf,domain_quality, andlexical_alignment. - Source files
src/search/evidence.tsandsrc/search/core/core-provider.tsimplement a three-stage pipeline that extracts, filters, and scores evidence while respecting token budgets. - Deterministic
stableCitationIdvalues and theisUsefulEvidenceExcerptfilter ensure that only high-quality, referenceable prose reaches downstream LLM agents. - Post-ranking boost modules in
score-floor.tsandrerank-fold.tstransparently update specific scoring components before final result delivery.
Frequently Asked Questions
What is the difference between relevance_score and evidence_score in Wigolo?
The relevance_score is a legacy flat numeric value in the range [0, 1] used for simple ranking and backward compatibility. The evidence_score is a structured object containing the same final aggregate plus a complete component breakdown—including factors like base_rrf, domain_quality, and recency_boost—that explains precisely why each result ranks where it does.
How does Wigolo ensure evidence excerpts are high quality?
Wigolo filters all potential evidence through the isUsefulEvidenceExcerpt function in src/search/evidence.ts (lines 31-38), which discards snippets shorter than 40 characters or those dominated by markdown link markup. This ensures that only substantive, meaningful prose reaches the LLM, avoiding boilerplate or navigation-heavy text.
How does the token budget affect evidence selection?
When users specify max_tokens_out, Wigolo's pipeline in applyEvidenceDefault monitors the cumulative token count of extracted evidence. If the budget would be exceeded, the system truncates or discards lower-priority excerpts while preserving the most relevant passages, guaranteeing predictable response sizes without sacrificing the highest-quality evidence.
Can downstream systems reconstruct the ranking logic using the evidence components?
Yes. Because the full component map travels end-to-end from src/search/core/core-provider.ts through to the final result, any consumer can visualize numeric contributions (e.g., "Domain quality = 0.12, Recency = 0.08") or re-run the pipeline with different boost configurations to get deterministic, comparable scores. The evidence_score.explanation field also provides a human-readable summary of the scoring factors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →