How Wigolo's Research Tool Decomposes Questions and Generates Synthesized Reports

Wigolo's research tool breaks complex user questions into targeted sub-queries, executes parallel searches across multiple engines, and synthesizes the results into citation-rich markdown reports using either LLM generation or deterministic fallback formatting.

The KnockOutEZ/wigolo repository provides an open-source research automation framework that transforms vague information needs into structured intelligence. At its core, wigolo's research tool implements a three-stage pipeline—question decomposition, parallel evidence gathering, and report synthesis—to deliver comprehensive answers without manual search curation.

Question Decomposition: From Broad Queries to Targeted Sub-Queries

The decomposition stage lives in src/research/decompose.ts and converts a single research question into a fan-out of specific search queries. This process combines deterministic pattern matching, template generation, and optional LLM sampling.

Query Type Detection and Classification

The detectQueryType function classifies incoming questions into four categories—comparison, how-to, concept, or general—using optimized regular expressions:

// src/research/decompose.ts
export function detectQueryType(question: string): QueryType {
  const q = question.trim().toLowerCase();
  if (/\bvs\.?\s/i.test(q) || /\bversus\b/i.test(q) || /^compare\b/i.test(q)) return 'comparison';
  if (/^how\s+(?:to|do|does|can|should)\b/i.test(q))               return 'how-to';
  if (/^(?:what\s+(?:is|are)|explain|overview\s+of|describe)\b/i.test(q))
    return 'concept';
  return 'general';
}

Entity Extraction and Template Generation

For comparison queries, extractComparisonEntities parses the question to isolate the entities being compared and any contextual nouns (like "runtime" or "library"). The system then generates a deterministic set of sub-queries:

// src/research/decompose.ts
if (queryType === 'comparison') {
  const { entities, context } = extractComparisonEntities(question);
  // Per-entity feature queries
  for (const e of entities) queries.push(`${e} ${context} features performance`);
  // Cross-comparison queries
  for (let i = 0; i < entities.length; i++)
    for (let j = i + 1; j < entities.length; j++)
      queries.push(`${entities[i]} vs ${entities[j]} ${context} comparison`);
  // Decision query
  queries.push(`${entities.join(' vs ')} which to choose ${context}`);
}

How-to and concept queries follow similar template logic, extracting core noun phrases to build targeted search variants.

LLM Sampling and Fallback Heuristics

If the client provides a SamplingCapableServer, the tool attempts decomposeWithSampling, requesting the LLM to return exactly N sub-queries in a JSON object with the shape { "subQueries": [...] }.

When sampling fails or is unavailable, decomposeWithFallback extracts noun phrases, splits the query at clause boundaries, and generates keyword variants until the target count is reached. The final list is deduplicated, and the original verbatim question is prepended to ensure the exact user phrasing always participates in the search (see runResearchPipeline lines 78-92).

Parallel Search Execution and Evidence Ranking

The runResearchPipeline function in src/research/pipeline.ts orchestrates the evidence-gathering stage. It receives the decomposed sub-queries and executes them against configured search engines with strict resource management.

Parallel Exploration and Time Budgeting

The pipeline calls exploreInParallel to execute sub-queries concurrently, respecting two critical parameters:

  • SEARCH_TOTAL_BUDGET_MS: Maximum time for the entire search phase
  • SEARCH_PER_QUERY_BUDGET_MS: Maximum time allocated to any single sub-query

Deduplication, Filtering, and Reranking

Raw results undergo three refinement steps:

  1. Deduplication: deduplicateResults removes redundant URLs across search engines
  2. Domain filtering: Allow/deny lists filter sources by domain reputation
  3. Reranking: A cross-encoder model in src/search/rerank.ts reorders results by relevance

Source validation prunes low-quality results using URL shape checks, content-gate detection, and configurable score floors (see src/research/pipeline.ts lines 158-190).

Synthesizing the Final Report

The synthesis stage in src/research/synthesize.ts transforms filtered sources into the final markdown output, offering two execution paths based on LLM availability.

LLM-Driven Synthesis

When a sampling server is available, synthesizeWithSampling constructs a prompt containing:

  • The original research question
  • A numbered list of source blocks (title, URL, trimmed markdown content)
  • Target length constraints (reportChars)

The LLM returns a comprehensive markdown report with inline citation markers such as [1], [2], which map back to the source list.

Deterministic Fallback Reporting

If LLM sampling fails or is disabled, buildFallbackReport assembles a structured markdown document:

  • Header: ## Research: <question>

  • Introduction noting the number of sources

  • Sections for each source containing the title, URL, and excerpt (respecting per-source and total character caps)

Both paths return a ResearchOutput object containing the final report, an array of citations, and a boolean samplingUsed flag.

Implementation Examples

Calling the Research Endpoint via JavaScript Client

import { WigoloClient } from '@wigolo/sdk';

const client = new WigoloClient({
  apiKey: process.env.WIGOGO_API_KEY,
  endpoint: 'https://api.wigolo.com/v1',
});

async function researchDemo() {
  const result = await client.research({
    question: 'React hooks vs Vue composition API',
    depth: 'standard',          // quick | standard | comprehensive
    max_sources: 12,
    include_full_markdown: true
  });

  console.log(result.report);
  result.citations.forEach(c => 
    console.log(`[${c.index}] ${c.title} – ${c.url}`)
  );
}

researchDemo().catch(console.error);

This internally calls src/tools/research.ts → handleResearch, which triggers the full pipeline.

Direct Pipeline Invocation for Custom Tooling

import { runResearchPipeline } from './src/research/pipeline.js';
import { defaultEngines } from './src/search/engines.js';
import { router } from './src/fetch/router.js';

const input = {
  question: 'How to deploy a Next.js app to AWS?',
  depth: 'quick',
  max_sources: 8,
};

const output = await runResearchPipeline(input, defaultEngines, router);
console.log(output.report); // markdown report with citations

Inspecting Generated Sub-Queries

import { decomposeQuestion } from './src/research/decompose.js';

const { subQueries, queryType } = await decomposeQuestion(
  'SQLite vs PostgreSQL vs DuckDB for analytics',
  'comprehensive'
);
console.log('Query type:', queryType);
console.log('Sub-queries:', subQueries);

Summary

  • Three-stage architecture: Question decomposition, parallel search with ranking, and report synthesis form the complete research pipeline in KnockOutEZ/wigolo.
  • Deterministic decomposition: The detectQueryType function in src/research/decompose.ts uses regex patterns to classify queries and generate templates for comparisons, how-to guides, and concept explanations.
  • Intelligent search orchestration: runResearchPipeline executes sub-queries in parallel with time budgets, then deduplicates, filters, and reranks results using cross-encoders.
  • Dual synthesis paths: Reports generate via LLM sampling (synthesizeWithSampling) with citation markers or deterministic fallback (buildFallbackReport) for offline environments.
  • Graceful degradation: Heuristic fallbacks for both decomposition and synthesis ensure the tool produces usable outputs even when LLM services are unavailable.

Frequently Asked Questions

What query types does wigolo's research tool recognize?

Wigolo classifies queries into comparison, how-to, concept, and general types using the detectQueryType function in src/research/decompose.ts. Comparison queries trigger entity extraction and side-by-side sub-query generation, while how-to queries focus on procedural decomposition.

How does the tool handle complex comparison questions?

For comparison queries, extractComparisonEntities isolates the compared items and context, then generates three query variants: per-entity feature searches, direct versus comparisons, and a decision-focused "which to choose" query. This ensures comprehensive coverage of both individual characteristics and comparative analysis.

What happens when LLM sampling is unavailable?

The tool implements deterministic fallbacks at both stages. During decomposition, decomposeWithFallback uses noun phrase extraction and clause splitting to generate sub-queries. During synthesis, buildFallbackReport assembles a structured markdown document from source excerpts without LLM assistance, ensuring reliable operation in offline or rate-limited scenarios.

How are search results ranked and filtered?

Results undergo deduplication via deduplicateResults, domain filtering based on allow/deny lists, and reranking through a cross-encoder model in src/search/rerank.ts. The pipeline applies score floors and content-gate detection (lines 158-190 in src/research/pipeline.ts) to exclude low-quality or inaccessible sources before synthesis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →