# How Wigolo's Research Tool Decomposes Questions and Generates Synthesized Reports

> Discover how Wigolo's research tool decomposes questions, executes parallel searches, and synthesizes results into citation-rich markdown reports with LLM generation or deterministic fallback.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: deep-dive
- Published: 2026-07-29

---

**Wigolo's research tool breaks complex user questions into targeted sub-queries, executes parallel searches across multiple engines, and synthesizes the results into citation-rich markdown reports using either LLM generation or deterministic fallback formatting.**

The `KnockOutEZ/wigolo` repository provides an open-source research automation framework that transforms vague information needs into structured intelligence. At its core, wigolo's research tool implements a three-stage pipeline—**question decomposition**, **parallel evidence gathering**, and **report synthesis**—to deliver comprehensive answers without manual search curation.

## Question Decomposition: From Broad Queries to Targeted Sub-Queries

The decomposition stage lives in [`src/research/decompose.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/decompose.ts) and converts a single research question into a fan-out of specific search queries. This process combines deterministic pattern matching, template generation, and optional LLM sampling.

### Query Type Detection and Classification

The `detectQueryType` function classifies incoming questions into four categories—**comparison**, **how-to**, **concept**, or **general**—using optimized regular expressions:

```typescript
// src/research/decompose.ts
export function detectQueryType(question: string): QueryType {
  const q = question.trim().toLowerCase();
  if (/\bvs\.?\s/i.test(q) || /\bversus\b/i.test(q) || /^compare\b/i.test(q)) return 'comparison';
  if (/^how\s+(?:to|do|does|can|should)\b/i.test(q))               return 'how-to';
  if (/^(?:what\s+(?:is|are)|explain|overview\s+of|describe)\b/i.test(q))
    return 'concept';
  return 'general';
}

```

### Entity Extraction and Template Generation

For **comparison** queries, `extractComparisonEntities` parses the question to isolate the entities being compared and any contextual nouns (like "runtime" or "library"). The system then generates a deterministic set of sub-queries:

```typescript
// src/research/decompose.ts
if (queryType === 'comparison') {
  const { entities, context } = extractComparisonEntities(question);
  // Per-entity feature queries
  for (const e of entities) queries.push(`${e} ${context} features performance`);
  // Cross-comparison queries
  for (let i = 0; i < entities.length; i++)
    for (let j = i + 1; j < entities.length; j++)
      queries.push(`${entities[i]} vs ${entities[j]} ${context} comparison`);
  // Decision query
  queries.push(`${entities.join(' vs ')} which to choose ${context}`);
}

```

**How-to** and **concept** queries follow similar template logic, extracting core noun phrases to build targeted search variants.

### LLM Sampling and Fallback Heuristics

If the client provides a `SamplingCapableServer`, the tool attempts `decomposeWithSampling`, requesting the LLM to return exactly *N* sub-queries in a JSON object with the shape `{ "subQueries": [...] }`.

When sampling fails or is unavailable, `decomposeWithFallback` extracts noun phrases, splits the query at clause boundaries, and generates keyword variants until the target count is reached. The final list is deduplicated, and the original verbatim question is **prepended** to ensure the exact user phrasing always participates in the search (see `runResearchPipeline` lines 78-92).

## Parallel Search Execution and Evidence Ranking

The `runResearchPipeline` function in [`src/research/pipeline.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/pipeline.ts) orchestrates the evidence-gathering stage. It receives the decomposed sub-queries and executes them against configured search engines with strict resource management.

### Parallel Exploration and Time Budgeting

The pipeline calls `exploreInParallel` to execute sub-queries concurrently, respecting two critical parameters:
- `SEARCH_TOTAL_BUDGET_MS`: Maximum time for the entire search phase
- `SEARCH_PER_QUERY_BUDGET_MS`: Maximum time allocated to any single sub-query

### Deduplication, Filtering, and Reranking

Raw results undergo three refinement steps:
1. **Deduplication**: `deduplicateResults` removes redundant URLs across search engines
2. **Domain filtering**: Allow/deny lists filter sources by domain reputation
3. **Reranking**: A cross-encoder model in [`src/search/rerank.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/rerank.ts) reorders results by relevance

Source validation prunes low-quality results using URL shape checks, content-gate detection, and configurable score floors (see [`src/research/pipeline.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/pipeline.ts) lines 158-190).

## Synthesizing the Final Report

The synthesis stage in [`src/research/synthesize.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/synthesize.ts) transforms filtered sources into the final markdown output, offering two execution paths based on LLM availability.

### LLM-Driven Synthesis

When a sampling server is available, `synthesizeWithSampling` constructs a prompt containing:
- The original research question
- A numbered list of source blocks (title, URL, trimmed markdown content)
- Target length constraints (`reportChars`)

The LLM returns a comprehensive markdown report with inline citation markers such as `[1]`, `[2]`, which map back to the source list.

### Deterministic Fallback Reporting

If LLM sampling fails or is disabled, `buildFallbackReport` assembles a structured markdown document:
- Header: `## Research: <question>`

- Introduction noting the number of sources
- Sections for each source containing the title, URL, and excerpt (respecting per-source and total character caps)

Both paths return a `ResearchOutput` object containing the final `report`, an array of `citations`, and a boolean `samplingUsed` flag.

## Implementation Examples

### Calling the Research Endpoint via JavaScript Client

```typescript
import { WigoloClient } from '@wigolo/sdk';

const client = new WigoloClient({
  apiKey: process.env.WIGOGO_API_KEY,
  endpoint: 'https://api.wigolo.com/v1',
});

async function researchDemo() {
  const result = await client.research({
    question: 'React hooks vs Vue composition API',
    depth: 'standard',          // quick | standard | comprehensive
    max_sources: 12,
    include_full_markdown: true
  });

  console.log(result.report);
  result.citations.forEach(c => 
    console.log(`[${c.index}] ${c.title} – ${c.url}`)
  );
}

researchDemo().catch(console.error);

```

*This internally calls `src/tools/research.ts → handleResearch`, which triggers the full pipeline.*

### Direct Pipeline Invocation for Custom Tooling

```typescript
import { runResearchPipeline } from './src/research/pipeline.js';
import { defaultEngines } from './src/search/engines.js';
import { router } from './src/fetch/router.js';

const input = {
  question: 'How to deploy a Next.js app to AWS?',
  depth: 'quick',
  max_sources: 8,
};

const output = await runResearchPipeline(input, defaultEngines, router);
console.log(output.report); // markdown report with citations

```

### Inspecting Generated Sub-Queries

```typescript
import { decomposeQuestion } from './src/research/decompose.js';

const { subQueries, queryType } = await decomposeQuestion(
  'SQLite vs PostgreSQL vs DuckDB for analytics',
  'comprehensive'
);
console.log('Query type:', queryType);
console.log('Sub-queries:', subQueries);

```

## Summary

- **Three-stage architecture**: Question decomposition, parallel search with ranking, and report synthesis form the complete research pipeline in `KnockOutEZ/wigolo`.
- **Deterministic decomposition**: The `detectQueryType` function in [`src/research/decompose.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/decompose.ts) uses regex patterns to classify queries and generate templates for comparisons, how-to guides, and concept explanations.
- **Intelligent search orchestration**: `runResearchPipeline` executes sub-queries in parallel with time budgets, then deduplicates, filters, and reranks results using cross-encoders.
- **Dual synthesis paths**: Reports generate via LLM sampling (`synthesizeWithSampling`) with citation markers or deterministic fallback (`buildFallbackReport`) for offline environments.
- **Graceful degradation**: Heuristic fallbacks for both decomposition and synthesis ensure the tool produces usable outputs even when LLM services are unavailable.

## Frequently Asked Questions

### What query types does wigolo's research tool recognize?

Wigolo classifies queries into **comparison**, **how-to**, **concept**, and **general** types using the `detectQueryType` function in [`src/research/decompose.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/decompose.ts). Comparison queries trigger entity extraction and side-by-side sub-query generation, while how-to queries focus on procedural decomposition.

### How does the tool handle complex comparison questions?

For comparison queries, `extractComparisonEntities` isolates the compared items and context, then generates three query variants: per-entity feature searches, direct versus comparisons, and a decision-focused "which to choose" query. This ensures comprehensive coverage of both individual characteristics and comparative analysis.

### What happens when LLM sampling is unavailable?

The tool implements deterministic fallbacks at both stages. During decomposition, `decomposeWithFallback` uses noun phrase extraction and clause splitting to generate sub-queries. During synthesis, `buildFallbackReport` assembles a structured markdown document from source excerpts without LLM assistance, ensuring reliable operation in offline or rate-limited scenarios.

### How are search results ranked and filtered?

Results undergo deduplication via `deduplicateResults`, domain filtering based on allow/deny lists, and reranking through a cross-encoder model in [`src/search/rerank.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/rerank.ts). The pipeline applies score floors and content-gate detection (lines 158-190 in [`src/research/pipeline.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/research/pipeline.ts)) to exclude low-quality or inaccessible sources before synthesis.