How Fusion Multi-Model Synthesis Aggregates Responses in FreeLLM API
The Fusion feature implements a "panel-of-models + judge" workflow that aggregates responses by running multiple LLMs in parallel, then either selecting the best answer or synthesizing them through a dedicated judge model.
The fusion multi-model synthesis feature in the tashfeenahmed/freellmapi repository transforms single requests into comprehensive answers by orchestrating diverse language models through a sophisticated aggregation pipeline. This system leverages parallel execution, intelligent fallback mechanisms, and optional judicial synthesis to produce higher-quality outputs than any single model could generate alone.
Understanding the Fusion Architecture
At its core, Fusion operates on a panel-of-models principle. When a client sends a fusion request, the system does not route to one provider—instead, it distributes the prompt across multiple independent models and aggregates their outputs. The architecture consists of four primary stages: configuration resolution, panel construction, parallel execution, and answer synthesis.
The implementation lives primarily in server/src/services/fusion.ts, which orchestrates the entire pipeline from request validation to final response assembly. Supporting infrastructure in server/src/services/router.ts handles provider-specific routing, while server/src/lib/request-log.ts ensures full observability through fusion tagged logging.
Step-by-Step Response Aggregation Process
Configuration Resolution
Before executing any model calls, Fusion determines the effective configuration through the resolveEffectiveConfig function in fusion.ts (lines 61-72). This merges the request's inline fusion object with dashboard-saved defaults to establish:
- The panel size
k(number of models to consult) - The specific panel models or auto-selection criteria
- The judge model identifier (if synthesis mode is enabled)
- The synthesis strategy (
synthesizeorbest_of)
This resolution ensures that even minimal requests inherit sensible defaults while allowing granular per-request overrides.
Panel Selection and Diversification
The selectPanel function (lines 49-86) constructs the execution panel through two distinct paths:
Explicit Panel Mode: When the request supplies a models array, each entry is validated for availability, tool/vision support, and deduplicated. Only valid, unique models enter the panel.
Auto-Panel Construction: Without explicit models, Fusion builds a diversified panel using diversifyChain (lines 17-36). This functionorders the fallback chain, filters for request-specific capabilities (tools, vision), and diversifies by both provider and model family to maximize perspective variety. The first k candidates become the active panel, while the next k form an overflow queue for refilling failed slots.
Parallel Model Execution
Each panel member executes via runModelCall (lines 17-30), which handles the complexity of production LLM interactions:
- Retry logic with exponential backoff for transient failures
- Rate-limit handling through
isRateLimitSignalchecks fromserver/src/lib/error-classify.ts - Provider cooldowns to prevent cascading failures
- Key rotation across a model's available API keys
Failed slots automatically refill from the overflow queue, ensuring the target okCount of successful responses is met even when individual providers error out.
Answer Collection Strategy
The runFusion function (lines 104-118) implements a wave-based collection system. Rather than awaiting all promises simultaneously, Fusion processes results in waves—first the initial k slots, then progressively drawing from the overflow list until sufficient successful answers accumulate.
Each completed slot reports through the optional onPanel hook, enabling real-time monitoring of panel performance. This wave approach optimizes latency by returning as soon as the minimum viable answer count is reached, without waiting for slow outliers.
Final Answer Selection
Fusion applies a cascading decision tree to select or generate the final response:
Tool-Call Shortcut: If any panel model returns structured tool_calls, Fusion immediately selects the first such answer verbatim (lines 40-45). The judge is bypassed entirely, and the response wraps the tool call directly. This prioritizes actionable structured output over synthesized text.
Best-of Strategy: When strategy: 'best_of' is configured or fewer than two viable answers exist, Fusion returns the longest panel answer as a proxy for completeness (lines 99-103). This requires no judge invocation and minimizes latency for simple aggregation needs.
Synthesize Mode (Default): For standard synthesis, surviving panel texts feed into a judge model:
-
Prompt Construction:
buildJudgeMessages(lines 99-118) assembles the judge prompt by injectingJUDGE_SYSTEM_PROMPTand formatting each panel answer as--- Response N ---within a user message appended to the original conversation context. -
Judge Execution: The judge invokes either through
runJudgeStreaming(for clients supplyingonJudgeDeltahooks) or standardrunModelCall. Streaming provides token-by-token delivery while regular calls return complete responses. -
Fallback Handling: If the judge fails, the system automatically falls back to the best-of answer, ensuring robustness against judicial model outages.
Implementation Details and Code Examples
The following TypeScript examples demonstrate practical usage patterns from server/src/services/fusion.ts:
import { runFusion, FUSION_MODEL_ID } from './services/fusion.js';
// Simple fusion call with auto-panel selection and synthesis
await runFusion({
messages: [{ role: 'user', content: 'Explain the theory of relativity.' }],
config: {}, // Empty config uses defaults
options: { temperature: 0.7 },
estimatedTokens: 256,
});
Explicit panel configuration with judicial synthesis and debug output:
await runFusion({
messages: [{ role: 'user', content: 'Summarise the last three news items.' }],
config: {
models: ['openrouter/google/gemini-1.5-flash', 'openrouter/anthropic/claude-3-opus-20240229'],
k: 2,
judge: 'openrouter/meta/llama-3.1-70b',
expose_panel: true, // Adds x_fusion field with full panel data
},
options: { temperature: 0.5 },
estimatedTokens: 512,
});
Best-of strategy without judge invocation:
await runFusion({
messages: [{ role: 'user', content: 'Write a haiku about rain.' }],
config: { strategy: 'best_of' },
options: { temperature: 0.9 },
estimatedTokens: 64,
});
The final response always includes a _fusion metadata field indicating which panel models participated and whether a judge was used. When expose_panel: true is set, the response contains x_fusion with the complete panel-wise payload for debugging.
Summary
- Configuration resolution merges inline requests with dashboard defaults via
resolveEffectiveConfigto establish panel parameters. - Panel selection uses
diversifyChainto maximize provider and model family diversity when auto-selecting models. - Parallel execution through
runModelCallimplements robust retry logic, rate-limit handling, and automatic overflow refilling. - Wave-based collection gathers answers progressively until
okCountis satisfied, optimizing for latency. - Three-tier synthesis prioritizes tool calls, falls back to longest-answer selection for
best_ofmode, or synthesizes via a judge model usingbuildJudgeMessages. - Observability is built-in through
fusionrequest-log tags and optionalonPanel/onJudgeDeltahooks.
Frequently Asked Questions
What is the difference between synthesize and best_of strategies in Fusion?
The synthesize strategy (default) sends all panel responses to a dedicated judge model that generates a coherent combined answer using buildJudgeMessages and JUDGE_SYSTEM_PROMPT. The best_of strategy skips the judge entirely and returns the longest panel answer as a proxy for completeness, reducing latency and token costs when simple aggregation suffices.
How does Fusion handle failures in individual panel models?
Fusion implements overflow queue refilling through the wave-based collection system in runFusion. When a panel slot fails, the system automatically substitutes the next candidate from the pre-calculated overflow list until the target okCount of successful responses is achieved. If the judge model fails during synthesis, Fusion falls back to the best-of answer to ensure request completion.
Can I see individual model responses when using Fusion?
Yes. Set expose_panel: true in the fusion configuration to receive the x_fusion field in the response, which contains the full panel-wise payload including each model's individual output. The standard _fusion field always indicates which models participated and whether judicial synthesis was applied.
What happens if a panel model returns a tool call instead of text?
Fusion immediately adopts the first tool call encountered among panel responses, bypassing both the judge and best-of selection logic. This tool-call shortcut (implemented in runFusion lines 40-45) returns the structured output verbatim, prioritizing actionable function calls over synthesized text responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →