How to Use the Fusion Model in FreeLLMAPI: Multi-Model Synthesis Guide
The fusion model is a virtual model in FreeLLMAPI that aggregates outputs from multiple free models in parallel and synthesizes them into a single coherent response using a judge model.
FreeLLMAPI provides a unique fusion virtual model that enables multi-model synthesis without requiring separate API calls. When you specify model: "fusion" in your request, the system orchestrates a panel of underlying models, executes them concurrently, and returns a unified answer. This approach improves reliability and answer quality by combining diverse perspectives from different providers and model families.
What Is the Fusion Model?
The fusion model is not a single underlying model but a meta-router implemented in server/src/services/fusion.ts. When invoked, it treats your prompt as a synthesis task and distributes it across a curated panel of available models. The system handles all complexity—including rate limiting, key rotation, retry logic, and response aggregation—transparently.
According to the FreeLLMAPI source code, the fusion workflow is orchestrated by the runFusion function. This function merges inline configuration with dashboard defaults, selects a diversified panel, executes calls in parallel, and manages the final synthesis step.
How the Fusion Architecture Works
The fusion pipeline executes three distinct phases for every request:
Panel Selection and Diversification
First, the system builds a panel of models to answer your query. The default panel size is 4 models with a hard cap of 8 models (as defined in server/src/services/fusion.ts lines 40–64).
The selectPanel function draws from either an explicit list you provide or the active fallback chain. To ensure diverse perspectives, the system calls diversifyChain (implemented in lines 17–36 of server/src/services/fusion.ts), which deduplicates models by provider first, then by model family. This two-pass algorithm prevents the panel from containing redundant models (e.g., two different GPT-3.5 variants).
If the panel cannot be filled immediately, remaining slots enter an overflow queue that refills failed slots until the required number of successful answers is reached.
Parallel Execution
Each panel member receives identical prompts and executes via runModelCall (provided by server/src/services/router.ts). This utility handles rate-limit management, key leasing, usage accounting, and request logging for every individual model call.
All fusion-related traffic is tagged with the constant FUSION_TAG = 'fusion' (lines 32–35), allowing analytics systems to distinguish fusion traffic from ordinary requests. The calls run in parallel, and the system waits for all successful responses before proceeding to synthesis.
Synthesis and Judging
Once panel responses arrive, the system evaluates them in server/src/services/fusion.ts (lines 75–78 and 99–104). The behavior depends on your chosen strategy:
synthesize(default): Responses pass to a judge model viabuildJudgeMessages. The judge (either the top-ranked available model or a user-specifiedfusion.judge) receives a system prompt instructing it to "combine the drafts into one self-contained answer."best_of: The judge step is skipped, and the system returns the longest panel answer directly.
Tool call handling: If any panel answer contains tool calls, the first such answer wins immediately and the judge is omitted, as tool calls must remain atomic.
Configuring Fusion Requests
You control fusion behavior through an inline fusion configuration object or dashboard-saved defaults resolved by resolveEffectiveConfig.
Inline Configuration Options
Pass these fields in the fusion object of your request body:
k: Panel size (default 4, maximum 8)models: Explicit array of model IDs to use (e.g.,["gpt-3.5-turbo", "mistral-medium"])strategy: Either"synthesize"(default) or"best_of"judge: Specific model ID to use as the judge (overrides top-ranked default)expose_panel: Boolean that, whentrue, includes detailed panel metadata in thex_fusionresponse header
Default Behavior and Fallback Chains
If you do not specify a models array, the system invokes getOrderedFusionChain from server/src/services/model-groups.ts to build the panel from the active fallback chain. This ensures that even without explicit configuration, fusion requests utilize the best available free models according to your deployment's capacity.
Implementation Examples
Basic Python Client Request
Use the OpenAI-compatible client to invoke the fusion model:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1",
api_key="freellmapi-your-unified-key",
)
resp = client.chat.completions.create(
model="fusion",
messages=[{"role": "user", "content": "Explain the Pythagorean theorem in simple terms."}],
extra_body={
"fusion": {
"k": 5,
"strategy": "synthesize",
"expose_panel": True
}
}
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))
if "x_fusion" in resp.headers:
print("Panel details:", resp.headers["x_fusion"])
Direct API Request with Custom Panel
Specify exact models for the panel using a direct HTTP request:
curl http://localhost:3001/v1/chat/completions \
-H "Authorization: Bearer freellmapi-your-unified-key" \
-H "Content-Type: application/json" \
-d '{
"model": "fusion",
"messages": [{"role": "user", "content": "Compare the climates of Tokyo and Vancouver."}],
"fusion": {
"models": ["gpt-3.5-turbo", "mistral-medium", "gemini-1.5-flash"],
"strategy": "synthesize",
"expose_panel": true
}
}'
The response includes an _fusion field containing the panel members and judge information. When expose_panel is enabled, the x_fusion header provides additional debugging metadata.
Streaming Fusion Responses
Fusion supports streaming for real-time synthesis:
stream = client.chat.completions.create(
model="fusion",
messages=[{"role": "user", "content": "Write a short story about a talking cat."}],
stream=True,
)
for chunk in stream:
# Tokens from the judge model appear as they are generated
print(chunk.choices[0].delta.content or "", end="", flush=True)
Summary
- The fusion virtual model in FreeLLMAPI enables multi-model synthesis by orchestrating parallel requests across diverse providers.
- The default panel contains 4 models (maximum 8) selected via
diversifyChainto ensure provider and family diversity. - Execution occurs in
server/src/services/fusion.tsviarunFusion, which handles configuration resolution, parallel execution viarunModelCall, and response synthesis. - Choose between
synthesize(judge model combines answers) andbest_of(longest answer wins) strategies. - Responses include an
_fusionmetadata field and optionalx_fusionheader for debugging whenexpose_panelis enabled.
Frequently Asked Questions
What is the maximum number of models in a fusion panel?
The fusion panel supports a default size of 4 models with an absolute maximum of 8 models. This limit is enforced in server/src/services/fusion.ts (lines 40–64) to balance synthesis quality against latency and token costs.
How does the fusion model handle tool calls?
If any panel member returns a response containing tool calls, the fusion system immediately returns that response and skips the judge synthesis step. This rule exists because tool calls must remain atomic and unmodified. The logic is implemented in the runFusion function within server/src/services/fusion.ts.
Can I see which models contributed to a fusion response?
Yes. Every fusion response includes an _fusion field in the JSON payload that lists the panel members and the judge model used (if applicable). Additionally, setting fusion.expose_panel: true in your request adds an x_fusion header with detailed debugging information about the execution pipeline.
What is the difference between synthesize and best_of strategies?
The synthesize strategy (default) sends all panel responses to a judge model that combines them into a single coherent answer. The best_of strategy skips synthesis and returns the longest panel answer directly. Use best_of when you want faster responses or when answers are expected to be structurally similar, and synthesize when you need coherent integration of diverse perspectives.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →