# How to Use the Fusion Model in FreeLLMAPI: Multi-Model Synthesis Guide

> Learn to use the fusion model in FreeLLMAPI to synthesize outputs from multiple free models into a single coherent response. Our guide makes multi-model synthesis easy.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-29

---

**The fusion model is a virtual model in FreeLLMAPI that aggregates outputs from multiple free models in parallel and synthesizes them into a single coherent response using a judge model.**

FreeLLMAPI provides a unique **fusion** virtual model that enables multi-model synthesis without requiring separate API calls. When you specify `model: "fusion"` in your request, the system orchestrates a panel of underlying models, executes them concurrently, and returns a unified answer. This approach improves reliability and answer quality by combining diverse perspectives from different providers and model families.

## What Is the Fusion Model?

The fusion model is not a single underlying model but a **meta-router** implemented in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts). When invoked, it treats your prompt as a synthesis task and distributes it across a curated panel of available models. The system handles all complexity—including rate limiting, key rotation, retry logic, and response aggregation—transparently.

According to the FreeLLMAPI source code, the fusion workflow is orchestrated by the `runFusion` function. This function merges inline configuration with dashboard defaults, selects a diversified panel, executes calls in parallel, and manages the final synthesis step.

## How the Fusion Architecture Works

The fusion pipeline executes three distinct phases for every request:

### Panel Selection and Diversification

First, the system builds a **panel** of models to answer your query. The default panel size is **4 models** with a hard cap of **8 models** (as defined in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) lines 40–64).

The `selectPanel` function draws from either an explicit list you provide or the active fallback chain. To ensure diverse perspectives, the system calls `diversifyChain` (implemented in lines 17–36 of [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts)), which deduplicates models by **provider** first, then by **model family**. This two-pass algorithm prevents the panel from containing redundant models (e.g., two different GPT-3.5 variants).

If the panel cannot be filled immediately, remaining slots enter an **overflow queue** that refills failed slots until the required number of successful answers is reached.

### Parallel Execution

Each panel member receives identical prompts and executes via `runModelCall` (provided by [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts)). This utility handles rate-limit management, key leasing, usage accounting, and request logging for every individual model call.

All fusion-related traffic is tagged with the constant `FUSION_TAG = 'fusion'` (lines 32–35), allowing analytics systems to distinguish fusion traffic from ordinary requests. The calls run in parallel, and the system waits for all successful responses before proceeding to synthesis.

### Synthesis and Judging

Once panel responses arrive, the system evaluates them in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) (lines 75–78 and 99–104). The behavior depends on your chosen **strategy**:

- **`synthesize`** (default): Responses pass to a **judge model** via `buildJudgeMessages`. The judge (either the top-ranked available model or a user-specified `fusion.judge`) receives a system prompt instructing it to "combine the drafts into one self-contained answer."
- **`best_of`**: The judge step is skipped, and the system returns the longest panel answer directly.

**Tool call handling**: If any panel answer contains tool calls, the first such answer wins immediately and the judge is omitted, as tool calls must remain atomic.

## Configuring Fusion Requests

You control fusion behavior through an inline `fusion` configuration object or dashboard-saved defaults resolved by `resolveEffectiveConfig`.

### Inline Configuration Options

Pass these fields in the `fusion` object of your request body:

- **`k`**: Panel size (default 4, maximum 8)
- **`models`**: Explicit array of model IDs to use (e.g., `["gpt-3.5-turbo", "mistral-medium"]`)
- **`strategy`**: Either `"synthesize"` (default) or `"best_of"`
- **`judge`**: Specific model ID to use as the judge (overrides top-ranked default)
- **`expose_panel`**: Boolean that, when `true`, includes detailed panel metadata in the `x_fusion` response header

### Default Behavior and Fallback Chains

If you do not specify a `models` array, the system invokes `getOrderedFusionChain` from [`server/src/services/model-groups.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-groups.ts) to build the panel from the active fallback chain. This ensures that even without explicit configuration, fusion requests utilize the best available free models according to your deployment's capacity.

## Implementation Examples

### Basic Python Client Request

Use the OpenAI-compatible client to invoke the fusion model:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="freellmapi-your-unified-key",
)

resp = client.chat.completions.create(
    model="fusion",
    messages=[{"role": "user", "content": "Explain the Pythagorean theorem in simple terms."}],
    extra_body={
        "fusion": {
            "k": 5,
            "strategy": "synthesize",
            "expose_panel": True
        }
    }
)

print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))
if "x_fusion" in resp.headers:
    print("Panel details:", resp.headers["x_fusion"])

```

### Direct API Request with Custom Panel

Specify exact models for the panel using a direct HTTP request:

```bash
curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "fusion",
    "messages": [{"role": "user", "content": "Compare the climates of Tokyo and Vancouver."}],
    "fusion": {
      "models": ["gpt-3.5-turbo", "mistral-medium", "gemini-1.5-flash"],
      "strategy": "synthesize",
      "expose_panel": true
    }
  }'

```

The response includes an `_fusion` field containing the panel members and judge information. When `expose_panel` is enabled, the `x_fusion` header provides additional debugging metadata.

### Streaming Fusion Responses

Fusion supports streaming for real-time synthesis:

```python
stream = client.chat.completions.create(
    model="fusion",
    messages=[{"role": "user", "content": "Write a short story about a talking cat."}],
    stream=True,
)

for chunk in stream:
    # Tokens from the judge model appear as they are generated

    print(chunk.choices[0].delta.content or "", end="", flush=True)

```

## Summary

- The **fusion** virtual model in FreeLLMAPI enables multi-model synthesis by orchestrating parallel requests across diverse providers.
- The default panel contains **4 models** (maximum **8**) selected via `diversifyChain` to ensure provider and family diversity.
- Execution occurs in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) via `runFusion`, which handles configuration resolution, parallel execution via `runModelCall`, and response synthesis.
- Choose between **`synthesize`** (judge model combines answers) and **`best_of`** (longest answer wins) strategies.
- Responses include an **`_fusion`** metadata field and optional **`x_fusion`** header for debugging when `expose_panel` is enabled.

## Frequently Asked Questions

### What is the maximum number of models in a fusion panel?

The fusion panel supports a default size of **4 models** with an absolute maximum of **8 models**. This limit is enforced in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) (lines 40–64) to balance synthesis quality against latency and token costs.

### How does the fusion model handle tool calls?

If any panel member returns a response containing tool calls, the fusion system immediately returns that response and skips the judge synthesis step. This rule exists because tool calls must remain atomic and unmodified. The logic is implemented in the `runFusion` function within [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts).

### Can I see which models contributed to a fusion response?

Yes. Every fusion response includes an **`_fusion`** field in the JSON payload that lists the panel members and the judge model used (if applicable). Additionally, setting **`fusion.expose_panel: true`** in your request adds an `x_fusion` header with detailed debugging information about the execution pipeline.

### What is the difference between synthesize and best_of strategies?

The **`synthesize`** strategy (default) sends all panel responses to a judge model that combines them into a single coherent answer. The **`best_of`** strategy skips synthesis and returns the longest panel answer directly. Use `best_of` when you want faster responses or when answers are expected to be structurally similar, and `synthesize` when you need coherent integration of diverse perspectives.