# How the FreeLLMAPI Router Selects Models: Contextual Bandit Scoring Explained

> Discover how the FreeLLMAPI router uses contextual bandit scoring to select the best AI model for your request based on reliability, speed, intelligence, and more. Automatic fallback included.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-09-04

---

**FreeLLMAPI employs a contextual bandit router that evaluates every candidate "(platform, model, key)" pair across five normalized axes—reliability, speed, intelligence, headroom, and rate-limit—to select the optimal model for each request, automatically falling back to the next highest-scoring candidate if the first fails.**

FreeLLMAPI is an open-source unified API gateway that aggregates multiple free-tier LLM providers under a single endpoint. Unlike simple round-robin load balancers, the FreeLLMAPI router selects models using a sophisticated multi-armed bandit algorithm implemented in `tashfeenahmed/freellmapi` that continuously learns from live traffic patterns and provider health.

## The Seven-Step Model Selection Pipeline

The routing logic resides primarily in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) and follows a deterministic seven-step process to ensure optimal provider selection.

### Step 1: Build the Fallback Chain

The router first resolves the active **fallback chain**—whether a named profile (e.g., `auto:coding`), a global fallback, or built-in aliases like `auto:smart`. Each chain row contains static metadata including model size, tier, rate limits, and capabilities.

In [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) (lines 22-31), the router constructs this chain by mapping the requested model identifier to its configured provider entries.

### Step 2: Gather Live Statistics

The system maintains decay-weighted counts of successes, failures, token throughput, and time-to-first-byte (TTFB) latency per model and per API key in a `statsCache`. Optional community priors are folded into these posteriors to improve cold-start performance.

This statistical aggregation occurs in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) (lines 78-100), where historical performance data is retrieved and weighted by recency.

### Step 3: Compute the Five Axes

For each candidate, the router calculates five normalized scores in [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts):

- **Reliability**: Thompson-sampled Beta posterior via `sampleBeta` (or expected value for dashboard views)
- **Speed**: Blended throughput and TTFB score via `speedScore`
- **Intelligence**: Tier-first composite rank via `intelligenceScore`
- **Headroom**: Quota guardrail based on monthly and window utilization
- **Rate-limit**: Penalty derived from recent 429 errors or failure spikes

The `scoreChainEntry()` function in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) orchestrates these calculations.

### Step 4: Apply Weighting Strategy

Operators select a preset strategy—`balanced`, `smartest`, `fastest`, `reliable`, `custom`, or `priority`—or define custom weights. The router multiplies the three core axes (reliability, speed, intelligence) by the chosen weight vector, then applies the headroom and rate-limit guardrails as multiplicative penalties.

Weight vectors are defined in [`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts) as `BANDIT_PRESETS` and fetched in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) via the `weightsFor(strategy)` helper (lines 14-16).

### Step 5: Apply Per-Model Overrides

Administrators can boost or demote specific models using the `MODEL_ROUTING_OVERRIDES` environment variable. The `applyModelWeightOverride()` function in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) (lines 30-34) scales the final score by the configured weight factor.

### Step 6: Key Selection

When a model has multiple API keys configured, the router selects the optimal key using either:

- **auto**: Selects the key with the highest per-key Thompson score
- **least-remaining**: Selects the key with the most remaining quota

This logic is implemented in `getKeySelectionStrategy()` within [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) (lines 86-108).

### Step 7: Final Selection and Fallback

The `orderChain()` function sorts the entire chain by match tier, composite score, and priority. The router attempts the top entry; if it encounters a 429, 5xx error, or timeout, the system retries with the next highest-scoring candidate, up to 20 attempts.

The fallback retry loop is handled in [`fallback-loop.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/fallback-loop.ts).

## Practical Implementation Examples

### Automatic Model Selection

Use the `auto` model identifier to let the router intelligently select the best available free model:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="YOUR_UNIFIED_KEY",
)

resp = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Explain quantum tunnelling in one paragraph."}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))

```

### Named Profile Routing

For coding-specific workloads, use a named fallback chain:

```bash
curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer YOUR_UNIFIED_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto:coding",
        "messages":[{"role":"user","content":"Write a quicksort in Rust."}]
      }' | jq .

```

### Runtime Model Weight Overrides

Programmatically adjust model priorities via environment variables:

```javascript
process.env.MODEL_ROUTING_OVERRIDES = JSON.stringify({
  "gpt-4o": {"weight": 0.5},
  "llama-3.3-70b": {"weight": 1.5}
});

const { OpenAI } = require("openai");
const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",
  apiKey: "YOUR_UNIFIED_KEY",
});

(async () => {
  const chat = await client.chat.completions.create({
    model: "auto",
    messages: [{ role: "user", content: "Summarise the plot of *The Matrix*." }],
  });
  console.log(chat.choices[0].message.content);
})();

```

## Summary

- **FreeLLMAPI uses a contextual bandit algorithm**, not round-robin, to route requests based on live performance data.
- **Five axes determine scores**: reliability (Thompson sampling), speed (throughput + TTFB), intelligence (tier ranking), headroom (quota utilization), and rate-limit penalties.
- **Selection happens in seven steps**: chain building, statistics gathering, axis scoring, weight application, overrides, key selection, and final ordering with fallback.
- **Configuration is flexible**: Operators choose from presets like `smartest` or `fastest`, or define custom weight vectors.
- **Automatic recovery**: Failed requests trigger fallback to the next highest-scoring model, up to 20 retries.

## Frequently Asked Questions

### How does FreeLLMAPI handle new models with no historical data?

For models without sufficient historical data, the router folds optional community priors into the posterior distribution and uses conservative default values in the Beta distribution parameters. This cold-start handling ensures new models receive traffic proportional to their theoretical capabilities until live statistics stabilize in `statsCache`.

### Can I force specific models to always be selected first?

Yes. Set the `MODEL_ROUTING_OVERRIDES` environment variable to boost specific model weights above 1.0, or create a named fallback chain in your configuration that lists preferred models in priority order. The `applyModelWeightOverride()` function in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) applies these multipliers after all other scoring completes.

### What happens if all models in the fallback chain fail?

If all candidates exhaust their 20-attempt fallback limit, the router returns the final error response to the client. However, the decay-weighted statistics in `statsCache` ensure that temporarily failing models are rapidly penalized and removed from consideration until they demonstrate recovery, preventing cascading failures.

### How does the router prevent exceeding free-tier quotas?

The **Headroom** axis in [`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts) tracks monthly and window-based utilization for each API key. When utilization approaches configured thresholds, the guardrail multiplies the model's score by a fraction approaching zero, effectively removing it from selection until the quota resets or the window slides.