# How the Thompson-Sampling Bandit Works for Model Selection in FreeLLMAPI

> Discover how FreeLLMAPI's Thompson-sampling bandit dynamically selects the best LLM for your requests by blending reliability, speed, and intelligence scores.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-08-30

---

**FreeLLMAPI uses a Thompson-sampling bandit that draws reliability samples from Beta distributions and blends them with deterministic speed and intelligence scores to dynamically select the optimal LLM for each request.**

Thompson sampling is a powerful reinforcement learning technique for balancing exploration and exploitation. In FreeLLMAPI, this approach solves the classic multi-armed bandit problem: which LLM model should serve a given request when multiple options exist with unknown performance characteristics? The implementation centers on two core services—[`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts) for probability calculations and [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts) for orchestration.

## Beta Posterior: Modeling Reliability from Observed Outcomes

Each model's performance history is tracked as decay-weighted pseudo-counts stored in `successes` and `failures` counters. The function `reliabilityPosterior()` in [[`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L13-L23) constructs a Beta distribution from these counts, optionally incorporating a community prior:

```typescript
// From scoring.ts (lines 13-23)
interface BetaParams {
  alpha: number;   // successes + priorAlpha
  beta: number;    // failures + priorBeta
}

export function reliabilityPosterior(
  successes: number,
  failures: number,
  communityPrior?: { alpha: number; beta: number; weight: number }
): BetaParams {
  // Combines observed counts with optional shared knowledge
}

```

The **Thompson-sampling** step occurs in `sampleBeta()` ([[`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L48-L53)), which draws a random value from this Beta distribution:

```typescript
// Manual invocation of the reliability sampler for debugging
import { reliabilityPosterior, sampleBeta } from './scoring';

// Model with 12 successes and 4 failures (no community prior)
const { alpha, beta } = reliabilityPosterior(12, 4);
const sampledReliability = sampleBeta(alpha, beta);
console.log(`Sampled reliability (Thompson): ${sampledReliability.toFixed(3)}`);

```

Models with higher uncertainty—those with fewer observations—produce wider Beta distributions, giving them greater probability of generating a high sample. This automatically allocates exploration to under-tested models without explicit exploration rules.

## Deterministic Scoring Dimensions: Speed and Intelligence

The bandit augments probabilistic reliability with two fixed signals defined in [[`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L34-L69):

- **Speed Score**: Derived from throughput metrics and time-to-first-byte measurements. Faster models receive higher values regardless of reliability uncertainty.
- **Intelligence Score**: Computed from catalog tier and model rank (`intelligenceComposite`), capturing capability estimates from benchmark data.

These three axes—reliability, speed, and intelligence—form the complete evaluation space for model selection.

## Convex Combination via Bandit Presets

The file [[`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L44-L53) defines `BANDIT_PRESETS` that control how each dimension contributes to the final score:

| Preset | Reliability | Speed | Intelligence | Use Case |
|--------|-------------|-------|--------------|----------|
| `balanced` | 0.50 | 0.30 | 0.20 | General-purpose routing |
| `smartest` | 0.30 | 0.20 | 0.50 | Complex reasoning tasks |
| `fastest` | 0.35 | 0.55 | 0.10 | Latency-sensitive applications |
| `reliable` | 0.70 | 0.20 | 0.10 | Production stability priority |

The `combineScore()` function ([[`scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/scoring.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L69-L76)) normalizes these weights and applies two guard-rails: **headroom factor** (remaining capacity) and **rate-limit factor** (current token budget availability).

Switching presets at runtime:

```typescript
// Switching to a different bandit preset (favour speed)
import { setRoutingStrategy } from './router';
setRoutingStrategy('fastest');   // reliability 0.35, speed 0.55

```

## Routing Flow: From Sampling to Model Selection

The [[`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L55-L64) service orchestrates the complete selection process through `scoreChainEntry()`:

```typescript
// From router.ts (lines 55-64)
async function scoreChainEntry(
  entry: ChainEntry,
  sampled: boolean
): Promise<ScoredChainEntry> {
  const stats = await getModelStats(entry.modelId);
  
  // Thompson sampling when sampled=true, deterministic when false
  const reliability = sampled 
    ? sampleBeta(reliabilityPosterior(stats.successes, stats.failures))
    : expectedReliability(stats);
    
  // Combine with speed, intelligence, and guard-rails
  const score = combineScore(reliability, speed, intelligence, guardRails);
}

```

The `orderChain()` function ([[`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L124-L132) sorts the fallback chain by descending bandit score. For a live request, this produces the ranked model list:

```typescript
// Route a request using the default balanced bandit strategy
import { routeRequest } from './router';

const result = await routeRequest({
  model: 'gpt-4',          // optional model hint
  max_tokens: 1024,
  temperature: 0.7,
  sampled: true,          // triggers Thompson sampling
});
console.log(`Chosen model: ${result.modelId} (score=${result.score.toFixed(3)})`);

```

## Exploration Guarantees and Anti-Starvation

The implementation includes explicit safeguards against model starvation in [[`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L87-L95):

- **`EXPLORE_CHANCE` (10%)**: Randomly overrides the bandit to try a low-confidence model
- **`EXPLORE_MIN_SAMPLES` (5)**: Forces exploration for models with fewer than 5 observations

These constants ensure that newly added models receive sufficient traffic to establish reliable statistics before the Thompson-sampling mechanism takes full control.

## Community Prior: Shared Knowledge Acceleration

An optional opt-in feature ([[`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L24-L34)) allows deployment operators to seed new models with up to 50 pseudo-samples from aggregate community data. This `communityPrior` shifts the Beta distribution toward historically observed performance patterns, reducing cold-start time without sacrificing the bandit's adaptability to local conditions.

## Summary

- **Thompson-sampling bandit** in FreeLLMAPI draws reliability samples from Beta distributions built on decay-weighted success/failure counts
- **Exploration emerges naturally** from uncertainty in the posterior—models with fewer observations have wider distributions and higher chance of selection
- **Three-dimensional scoring** combines sampled reliability with deterministic speed and intelligence metrics via configurable convex weights
- **Guard-rails** (headroom, rate limits) and **explicit exploration constants** prevent starvation and respect operational constraints
- **Community priors** optionally accelerate cold-start for new models without compromising local adaptability

## Frequently Asked Questions

### What makes Thompson sampling superior to epsilon-greedy for LLM selection?

Thompson sampling allocates exploration **proportionally to uncertainty** rather than randomly. In FreeLLMAPI's implementation, a model with 1 success and 0 failures has high variance in its Beta distribution and thus substantial probability of generating a top sample. Epsilon-greedy would explore randomly, potentially wasting trials on already-proven poor performers. The Bayesian approach also enables natural incorporation of the community prior for faster convergence.

### How does the decay weighting affect long-term reliability estimates?

The pseudo-counts (`successes` and `failures`) use exponential decay, meaning older observations contribute less to the Beta parameters. This allows FreeLLMAPI to adapt to model degradation or improvement over time without manual intervention. The decay rate is configurable per deployment, balancing responsiveness against noise sensitivity.

### Can the bandit be disabled for deterministic routing?

Yes. The `sampled` parameter controls this behavior. When `false`, `expectedReliability()` returns the mean of the Beta distribution (α / (α + β)) rather than a random sample. The dashboard view uses this deterministic mode, while live requests typically use `sampled: true`. Setting the strategy to `'priority'` additionally bypasses scoring entirely, using manual priority values instead.

### What happens when all models hit rate limits?

The `rateLimitFactor` guard-rail multiplies the bandit score by a value between 0 and 1 based on remaining token capacity. When all models approach limits, this factor shrinks scores uniformly, but the relative ranking persists. If hard limits are exceeded, [[`rate-limit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/rate-limit.ts)](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) triggers graceful degradation through the fallback chain ordered by `orderChain()`.