How the Thompson-Sampling Bandit Works for Model Selection in FreeLLMAPI
FreeLLMAPI uses a Thompson-sampling bandit that draws reliability samples from Beta distributions and blends them with deterministic speed and intelligence scores to dynamically select the optimal LLM for each request.
Thompson sampling is a powerful reinforcement learning technique for balancing exploration and exploitation. In FreeLLMAPI, this approach solves the classic multi-armed bandit problem: which LLM model should serve a given request when multiple options exist with unknown performance characteristics? The implementation centers on two core services—scoring.ts for probability calculations and router.ts for orchestration.
Beta Posterior: Modeling Reliability from Observed Outcomes
Each model's performance history is tracked as decay-weighted pseudo-counts stored in successes and failures counters. The function reliabilityPosterior() in [server/src/services/scoring.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L13-L23) constructs a Beta distribution from these counts, optionally incorporating a community prior:
// From scoring.ts (lines 13-23)
interface BetaParams {
alpha: number; // successes + priorAlpha
beta: number; // failures + priorBeta
}
export function reliabilityPosterior(
successes: number,
failures: number,
communityPrior?: { alpha: number; beta: number; weight: number }
): BetaParams {
// Combines observed counts with optional shared knowledge
}
The Thompson-sampling step occurs in sampleBeta() ([scoring.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L48-L53)), which draws a random value from this Beta distribution:
// Manual invocation of the reliability sampler for debugging
import { reliabilityPosterior, sampleBeta } from './scoring';
// Model with 12 successes and 4 failures (no community prior)
const { alpha, beta } = reliabilityPosterior(12, 4);
const sampledReliability = sampleBeta(alpha, beta);
console.log(`Sampled reliability (Thompson): ${sampledReliability.toFixed(3)}`);
Models with higher uncertainty—those with fewer observations—produce wider Beta distributions, giving them greater probability of generating a high sample. This automatically allocates exploration to under-tested models without explicit exploration rules.
Deterministic Scoring Dimensions: Speed and Intelligence
The bandit augments probabilistic reliability with two fixed signals defined in [scoring.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L34-L69):
- Speed Score: Derived from throughput metrics and time-to-first-byte measurements. Faster models receive higher values regardless of reliability uncertainty.
- Intelligence Score: Computed from catalog tier and model rank (
intelligenceComposite), capturing capability estimates from benchmark data.
These three axes—reliability, speed, and intelligence—form the complete evaluation space for model selection.
Convex Combination via Bandit Presets
The file [scoring.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L44-L53) defines BANDIT_PRESETS that control how each dimension contributes to the final score:
| Preset | Reliability | Speed | Intelligence | Use Case |
|---|---|---|---|---|
balanced |
0.50 | 0.30 | 0.20 | General-purpose routing |
smartest |
0.30 | 0.20 | 0.50 | Complex reasoning tasks |
fastest |
0.35 | 0.55 | 0.10 | Latency-sensitive applications |
reliable |
0.70 | 0.20 | 0.10 | Production stability priority |
The combineScore() function ([scoring.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts#L69-L76)) normalizes these weights and applies two guard-rails: headroom factor (remaining capacity) and rate-limit factor (current token budget availability).
Switching presets at runtime:
// Switching to a different bandit preset (favour speed)
import { setRoutingStrategy } from './router';
setRoutingStrategy('fastest'); // reliability 0.35, speed 0.55
Routing Flow: From Sampling to Model Selection
The [router.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L55-L64) service orchestrates the complete selection process through scoreChainEntry():
// From router.ts (lines 55-64)
async function scoreChainEntry(
entry: ChainEntry,
sampled: boolean
): Promise<ScoredChainEntry> {
const stats = await getModelStats(entry.modelId);
// Thompson sampling when sampled=true, deterministic when false
const reliability = sampled
? sampleBeta(reliabilityPosterior(stats.successes, stats.failures))
: expectedReliability(stats);
// Combine with speed, intelligence, and guard-rails
const score = combineScore(reliability, speed, intelligence, guardRails);
}
The orderChain() function ([router.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L124-L132) sorts the fallback chain by descending bandit score. For a live request, this produces the ranked model list:
// Route a request using the default balanced bandit strategy
import { routeRequest } from './router';
const result = await routeRequest({
model: 'gpt-4', // optional model hint
max_tokens: 1024,
temperature: 0.7,
sampled: true, // triggers Thompson sampling
});
console.log(`Chosen model: ${result.modelId} (score=${result.score.toFixed(3)})`);
Exploration Guarantees and Anti-Starvation
The implementation includes explicit safeguards against model starvation in [router.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L87-L95):
EXPLORE_CHANCE(10%): Randomly overrides the bandit to try a low-confidence modelEXPLORE_MIN_SAMPLES(5): Forces exploration for models with fewer than 5 observations
These constants ensure that newly added models receive sufficient traffic to establish reliable statistics before the Thompson-sampling mechanism takes full control.
Community Prior: Shared Knowledge Acceleration
An optional opt-in feature ([router.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L24-L34)) allows deployment operators to seed new models with up to 50 pseudo-samples from aggregate community data. This communityPrior shifts the Beta distribution toward historically observed performance patterns, reducing cold-start time without sacrificing the bandit's adaptability to local conditions.
Summary
- Thompson-sampling bandit in FreeLLMAPI draws reliability samples from Beta distributions built on decay-weighted success/failure counts
- Exploration emerges naturally from uncertainty in the posterior—models with fewer observations have wider distributions and higher chance of selection
- Three-dimensional scoring combines sampled reliability with deterministic speed and intelligence metrics via configurable convex weights
- Guard-rails (headroom, rate limits) and explicit exploration constants prevent starvation and respect operational constraints
- Community priors optionally accelerate cold-start for new models without compromising local adaptability
Frequently Asked Questions
What makes Thompson sampling superior to epsilon-greedy for LLM selection?
Thompson sampling allocates exploration proportionally to uncertainty rather than randomly. In FreeLLMAPI's implementation, a model with 1 success and 0 failures has high variance in its Beta distribution and thus substantial probability of generating a top sample. Epsilon-greedy would explore randomly, potentially wasting trials on already-proven poor performers. The Bayesian approach also enables natural incorporation of the community prior for faster convergence.
How does the decay weighting affect long-term reliability estimates?
The pseudo-counts (successes and failures) use exponential decay, meaning older observations contribute less to the Beta parameters. This allows FreeLLMAPI to adapt to model degradation or improvement over time without manual intervention. The decay rate is configurable per deployment, balancing responsiveness against noise sensitivity.
Can the bandit be disabled for deterministic routing?
Yes. The sampled parameter controls this behavior. When false, expectedReliability() returns the mean of the Beta distribution (α / (α + β)) rather than a random sample. The dashboard view uses this deterministic mode, while live requests typically use sampled: true. Setting the strategy to 'priority' additionally bypasses scoring entirely, using manual priority values instead.
What happens when all models hit rate limits?
The rateLimitFactor guard-rail multiplies the bandit score by a value between 0 and 1 based on remaining token capacity. When all models approach limits, this factor shrinks scores uniformly, but the relative ranking persists. If hard limits are exceeded, [rate-limit.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts) triggers graceful degradation through the fallback chain ordered by orderChain().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →