How FreeLLMAPI Uses the Thompson-Sampling Bandit Algorithm for Intelligent Model Routing

FreeLLMAPI treats model selection as a contextual bandit problem, using Thompson sampling on a Beta distribution to balance exploration of uncertain models against exploitation of proven performers, while combining this reliability score with speed and intelligence metrics to route each request optimally.

FreeLLMAPI is an open-source LLM routing layer that dynamically selects the best (platform, model, key) tuple for every incoming request. At its core, the system implements a Thompson-sampling bandit algorithm to handle the exploration-exploitation trade-off, ensuring high availability while continuously discovering optimal model configurations. The algorithm is primarily implemented in server/src/services/scoring.ts and orchestrated through server/src/services/router.ts.

Core Architecture of the Bandit Router

The router evaluates every candidate model using a convex combination of five normalized axes. The reliability axis is the probabilistic component where Thompson sampling operates, while the remaining axes provide deterministic signals.

The Five-Axis Scoring Framework

Each candidate receives a base score calculated from weighted components:


base = w_rel·reliability + w_spd·speed + w_int·intelligence
effective = base × headroomFactor × rateLimitFactor

The weights (w_rel, w_spd, w_int) derive from selectable bandit presets defined in server/src/services/scoring.ts (lines 44-53). Available presets include balanced, smartest, fastest, and reliable, allowing operators to prioritize different routing strategies without changing the underlying Thompson-sampling logic.

Beta Posterior and Decay-Weighted Updates

The reliability score follows a Bayesian updating process. Each model accumulates decay-weighted successes and failures over a rolling 7-day window. The system maintains two count arrays per model:

  • Alpha (α) = successes + community_successes + 1
  • Beta (β) = failures + community_failures + 1

These parameters define the Beta distribution posterior, implemented in server/src/services/scoring.ts (lines 73-82). The addition of 1 to each parameter creates a uniform prior, while optional community-wide counts provide shared prior knowledge across similar models.

How Thompson Sampling Works in Practice

Thompson sampling solves the contextual bandit by maintaining a probability distribution over the expected reward (reliability) of each model and sampling from these distributions to make decisions.

Sampling from the Beta Distribution

During live routing, the system draws a random sample from Beta(α, β) using the sampleBeta function (scoring.ts, lines 88-94). This sample serves as the instantaneous reliability value in the convex combination:

import { reliabilityPosterior, sampleBeta } from './scoring';

// Decay-weighted counts for a specific model/key
const successes = 12;
const failures = 3;
const community = { successes: 5, failures: 2 };

const { alpha, beta } = reliabilityPosterior(successes, failures, community);
const reliabilitySample = sampleBeta(alpha, beta);  // Thompson sample

The variance of the Beta distribution naturally decreases as sample counts increase. Models with few observations maintain high variance, causing the sampler to occasionally select them for exploration. High-performing models with thousands of observations have tight distributions around their true mean, ensuring exploitation of proven reliability.

Per-Key Bandit Ordering

When a single model has multiple API keys, FreeLLMAPI extends the bandit to the key level. Each key maintains independent success and failure counts, and the router samples a separate Thompson score per key (sampleBeta on key-level counts) to determine the optimal key for that specific request (router.ts, lines 388-415). This prevents a single failing key from degrading the perceived reliability of an otherwise healthy model.

Dynamic Adjustments and Guard Rails

The bandit algorithm operates within a framework of guard rails that respect operational constraints and traffic patterns.

Peak-Hours Reliability Boosting

During operator-defined peak windows, the router transfers a portion of the speed weight to the reliability weight (scoring.ts, lines 89-95). This peak-hours adjustment reinforces the Thompson-sampling effect during high-traffic periods, prioritizing stability over latency when capacity is constrained.

Task-Type Biasing

The system applies small weight adjustments based on request classification. For code versus chat tasks, the router shifts marginal weight between axes (scoring.ts, lines 22-28). However, the underlying Thompson-sampled reliability remains the stochastic anchor of the scoring function, ensuring that exploration continues regardless of task type.

Configuration and Presets

Operators can inspect and modify bandit behavior through the preset configuration:

import { DEFAULT_STRATEGY, BANDIT_PRESETS } from './scoring';

const strategy = DEFAULT_STRATEGY;  // 'balanced' default
const weights = BANDIT_PRESETS[strategy];
// Returns: { reliability: 0.5, speed: 0.25, intelligence: 0.25 }

While the dashboard displays the deterministic expectation α/(α+β) for stable sorting (scoring.ts, lines 84-92), the live router always uses the stochastic sample for request distribution.

Summary

  • Bayesian foundation: FreeLLMAPI models reliability as a Beta distribution with decay-weighted success/failure counts over a 7-day window.
  • Thompson sampling: The router draws random samples from Beta(α, β) via sampleBeta in scoring.ts, automatically balancing exploration of uncertain models against exploitation of proven performers.
  • Multi-axis scoring: Reliability samples combine with deterministic speed and intelligence scores, weighted by selectable presets (balanced, fastest, smartest, reliable).
  • Hierarchical bandits: Per-key Thompson sampling in router.ts (lines 388-415) ensures optimal API key selection within each model.
  • Dynamic adaptation: Peak-hours and task-type adjustments shift weight profiles without disrupting the core probabilistic exploration mechanism.

Frequently Asked Questions

What is Thompson sampling and why does FreeLLMAPI use it?

Thompson sampling is a Bayesian algorithm for solving multi-armed bandit problems that maintains a probability distribution over the expected reward of each option and samples from these distributions to make decisions. FreeLLMAPI uses it because it provides a principled, probabilistic approach to the exploration-exploitation trade-off, naturally exploring uncertain models (high variance in the Beta distribution) while exploiting reliable ones (low variance) without requiring manual tuning of exploration parameters.

How does FreeLLMAPI prevent newly added models from being ignored?

New models start with a uniform prior (α=1, β=1) plus any community-wide prior counts, resulting in a Beta distribution with high variance. Because Thompson sampling draws random samples from these distributions, models with few observations will occasionally produce high reliability samples due to their wide confidence intervals. This automatic exploration ensures new models receive traffic proportional to their uncertainty until sufficient data tightens their posterior distributions.

Where is the core bandit logic implemented in the codebase?

The mathematical core resides in server/src/services/scoring.ts, specifically in the reliabilityPosterior function (lines 73-82) for updating Bayesian beliefs and the sampleBeta function (lines 88-94) for stochastic sampling. The orchestration logic that applies these scores to routing decisions lives in server/src/services/router.ts (lines 388-415), which handles per-key bandit ordering and the final model selection chain.

Can I disable exploration and always pick the most reliable model?

While you cannot disable Thompson sampling entirely without code modification, you can minimize exploration by selecting the reliable preset, which maximizes the weight of the reliability axis. Additionally, the system uses decay-weighted counts, so as models accumulate thousands of observations, their Beta distributions become extremely tight (low variance), causing the sampler to behave nearly deterministically. For complete determinism, you would need to modify sampleBeta in server/src/services/scoring.ts to return the mean α/(α+β) instead of a random sample.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →