FreeLLMAPI Router Model Selection Algorithm: Thompson Sampling Explained

The FreeLLMAPI Router uses a Thompson sampling-based multi-armed bandit algorithm that scores candidate models by combining Bayesian reliability estimates, speed metrics, and intelligence rankings, then dynamically penalizes recent failures and respects quota head-room to select the optimal model for each request.

The FreeLLMAPI Router is the intelligent request dispatcher in the tashfeenahmed/freellmapi open-source project. Unlike simple round-robin or random selection, the router treats every enabled language model as an arm in a multi-armed bandit problem, continuously learning from real-world traffic to maximize successful completions while respecting free-tier rate limits.

Core Components of the Bandit Algorithm

The router's model selection logic relies on four interconnected systems that evaluate, penalize, and rank candidate models before each request.

Thompson Sampling and Bayesian Reliability

At the heart of the FreeLLMAPI Router algorithm is Thompson sampling implemented via Beta distributions. For every candidate model, the system maintains a Bayesian posterior over its reliability using historic success and failure counts.

In server/src/services/scoring.ts, the reliabilityPosterior function (lines 21-23) constructs the Beta parameters, while sampleBeta (lines 448-452) performs the actual sampling:

// From scoring.ts - generating a sampled reliability estimate
const reliability = sampled
    ? sampleBeta(alpha, beta)                // Thompson sampling step
    : expectedReliability(successes, failures, community);

When live routing occurs with sampled = true, the router draws a random value from each model's Beta posterior. This injects statistically-sound exploration: models with limited data have high variance and occasionally win selection, ensuring the system discovers new high-performing models while exploiting proven ones.

Speed and Intelligence Weighting

Beyond reliability, the algorithm evaluates two additional quality axes. The speed score derives from observed token-per-second throughput and time-to-first-byte latency via speedScore. The intelligence score is a static composite based on model size labels and intelligence rankings via intelligenceScore.

The combineScore function (scoring.ts lines 40-45) multiplies these three axes—reliability, speed, and intelligence—by a configurable weight vector:

const score = combineScore(
    { reliability, speed, intelligence, headroom, rateLimit }, 
    weights
);

Users can select preset weight configurations (balanced, fastest, reliable, smartest) or provide custom vectors through the routing strategy API.

Dynamic Penalty System

After each request, the router updates per-model penalty counters to reflect recent failures. The system differentiates between error types in server/src/services/router.ts (lines 86-106):

  • Rate-limit hits (429 errors): Trigger recordRateLimitHit and apply heavier penalties
  • Upstream failures: Trigger recordModelFailure with lighter penalties

These penalties decay automatically every two minutes, allowing models to recover once their error rates improve. During scoring, the penalty feeds into a rateLimitFactor that multiplicatively reduces the model's final selection probability.

Quota Guards and Head-Room Checks

Before a model can win selection, the FreeLLMAPI Router enforces two quota-based guard-rails calculated in server/src/services/router.ts (lines 979-991):

  1. Monthly head-room: The fraction of free-tier token budget remaining (headroomFactor)
  2. Window head-room: The remaining capacity within current RPM/RPD/TPM/TPD limits (rateWindowHeadroomFactor)

The router uses Math.min of these two values, ensuring models approaching any quota ceiling are demoted regardless of their quality scores. This prevents hard failures and optimizes for successful request completion within provider constraints.

Chain Ordering and Final Selection

The orderChain function (router.ts lines 1025-1064) implements the final selection logic. For strategies other than priority, the router uses Thompson-sampled scores with a small EXPLORE_CHANCE floor to guarantee exploration of rarely-used models.

// Candidate ordering with exploration guarantee
const scored = chain.map(row => ({
    modelId: row.model_id,
    score: scoreChainEntry(row, weights, 0, 10, /*sampled=*/true).score,
}));
scored.sort((a, b) => b.score - a.score);

The request handler consumes this ordered list and picks the first model satisfying token requirements, provider capabilities, and cooldown constraints.

Practical Implementation Examples

Basic Request Routing

To route a chat request through the bandit algorithm:

import { routeRequest } from './services/router.js';

const payload = {
  model: 'gpt-4o',
  max_tokens: 2000,
  messages: [{ role: 'user', content: 'Hello!' }],
};

// Router returns resolved provider, API key, and model ID
const route = await routeRequest(payload, {
  strategy: 'balanced',
  explore: true
});

await route.provider.sendChat(route.apiKey, {
  model: route.modelId,
  ...payload,
});

Inspecting Model Scores Deterministically

For dashboards or debugging, compute deterministic (non-sampled) scores:

import { scoreChainEntry, getActiveRoutingWeights } from './services/router.js';

const chain = await getEnabledChain();
const weights = getActiveRoutingWeights().weights;

const scored = chain.map(row => ({
  modelId: row.model_id,
  score: scoreChainEntry(row, weights, 0, 10, /*sampled=*/false).score,
}));

scored.sort((a, b) => b.score - a.score);
console.log('Top models by bandit score:', scored.slice(0, 5));

Changing Routing Strategies

Switch between preset weight configurations at runtime:

import { setRoutingStrategy } from './services/router.js';

// Options: 'balanced', 'fastest', 'reliable', 'smartest', 'custom', 'priority'
setRoutingStrategy('fastest');

Summary

  • The FreeLLMAPI Router implements a Thompson sampling-based multi-armed bandit where each model represents an arm with unknown payoff probability.
  • Bayesian reliability scoring in scoring.ts uses Beta posteriors sampled via sampleBeta to balance exploration of new models against exploitation of proven ones.
  • The algorithm combines reliability with speed and intelligence metrics using configurable weight vectors managed by combineScore.
  • Dynamic penalties for 429 errors and upstream failures automatically demote unstable models, decaying over two-minute windows.
  • Quota head-room checks prevent selection of models nearing rate limits or monthly budgets, ensuring higher request success rates.
  • Final selection occurs in orderChain (router.ts lines 1025-1064), which respects manual priority lists or returns Thompson-sampled rankings with guaranteed exploration floors.

Frequently Asked Questions

What makes Thompson sampling better than random selection for LLM routing?

Thompson sampling provides a mathematically optimal solution to the exploration-exploitation tradeoff inherent in multi-armed bandit problems. Unlike random selection, it naturally favors models with higher historical success rates while still exploring uncertain candidates proportional to their probability of being optimal. According to the FreeLLMAPI source code in scoring.ts, this Bayesian approach adapts faster to provider outages and performance changes than epsilon-greedy or purely random strategies.

How does the router handle sudden provider outages or rate limits?

When a model returns a 429 error or upstream failure, the recordRateLimitHit or recordModelFailure functions (router.ts lines 86-106) immediately increment penalty counters. These penalties factor into the rateLimitFactor during the next scoring cycle, effectively removing the problematic model from contention for several minutes. Because penalties decay every two minutes, the system automatically recovers once the provider stabilizes, without manual intervention.

Can I customize which factors matter most for my use case?

Yes. The router supports preset strategies including balanced, fastest, reliable, and smartest, each providing different weight vectors to combineScore. Advanced users can inject custom weight vectors through the routing configuration or modify model-weight-overrides.ts to adjust per-model multipliers. This flexibility allows latency-sensitive applications to prioritize speed scores while reliability-critical workloads can weight Bayesian posterior samples more heavily.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →