How the FreeLLMAPI Router Selects Models: Contextual Bandit Scoring Explained

FreeLLMAPI employs a contextual bandit router that evaluates every candidate "(platform, model, key)" pair across five normalized axes—reliability, speed, intelligence, headroom, and rate-limit—to select the optimal model for each request, automatically falling back to the next highest-scoring candidate if the first fails.

FreeLLMAPI is an open-source unified API gateway that aggregates multiple free-tier LLM providers under a single endpoint. Unlike simple round-robin load balancers, the FreeLLMAPI router selects models using a sophisticated multi-armed bandit algorithm implemented in tashfeenahmed/freellmapi that continuously learns from live traffic patterns and provider health.

The Seven-Step Model Selection Pipeline

The routing logic resides primarily in server/src/services/router.ts and follows a deterministic seven-step process to ensure optimal provider selection.

Step 1: Build the Fallback Chain

The router first resolves the active fallback chain—whether a named profile (e.g., auto:coding), a global fallback, or built-in aliases like auto:smart. Each chain row contains static metadata including model size, tier, rate limits, and capabilities.

In server/src/services/router.ts (lines 22-31), the router constructs this chain by mapping the requested model identifier to its configured provider entries.

Step 2: Gather Live Statistics

The system maintains decay-weighted counts of successes, failures, token throughput, and time-to-first-byte (TTFB) latency per model and per API key in a statsCache. Optional community priors are folded into these posteriors to improve cold-start performance.

This statistical aggregation occurs in router.ts (lines 78-100), where historical performance data is retrieved and weighted by recency.

Step 3: Compute the Five Axes

For each candidate, the router calculates five normalized scores in server/src/services/scoring.ts:

  • Reliability: Thompson-sampled Beta posterior via sampleBeta (or expected value for dashboard views)
  • Speed: Blended throughput and TTFB score via speedScore
  • Intelligence: Tier-first composite rank via intelligenceScore
  • Headroom: Quota guardrail based on monthly and window utilization
  • Rate-limit: Penalty derived from recent 429 errors or failure spikes

The scoreChainEntry() function in router.ts orchestrates these calculations.

Step 4: Apply Weighting Strategy

Operators select a preset strategy—balanced, smartest, fastest, reliable, custom, or priority—or define custom weights. The router multiplies the three core axes (reliability, speed, intelligence) by the chosen weight vector, then applies the headroom and rate-limit guardrails as multiplicative penalties.

Weight vectors are defined in scoring.ts as BANDIT_PRESETS and fetched in router.ts via the weightsFor(strategy) helper (lines 14-16).

Step 5: Apply Per-Model Overrides

Administrators can boost or demote specific models using the MODEL_ROUTING_OVERRIDES environment variable. The applyModelWeightOverride() function in router.ts (lines 30-34) scales the final score by the configured weight factor.

Step 6: Key Selection

When a model has multiple API keys configured, the router selects the optimal key using either:

  • auto: Selects the key with the highest per-key Thompson score
  • least-remaining: Selects the key with the most remaining quota

This logic is implemented in getKeySelectionStrategy() within router.ts (lines 86-108).

Step 7: Final Selection and Fallback

The orderChain() function sorts the entire chain by match tier, composite score, and priority. The router attempts the top entry; if it encounters a 429, 5xx error, or timeout, the system retries with the next highest-scoring candidate, up to 20 attempts.

The fallback retry loop is handled in fallback-loop.ts.

Practical Implementation Examples

Automatic Model Selection

Use the auto model identifier to let the router intelligently select the best available free model:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="YOUR_UNIFIED_KEY",
)

resp = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Explain quantum tunnelling in one paragraph."}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))

Named Profile Routing

For coding-specific workloads, use a named fallback chain:

curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer YOUR_UNIFIED_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model":"auto:coding",
        "messages":[{"role":"user","content":"Write a quicksort in Rust."}]
      }' | jq .

Runtime Model Weight Overrides

Programmatically adjust model priorities via environment variables:

process.env.MODEL_ROUTING_OVERRIDES = JSON.stringify({
  "gpt-4o": {"weight": 0.5},
  "llama-3.3-70b": {"weight": 1.5}
});

const { OpenAI } = require("openai");
const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",
  apiKey: "YOUR_UNIFIED_KEY",
});

(async () => {
  const chat = await client.chat.completions.create({
    model: "auto",
    messages: [{ role: "user", content: "Summarise the plot of *The Matrix*." }],
  });
  console.log(chat.choices[0].message.content);
})();

Summary

  • FreeLLMAPI uses a contextual bandit algorithm, not round-robin, to route requests based on live performance data.
  • Five axes determine scores: reliability (Thompson sampling), speed (throughput + TTFB), intelligence (tier ranking), headroom (quota utilization), and rate-limit penalties.
  • Selection happens in seven steps: chain building, statistics gathering, axis scoring, weight application, overrides, key selection, and final ordering with fallback.
  • Configuration is flexible: Operators choose from presets like smartest or fastest, or define custom weight vectors.
  • Automatic recovery: Failed requests trigger fallback to the next highest-scoring model, up to 20 retries.

Frequently Asked Questions

How does FreeLLMAPI handle new models with no historical data?

For models without sufficient historical data, the router folds optional community priors into the posterior distribution and uses conservative default values in the Beta distribution parameters. This cold-start handling ensures new models receive traffic proportional to their theoretical capabilities until live statistics stabilize in statsCache.

Can I force specific models to always be selected first?

Yes. Set the MODEL_ROUTING_OVERRIDES environment variable to boost specific model weights above 1.0, or create a named fallback chain in your configuration that lists preferred models in priority order. The applyModelWeightOverride() function in router.ts applies these multipliers after all other scoring completes.

What happens if all models in the fallback chain fail?

If all candidates exhaust their 20-attempt fallback limit, the router returns the final error response to the client. However, the decay-weighted statistics in statsCache ensure that temporarily failing models are rapidly penalized and removed from consideration until they demonstrate recovery, preventing cascading failures.

How does the router prevent exceeding free-tier quotas?

The Headroom axis in scoring.ts tracks monthly and window-based utilization for each API key. When utilization approaches configured thresholds, the guardrail multiplies the model's score by a fraction approaching zero, effectively removing it from selection until the quota resets or the window slides.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →