# FreeLLMAPI Routing Strategies: 6 Built-In Methods for Model Selection

> Explore FreeLLMAPI's 6 routing strategies for intelligent LLM model selection. Learn about priority, balanced, smartest, fastest, reliable, and custom methods to optimize your AI requests.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-08-30

---

**FreeLLMAPI provides six routing strategies—priority, balanced, smartest, fastest, reliable, and custom—that determine how requests are distributed across a fallback chain of LLM models.**

All routing decisions in FreeLLMAPI flow through a **fallback chain** of models. The order of that chain is governed by a configurable strategy stored in the SQLite `settings` table under the key `routing_strategy`. This value is retrieved by `getRoutingStrategy()` in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) [L996-L1000](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L996-L1000), which then determines how incoming requests are matched to available models.

## The Six Routing Strategies in FreeLLMAPI

FreeLLMAPI ships with six built-in routing strategies. Five rely on **multi-armed bandit scoring** with Thompson sampling, while one uses explicit operator-defined priorities.

### Priority: Manual Operator Control

The `priority` strategy ignores bandit scores entirely. Instead, it sorts models by the static `priority` value set in the dashboard, applying only the temporary penalty from `getPenalty()` when models return HTTP 429 rate-limit errors.

This strategy is evaluated in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) [L1030-L1064](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L1030-L1064). It guarantees predictable, deterministic routing based purely on operator preference.

### Balanced: Equal Weight to All Factors

The `balanced` strategy uses the **bandit preset** that distributes weight evenly across three axes:

- **Reliability** – success rate and error frequency
- **Speed** – throughput and latency
- **Intelligence** – benchmark performance on reasoning tasks

This preset is defined in [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) as `BANDIT_PRESETS.balanced`. Scores are computed and then **Thompson-sampled** for each live request to introduce beneficial exploration.

### Smartest: Intelligence-First Selection

The `smartest` strategy elevates the **intelligence** axis while still considering reliability and speed. It uses the `smartest` preset from `BANDIT_PRESETS`, assigning higher weight to benchmark-derived intelligence scores.

Ideal for applications where response quality outweighs latency or occasional failures.

### Fastest: Latency-Optimized Routing

The `fastest` strategy prioritizes **speed** (throughput and latency) over other concerns. It applies the `fastest` preset from `BANDIT_PRESETS` with elevated speed weights.

Use this for real-time applications like chat interfaces where low latency is critical.

### Reliable: Uptime-Focused Fallback

The `reliable` strategy favors **reliability** (low failure rate) above speed or intelligence. It pulls the `reliable` preset from `BANDIT_PRESETS`, weighting historical success rates heavily.

Best for production workloads where consistent availability matters more than peak performance.

### Custom: Operator-Defined Weight Vectors

The `custom` strategy allows operators to supply their own weight vector across the three axes. These values are persisted under the `routing_custom_weights` setting and read by `getCustomWeights()` in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) [L889-L906](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts#L889-L906).

The vector is automatically normalized before scoring, so relative proportions matter more than absolute values.

## How Bandit-Based Strategies Work

All strategies except `priority` share the same underlying mechanism:

1. **Score calculation** – Each model's historical reliability, speed, and intelligence are combined according to the selected weight vector
2. **Thompson sampling** – A random draw (`sampleBeta`) generates a **score** for that specific request
3. **Selection** – The highest-scoring model is chosen; if rate-limited or constrained, the next-best is selected

This stochastic approach automatically balances exploitation of known-good models with exploration of potentially better alternatives.

Two orthogonal features modify this behavior:

- **Exploration toggle** (`EXPLORE_CHANCE`) – Forces random model selection for a percentage of requests
- **Peak-hour adjustments** – Dynamically shift effective weights during high-traffic periods

## Programmatic Control of Routing Strategies

### Retrieve the Current Strategy

```typescript
import { getRoutingStrategy } from '../server/src/services/router';

// Returns: "priority" | "balanced" | "smartest" | "fastest" | "reliable" | "custom"
const current = getRoutingStrategy();
console.log('Active routing strategy:', current);

```

### Change the Active Strategy

```typescript
import { setRoutingStrategy } from '../server/src/services/router';

// Switch to latency-optimized routing
await setRoutingStrategy('fastest');

```

### Configure Custom Weights

```typescript
import { setCustomWeights, getCustomWeights } from '../server/src/services/router';

// 70% speed, 20% reliability, 10% intelligence
await setCustomWeights({
  reliability: 0.2,
  speed: 0.7,
  intelligence: 0.1
});

// Retrieved values are normalized to sum to 1.0
console.log(getCustomWeights());
// { reliability: 0.2, speed: 0.7, intelligence: 0.1 }

```

### HTTP API Examples

```bash

# Get current routing strategy

curl -s https://your-relay.example.com/api/routing/strategy | jq .

# Set strategy to emphasize reliability

curl -X PUT \
  -H "Content-Type: application/json" \
  -d '"reliable"' \
  https://your-relay.example.com/api/routing/strategy

# Update custom weights

curl -X PUT \
  -H "Content-Type: application/json" \
  -d '{"reliability":0.5,"speed":0.3,"intelligence":0.2}' \
  https://your-relay.example.com/api/routing/strategy/custom

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) | Core routing logic, strategy getters/setters, penalty calculation, and weight normalization |
| [`server/src/services/scoring.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.ts) | Bandit preset definitions (`BANDIT_PRESETS`) and score computation functions |
| [`docs/architecture.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/architecture.md) | High-level routing architecture documentation |
| [`README.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/README.md) | User-facing router configuration guide |

## Summary

- **Six strategies** control FreeLLMAPI request routing: `priority`, `balanced`, `smartest`, `fastest`, `reliable`, and `custom`
- **`priority`** uses static operator-defined order; all others use **Thompson-sampled bandit scoring**
- **Weight vectors** determine the tradeoff between reliability, speed, and intelligence in bandit-based strategies
- **Custom weights** allow fine-grained control via `setCustomWeights()` and are normalized automatically
- Strategy persistence uses SQLite `settings` table; runtime retrieval goes through `getRoutingStrategy()` in [`router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/router.ts)

## Frequently Asked Questions

### How do I switch from manual priority to automatic bandit routing?

Change the `routing_strategy` setting to `balanced`, `smartest`, `fastest`, or `reliable` using `setRoutingStrategy()` or the HTTP PUT `/api/routing/strategy` endpoint. The router immediately begins computing scores from historical model performance rather than following your static priority list.

### What happens when my highest-priority model returns a 429 error?

The router applies a temporary penalty via `getPenalty()` and falls back to the next-viable model in the chain. For `priority` strategy, this is the next-highest static priority. For bandit strategies, it's the next-highest Thompson-sampled score after applying the penalty.

### Can I use different strategies for different API endpoints or users?

The current implementation stores one global `routing_strategy` value. Per-endpoint or per-user strategies would require extending `getRoutingStrategy()` in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) to accept a context parameter and route accordingly. The underlying scoring infrastructure supports this, but the setting retrieval layer does not yet implement it.