How FreeLLMAPI Routing Strategies Work: A Complete Guide to Model Selection
FreeLLMAPI uses six distinct routing strategies—from balanced to priority—that apply different weight vectors to reliability, speed, and intelligence scores to determine which LLM handles each request.
The tashfeenahmed/freellmapi router implements a contextual bandit system that evaluates every enabled model across five normalized axes. Your chosen FreeLLMAPI routing strategy supplies the weight vector that ultimately decides which model serves incoming traffic, making strategy selection critical for latency, accuracy, and uptime.
The Six FreeLLMAPI Routing Strategies Explained
Each strategy corresponds to a specific weight distribution across three primary axes: reliability, speed, and intelligence. The router normalizes these weights to ensure they sum to 1.0 before scoring.
Balanced (Default)
- Weight Vector: 0.50 reliability / 0.25 speed / 0.25 intelligence
- Use Case: General-purpose traffic where no single attribute should dominate
- Behavior: Provides moderate bias toward reliable models while still rewarding performance and capability
The balanced strategy is defined as DEFAULT_STRATEGY in the codebase and serves as the fallback when no explicit configuration exists.
Smartest
- Weight Vector: 0.35 reliability / 0.10 speed / 0.55 intelligence
- Use Case: Code generation, complex reasoning, or any task where answer quality trumps latency
- Behavior: Pushes the intelligence axis to the forefront; fast but less capable models are deprioritized significantly
Fastest
- Weight Vector: 0.35 reliability / 0.55 speed / 0.10 intelligence
- Use Case: Real-time chat applications, UI-driven assistants, or latency-sensitive streaming
- Behavior: Emphasizes speed above all else; the router prefers models that return tokens quickly, even if they offer lower capability
Reliable
- Weight Vector: 0.70 reliability / 0.15 speed / 0.15 intelligence
- Use Case: Production pipelines or critical workflows where uptime and error-free responses are mandatory
- Behavior: Strongly favors models with healthy uptime signals; demotes instances showing recent 429 or 5xx errors
Custom
- Weight Vector: User-provided (normalized to sum = 1.0)
- Use Case: Operators requiring a bespoke balance not covered by presets
- Behavior: Uses the exact vector stored in the
routing_custom_weightssetting, accessible viagetCustomWeights()inserver/src/services/router.ts
Priority
- Weight Vector: N/A (manual ordering + 429 penalty)
- Use Case: Legacy fallback chains or situations requiring explicit static ordering
- Behavior: Ignores bandit scores entirely; models are ordered by the explicit
priorityfield with a penalty added for recent 429 responses
How the Router Calculates Model Scores
In server/src/services/router.ts, the router constructs a composite score for every chain entry (enabled model) using five normalized axes:
- Reliability: Historical uptime and error rates
- Speed: Token-generation latency
- Intelligence: Capability benchmarks for the target task
- Headroom: Remaining capacity relative to rate limits
- Rate-limit: Current throttle status
The helper getActiveRoutingWeights() (lines 74-76) retrieves the current strategy’s weight vector, which is then multiplied against these normalized scores. The model with the highest final score receives the request.
Advanced Routing Features
Peak-Hours Adjustments
For balanced and smartest strategies, the raw preset weights can be dynamically re-weighted during operator-defined peak windows (e.g., 18:00–06:00 UTC). As implemented in server/src/services/scoring.ts (lines 54-66), the constant PEAK_SPEED_TO_RELIABILITY = 0.6 shifts a portion of the speed weight onto reliability during high-traffic periods.
Strategies fastest and reliable are exempt from this adjustment to preserve their original intent.
Exploration Floor for New Models
When routing_explore_enabled = 1, the system prevents "starvation" of new or rarely-used models. As defined in server/src/services/router.ts (lines 14-22), any model with fewer than EXPLORE_MIN_SAMPLES = 5 decay-weighted samples receives a guaranteed 10% chance of being selected first, ensuring the bandit gathers sufficient data on all candidates.
Per-Model Overrides
The environment variable MODEL_ROUTING_OVERRIDES allows operators to promote (>1) or demote (<1) specific models after guard-rail multipliers are applied. This override works with every strategy except priority, where manual ordering already handles such tweaks. The logic resides in server/src/services/router.ts (lines 72-80).
Configuring Routing Strategies via API and Code
You can inspect and modify routing configurations through the REST API or internal TypeScript functions.
Retrieving the Current Strategy
import fetch from 'node-fetch';
async function getStrategy() {
const resp = await fetch('http://localhost:8080/api/fallback/routing', {
headers: { 'Authorization': `Bearer ${process.env.FREEAPI_KEY}` },
});
const data = await resp.json();
console.log('Current strategy:', data.strategy); // e.g. "balanced"
}
getStrategy();
Switching to Fastest Mode
import fetch from 'node-fetch';
async function setFastest() {
await fetch('http://localhost:8080/api/fallback/routing', {
method: 'PUT',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${process.env.FREEAPI_KEY}`,
},
body: JSON.stringify({ strategy: 'fastest' }),
});
console.log('Routing strategy updated to fastest');
}
setFastest();
Internal API Usage
When extending the server, import directly from the router service:
import { setRoutingStrategy, getRoutingStrategy } from '../services/router.js';
function toggleReliableMode(enable: boolean) {
const newStrategy = enable ? 'reliable' : 'balanced';
setRoutingStrategy(newStrategy as any);
console.log('Switched routing strategy to', getRoutingStrategy());
}
Reading Custom Weights
import { getCustomWeights } from '../services/router.js';
const weights = getCustomWeights();
console.log('Custom weights →', weights);
// { reliability: 0.4, speed: 0.3, intelligence: 0.3 }
Key Source Files and Implementation Details
| File | Purpose |
|---|---|
server/src/services/router.ts |
Core routing logic, strategy getters/setters (getRoutingStrategy, setRoutingStrategy), weight handling, and override processing |
server/src/services/scoring.ts |
Scoring functions including reliabilityPosterior, speedScore, headroomFactor, and peakAdjustedWeights |
docs/architecture/01-routing-and-bandit-scoring.md |
Architectural deep-dive explaining each axis, presets, and guard-rail interactions |
The getRoutingStrategy() helper validates all inputs against VALID_STRATEGIES (['priority','balanced','smartest','fastest','reliable','custom']) as shown in lines 22-28 of the router service.
Summary
- FreeLLMAPI routing strategies determine model selection through weighted scoring across reliability, speed, intelligence, headroom, and rate-limit axes.
- Choose
balancedfor general workloads,smartestfor quality-critical tasks,fastestfor latency-sensitive applications, andreliablefor uptime-critical production systems. - Use
customwhen you need precise control over the weight vector viarouting_custom_weights. - Select
priorityonly when you require static manual ordering without bandit scoring. - Enable
routing_explore_enabledto prevent new models from starving while the system gathers performance data. - Configure strategies via
GET/PUT /api/fallback/routingor programmatically throughserver/src/services/router.ts.
Frequently Asked Questions
How do I know which routing strategy is currently active?
Send a GET request to /api/fallback/routing or call getRoutingStrategy() from server/src/services/router.ts. The response includes the active strategy name and current weight vector. According to the source code, the system defaults to balanced if no strategy has been explicitly configured.
Can I use custom weights without creating a permanent preset?
Yes. Set the strategy to custom and provide your specific weight distribution through the routing_custom_weights setting. The router normalizes these values to sum to 1.0 automatically. Access your current custom configuration programmatically using getCustomWeights() in the router service.
Why does the fastest strategy still allocate 35% weight to reliability?
Even in fastest mode, the router maintains a baseline reliability weight to avoid routing requests to completely unhealthy or throttled models. This prevents cascading failures while still strongly favoring low-latency candidates. If you need absolute speed without reliability checks, you would need to modify the weight constants in server/src/services/scoring.ts or use per-model overrides.
Do routing strategies affect the exploration floor setting?
No. The exploration floor (EXPLORE_MIN_SAMPLES = 5) operates independently of your chosen strategy. When routing_explore_enabled = 1, models with fewer than five decay-weighted samples receive a 10% selection probability regardless of whether you use balanced, smartest, or any other strategy. This ensures the bandit gathers sufficient data on new models without interfering with your primary routing logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →