How FreeLLMAPI Calculates Reliability, Speed, and Capability Scores for Model Routing
FreeLLMAPI calculates reliability using a Bayesian Beta posterior from success/failure counters, measures speed through normalized throughput with optional weight shifting, and assigns capability via preset intelligence weights, combining all three into a weighted score for routing decisions.
FreeLLMAPI implements a three-axis scoring system to determine which large language model (LLM) should handle each request. The system quantifies reliability, speed, and capability (intelligence) through statistical methods and configurable presets defined in the tashfeenahmed/freellmapi repository.
The Three-Axis Scoring Architecture
FreeLLMAPI evaluates every model candidate across three distinct dimensions before routing traffic. Each axis serves a specific purpose in ensuring quality of service.
Reliability: Bayesian Beta Posterior
Reliability represents the probability that a model returns a successful response without timeout or error. The system calculates this using a Beta distribution posterior derived from per-model telemetry counters.
In server/src/services/scoring.ts (lines 273‑285), the reliabilityPosterior function builds the distribution by combining observed outcomes with a community-wide prior:
// Pseudo-code based on scoring.ts implementation
function reliabilityPosterior(modelSuccesses, modelFailures, communityPrior) {
const alpha = modelSuccesses + communityPrior.successes;
const beta = modelFailures + communityPrior.failures;
return { alpha, beta };
}
The system uses this posterior in two ways:
sampleBeta(alpha, beta): Stochastically samples a reliability value for routing decisions, adding exploration to the selection process.expectedReliability(alpha, beta): Computesα / (α + β)for dashboard displays and deterministic comparisons.
The community prior (community.successes, community.failures) smooths sparse data for new or rarely-used models.
Speed: Normalized Throughput with Deliberate Degradation
Speed measures normalized request throughput, capped to prevent extremely fast but unreliable models from dominating the routing pool. The raw throughput value is transformed into a [0, 1] scale based on observed maximums.
According to the implementation commentary in scoring.ts around lines 63‑66, the system deliberately downgrades raw speed scores because speed alone does not guarantee user satisfaction. Additionally, the architecture supports a dynamic speed→reliability shift that transfers a configurable portion of the speed weight to reliability during peak traffic windows.
Capability: Preset Intelligence Weights
Unlike reliability and speed, capability scores (labeled "intelligence" in the codebase) are preset weights rather than inferred telemetry values. Each routing preset defines how much emphasis to place on model intelligence versus other factors.
The preset definitions in scoring.ts (lines 46‑51) establish fixed weight distributions:
| Preset | Reliability | Speed | Intelligence |
|---|---|---|---|
| balanced | 0.50 | 0.25 | 0.25 |
| smartest | 0.35 | 0.10 | 0.55 |
| fastest | 0.35 | 0.55 | 0.10 |
| reliable | 0.70 | 0.15 | 0.15 |
The actual capability value for a specific model defaults to 1.0 (normalized) for most LLMs, meaning the intelligence axis primarily acts as a modifier controlled by preset selection.
Combining Scores for Routing Decisions
FreeLLMAPI merges the three axes into a single composite score using the combineScore function implemented in scoring.ts (lines 530‑532):
const wSum = weights.reliability + weights.speed + weights.intelligence || 1;
const score =
(weights.reliability * inputs.reliability +
weights.speed * inputs.speed +
weights.intelligence * inputs.intelligence) / wSum;
The weights object comes from BANDIT_PRESETS, which exports configurations for the four routing strategies listed above.
Dynamic Weight Shifting
During high-traffic periods, the system can execute a speed→reliability shift to prioritize stable models over fast ones. Controlled by the speedReliabilityShift parameter (default 0), this logic in scoring.ts (lines 315‑327) transfers a fraction of the speed weight onto reliability:
// Conceptual implementation from lines 315-327
if (speedReliabilityShift > 0) {
weights.reliability += weights.speed * speedReliabilityShift;
weights.speed *= (1 - speedReliabilityShift);
}
This prevents flaky but high-throughput models from receiving disproportionate traffic during peak hours.
Implementation in the Codebase
The scoring system resides primarily in server/src/services/scoring.ts, while the routing logic consumes these scores in server/src/services/router.ts.
When a request arrives, the router evaluates eligible models by calling combineScore for each candidate and selects the highest-scoring instance. The router integration appears around lines 1033‑1036 in router.ts, where the bandit selection algorithm uses the composite scores to distribute traffic optimally.
The server/src/services/declarative-config.ts file validates incoming model configurations, ensuring that reliability and speed fields conform to expected schemas before entering the scoring pipeline.
Practical Usage Example
import { combineScore, BANDIT_PRESETS, reliabilityPosterior, sampleBeta } from './services/scoring';
// 1. Calculate Bayesian reliability for a model with 50 successes and 2 failures
const { alpha, beta } = reliabilityPosterior(50, 2, { successes: 1000, failures: 50 });
const reliabilityScore = sampleBeta(alpha, beta); // stochastic sampling for routing
// 2. Normalize observed throughput (e.g., 150 req/s out of 200 max)
const speedScore = 150 / 200; // 0.75
// 3. Combine using the "smartest" preset to prioritize capability
const finalScore = combineScore(
{ reliability: reliabilityScore, speed: speedScore, intelligence: 1.0 },
BANDIT_PRESETS.smartest
);
// The router selects the model with the highest finalScore
Unit tests in server/src/__tests__/services/scoring.test.ts and server/src/__tests__/services/router-bandit.test.ts validate that these calculations produce expected routing behaviors across different presets and traffic conditions.
Summary
- Reliability derives from a Bayesian Beta posterior (
reliabilityPosterior) using per-model success/failure counters and a community prior, with values sampled viasampleBetafor routing. - Speed represents normalized throughput deliberately capped to prevent unreliable fast models from dominating, with optional dynamic weight shifting to reliability during peak traffic.
- Capability (intelligence) uses preset-defined weights from
BANDIT_PRESETSrather than telemetry, allowing operators to prioritize quality via the "smartest" preset or speed via the "fastest" preset. - Combination occurs through
combineScoreinscoring.ts, producing a weighted average used byrouter.ts(lines 1033‑1036) to select the optimal model for each request.
Frequently Asked Questions
How does FreeLLMAPI handle new models with no historical reliability data?
FreeLLMAPI applies a community prior to smooth sparse data for new models. The reliabilityPosterior function adds community-wide success and failure counts to the model's observed counters before calculating the Beta distribution. This ensures new models start with moderate reliability scores rather than extreme values, preventing them from immediately dominating or being excluded from routing.
Can I create custom routing presets beyond the four built-in options?
The codebase defines weights through the BANDIT_PRESETS constant in scoring.ts (lines 46‑51). While the current implementation provides balanced, smartest, fastest, and reliable presets, the combineScore function accepts any arbitrary weight object. You can pass custom { reliability, speed, intelligence } weights directly to combineScore to create specialized routing strategies for specific use cases.
What triggers the speed-to-reliability weight shift?
The speed→reliability shift activates via the speedReliabilityShift parameter, defaulting to 0 (no shift). During high-traffic periods, operators or automated systems can increase this value (typically between 0 and 1) to transfer proportionally more weight from speed to reliability. This prevents routing storms where fast but unstable models fail repeatedly under load, degrading overall system reliability.
How does the router actually use the calculated scores?
The router in server/src/services/router.ts calls combineScore for every eligible model candidate during request processing (around lines 1033‑1036). It then selects the model with the highest composite score, using the Beta-sampled reliability values to introduce controlled randomness that prevents overloading single models while still favoring high-performing candidates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →