FreellMAPI Routing Strategies: How the Six Methods Select and Score LLM Models

FreellMAPI uses six distinct routing strategies—priority, balanced, smartest, fastest, reliable, and custom—to select the optimal language model for each request, scoring candidates with a weighted bandit algorithm that factors in success rate, latency, cost, and peak-hour adjustments.

The tashfeenahmed/freellmapi repository implements an intelligent request router that dynamically selects between multiple LLM providers. Understanding these routing strategies and the underlying scoring mechanics is essential for optimizing performance, cost, and reliability in production deployments.

The Six Routing Strategies Explained

The RoutingStrategy enumeration is defined in server/src/services/scoring.ts at line 34, and the list of accepted values is validated in server/src/services/router.ts at line 395. Each strategy maps to a specific weight vector that determines how the router evaluates candidate models.

Priority Strategy

The priority strategy selects the model with the highest raw success rate, prioritizing reliability above all other factors. This approach ignores latency and cost metrics, making it ideal when you need the most dependable answer regardless of speed or expense.

Balanced Strategy (Default)

The balanced strategy mixes success rate, latency, and cost according to a preset weight set defined in BANDIT_PRESETS at line 44 of scoring.ts. As the default configuration, this strategy provides a general-purpose routing solution that optimizes the trade-off between speed, quality, and expense.

Smartest Strategy

The smartest strategy assigns extra weight to models that have historically produced higher intelligence scores, such as higher-quality completions or better instruction following. This method favors capability over cost, selecting the most powerful available model even when more economical options exist.

Fastest Strategy

The fastest strategy emphasizes low latency above all else, selecting the quickest responding model even if it offers lower capability. This approach suits latency-sensitive workloads such as real-time chat applications or voice transcription services.

Reliable Strategy

The reliable strategy focuses on models with consistent uptime and low error rates, applying a moderate cost bias to avoid expensive failures. This method targets mission-critical services that cannot tolerate intermittent provider outages.

Custom Strategy

The custom strategy allows operators to supply their own weight vector via configuration endpoints. This flexibility enables A/B testing new models or fine-tuning the scoring formula for specific business requirements, bypassing the preset configurations in BANDIT_PRESETS.

How the Router Scores Models

The scoring mechanism operates as a multi-armed bandit system that combines historical telemetry with configurable weights. The core logic resides in server/src/services/router.ts, where the orderChain function at line 1025 returns models sorted by their computed scores.

Weight Presets in BANDIT_PRESETS

For each non-custom strategy, the code maps a preset weight vector stored in BANDIT_PRESETS at line 44 of scoring.ts. These vectors define the relative importance of metrics such as success rate, average latency, and cost per token. The router applies these weights consistently unless overridden by peak-hour adjustments or custom configurations.

Peak-Hour Adjustment Logic

During configured peak hours, the router may down-weight cost-heavy models to preserve quota and budget. The weightsWithPeak adjustment occurs at line 634 of router.ts, where the system detects peak traffic periods and temporarily modifies the weight vectors to favor less expensive alternatives without completely sacrificing quality.

The Bandit Score Calculation in orderChain

The orderChain function at line 1025 of router.ts computes the final bandit score for every eligible model by combining the preset weights with live telemetry data. This includes recent success and failure counts, average response latency, token usage statistics, and historical error rates. The function returns a sorted array of models, with the highest-scoring candidate receiving the request.

Key-Selection Alignment

The scoring system respects the current key-selection strategy, ensuring that models are only scored if their associated API keys are enabled and available. This alignment prevents the router from selecting providers with exhausted quotas or disabled credentials, effectively filtering the candidate pool before the bandit algorithm runs.

Retrieving Debug Data with getRoutingScores

For observability and debugging, the getRoutingScores() function at line 2079 of router.ts returns a complete breakdown of the routing decision. This includes the active strategy, applied weight vectors, any custom weights, computed scores for each model, and the peak-adjustment flag. This data enables operators to audit why specific models were selected and fine-tune their configurations accordingly.

Configuring Routing in Practice

You can switch strategies and inspect scores programmatically using the router service exports.

Change to the fastest strategy and view current diagnostics:

import { setRoutingStrategy, getRoutingScores } from './services/router.js';

// Switch to lowest-latency routing
setRoutingStrategy('fastest');

// Retrieve scoring table for debugging
const scores = getRoutingScores();
console.log('Active strategy:', scores.strategy);
console.table(scores.scores);

Implement a custom weight vector that doubles the importance of latency:

import { setCustomWeights } from './services/router.js';

setCustomWeights({
  successRate: 1,
  latencyMs: 2,   // prioritize speed
  costUsd: 0.5,
  intelligence: 0.8
});

Integrate the router into an API endpoint handler:

import { routeRequest } from './services/router.js';

app.post('/v1/chat/completions', async (req, res) => {
  const result = await routeRequest(req.body);
  
  if (result.error) {
    return res.status(400).json(result.error);
  }
  
  const response = await fetch(result.endpoint, {
    method: 'POST',
    headers: result.headers,
    body: JSON.stringify(result.payload)
  });
  
  res.json(await response.json());
});

Summary

  • Six strategies control model selection: priority, balanced, smartest, fastest, reliable, and custom, defined in scoring.ts and validated in router.ts.
  • Balanced serves as the default strategy, while custom allows operator-defined weight vectors for specialized use cases.
  • Bandit scoring combines preset weights from BANDIT_PRESETS with live telemetry to rank models via the orderChain function.
  • Peak-hour adjustments automatically modify weights to reduce costs during high-traffic periods, implemented at line 634 of router.ts.
  • Key-selection alignment ensures only models with available API credentials are considered in the scoring pool.
  • Diagnostic visibility comes through getRoutingScores(), which exposes the full scoring breakdown for debugging and UI consumption.

Frequently Asked Questions

What is the default routing strategy in FreellMAPI?

The balanced strategy is the default configuration. It employs a preset weight vector that equally considers success rate, latency, and cost to provide optimal performance for general-purpose workloads without requiring manual tuning.

How do I create a custom routing strategy?

Enable the custom strategy by calling setCustomWeights() with your desired metric weights. This overrides the BANDIT_PRESETS configuration and allows you to define specific priorities, such as maximizing intelligence scores for research tasks or minimizing latency for real-time applications.

Does the router consider API key availability when scoring models?

Yes, the router filters the candidate pool before scoring based on the current key-selection strategy. Models associated with disabled, exhausted, or invalid API keys are excluded from the bandit calculation, ensuring the system never selects a provider it cannot authenticate with.

How can I view the current routing scores for debugging?

Call the getRoutingScores() function exposed from server/src/services/router.ts. This returns the active strategy, applied weights (including any peak-hour adjustments), individual model scores, and the selection rationale, enabling comprehensive auditing of routing decisions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →