How FreeLLMAPI Handles Smart Routing for LLM Requests: Anatomy of the Routing Engine

FreeLLMAPI routes LLM requests through a multi-stage decision pipeline that ranks candidate models by configurable strategy, filters by capability requirements, and selects viable API keys while respecting rate limits and quotas.

FreeLLMAPI is an open-source aggregation layer that unifies multiple large language model providers behind a single interface. Its approach to smart routing for LLM requests dynamically selects the optimal provider, model, and API key for every prompt, balancing reliability, latency, and capacity constraints. The core logic resides in server/src/services/router.ts, supported by specialized modules for scoring, rate limiting, and quota management.

Building and Sorting the Candidate Chain

The routing process begins by constructing a prioritized list of eligible models. The system first invokes getActiveChain to retrieve the current model catalog from the SQLite database, filtering out disabled entries. This produces the initial candidate chain containing all potentially available models.

Once populated, the chain is ordered via orderChain(chain, strategy). The routing strategy determines the sort order and can be configured per deployment:

  • Bandit-based: Dynamically balances exploration and exploitation using performance history
  • Balanced: Distributes load evenly across capable providers
  • Manual priority: Respects static administrator-defined rankings

According to server/src/services/scoring.js, the bandit-based strategy relies on scoring functions including speedScore, intelligenceScore, and combineScore to rank models by historical performance.

Exploration and Session Persistence

When operating in bandit-based mode with exploration enabled, the router injects randomness to optimize long-term performance while honoring client preferences.

Exploration for Unmeasured Models

If the strategy is bandit-based and exploration is enabled, the router identifies unmeasured models—those lacking sufficient reliability or speed data—and randomly bumps one to the front of the chain. This ensures new models collect performance samples without compromising overall system availability.

Sticky-Session Pinning

Clients may request consistency by passing a preferredModelDbId parameter. When provided, the router moves that specific model to the front of the candidate chain, injecting it if absent. This sticky-session mechanism ensures that multi-turn conversations or specific workloads remain pinned to a consistent backend, falling back automatically if the model becomes disabled.

Capability Filtering and Context Validation

Before attempting a provider, the router applies fast-path filters to eliminate unsuitable candidates. These checks ensure functional compatibility and prevent guaranteed failures.

Feature Requirements

The router validates specialized capabilities through boolean flags:

  • requireVision: Ensures the model supports image inputs
  • requireTools: Confirms function-calling capabilities
  • requireStructured: Verifies JSON schema or structured output support

Context Window and Token Budgets

The fitsContextWindowStrict function calculates the request's total token footprint (input plus a capped output reserve) and validates it against the model's advertised context window. This check applies a safety factor for "margin-deferred" models to prevent overflow errors.

Additionally, the router enforces per-minute token budgets (tpm_limit) to avoid predictable 413 Payload Too Large errors. Administrators may also specify skipPlatforms as a Set of provider names to exclude specific backends from individual requests.

API Key Selection and Rate Limit Enforcement

Once a model passes capability filters, the router must select a usable credential.

The selectKeyForModel Function

For each surviving candidate, the router calls selectKeyForModel (implemented in server/src/services/select-key.ts). This function:

  1. Retrieves all API keys associated with the model's platform
  2. Applies rate-limit and concurrency checks via canMakeRequest, canUseTokens, and canUseProviderMinute from server/src/services/ratelimit.js
  3. Honors key-level quotas, cooldown periods, and proxy overrides defined in server/src/services/provider-quota.ts
  4. Returns the first viable key or marks the model as exhausted

The provider-quota module distinguishes between free and paid tiers to calculate available headroom before attempting key selection.

Exhaustion Handling and Diagnostic Reporting

When no models or keys satisfy the request constraints, the router throws a RouteError. The error message is constructed by summarizeExhaustion, which groups diagnostic information into specific categories: rate-limiting, insufficient quota, oversized prompts, or absent API keys. This diagnostic grouping helps operators identify whether they need to add more keys, adjust quotas, or wait for rate-limit windows to reset.

Practical Routing Examples

The routeRequest function in server/src/services/router.ts serves as the primary entry point. Below are common configuration patterns.

Automatic Provider Selection

For standard requests with no special constraints, call routeRequest with an estimated token count:

import { routeRequest } from '@/server/src/services/router';

const estimatedTokens = 1200;
const result = routeRequest(estimatedTokens);

console.log({
  provider: result.provider.name,
  modelId: result.modelId,
  apiKey: result.apiKey,
  endpointScope: result.endpointScope,
});

Capability Filtering and Provider Exclusion

To require vision support while excluding a specific provider:

import { routeRequest } from '@/server/src/services/router';

const result = routeRequest(
  800,                     // estimated tokens
  undefined,               // keyBlacklist
  undefined,               // explicit model id
  true,                    // requireVision
  false,                   // requireTools
  undefined,               // requireStructured
  undefined,               // sticky session id
  new Set(['groq'])        // skipPlatforms
);

Sticky Session with Fallback

To maintain session persistence using a previously selected model:

import { routeRequest } from '@/server/src/services/router';

const stickyModelId = 42; // DB id of the previously-used model
const result = routeRequest(
  1000,                    // estimated tokens
  undefined,
  stickyModelId,           // preferredModelDbId
  false,
  false
);

Summary

  • Candidate construction uses getActiveChain and orderChain to rank models by configurable strategies defined in shared/types.ts and scored via server/src/services/scoring.js.
  • Exploration logic ensures new models receive traffic for performance sampling, while sticky-session pinning respects client preferences via preferredModelDbId.
  • Capability validation filters by vision, tools, structured output, context window fit, and token budgets before attempting key selection.
  • Key selection via selectKeyForModel enforces rate limits through server/src/services/ratelimit.js and quota limits through server/src/services/provider-quota.ts.
  • Failure diagnostics aggregate exhaustion reasons through summarizeExhaustion to guide operational responses.

Frequently Asked Questions

How does FreeLLMAPI balance load across multiple providers?

FreeLLMAPI applies the configured routing strategy—bandit-based, balanced, or manual priority—via the orderChain function. Bandit-based strategies use historical performance data from server/src/services/scoring.js to weight models, while balanced mode distributes requests evenly. The system dynamically reorders the candidate chain for each request based on real-time capacity and performance metrics.

What happens when no providers can handle a request?

When all candidates fail validation or key selection, the router throws a RouteError constructed by summarizeExhaustion. This error categorizes the failure reason—such as rate-limiting, exhausted quotas, or oversized context windows—and suggests specific remediation steps like adding new API keys or waiting for rate-limit resets.

Can I force specific models for certain request types?

Yes. You can enforce sticky sessions by passing a preferredModelDbId to routeRequest, which moves that model to the front of the candidate chain. Additionally, you can filter by capabilities (requireVision, requireTools, requireStructured) or exclude entire providers using the skipPlatforms parameter to guarantee specific routing behavior.

How does the router handle new models with no historical data?

Through the exploration step in bandit-based routing mode. When enabled, the router randomly promotes unmeasured models—those lacking reliability or speed samples—to the front of the chain with a certain probability. This ensures new integrations collect performance data without requiring manual traffic allocation or compromising system reliability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →