How FreeLLMAPI Handles Smart Routing for LLM Requests: Anatomy of the Routing Engine
FreeLLMAPI routes LLM requests through a multi-stage decision pipeline that ranks candidate models by configurable strategy, filters by capability requirements, and selects viable API keys while respecting rate limits and quotas.
FreeLLMAPI is an open-source aggregation layer that unifies multiple large language model providers behind a single interface. Its approach to smart routing for LLM requests dynamically selects the optimal provider, model, and API key for every prompt, balancing reliability, latency, and capacity constraints. The core logic resides in server/src/services/router.ts, supported by specialized modules for scoring, rate limiting, and quota management.
Building and Sorting the Candidate Chain
The routing process begins by constructing a prioritized list of eligible models. The system first invokes getActiveChain to retrieve the current model catalog from the SQLite database, filtering out disabled entries. This produces the initial candidate chain containing all potentially available models.
Once populated, the chain is ordered via orderChain(chain, strategy). The routing strategy determines the sort order and can be configured per deployment:
- Bandit-based: Dynamically balances exploration and exploitation using performance history
- Balanced: Distributes load evenly across capable providers
- Manual priority: Respects static administrator-defined rankings
According to server/src/services/scoring.js, the bandit-based strategy relies on scoring functions including speedScore, intelligenceScore, and combineScore to rank models by historical performance.
Exploration and Session Persistence
When operating in bandit-based mode with exploration enabled, the router injects randomness to optimize long-term performance while honoring client preferences.
Exploration for Unmeasured Models
If the strategy is bandit-based and exploration is enabled, the router identifies unmeasured models—those lacking sufficient reliability or speed data—and randomly bumps one to the front of the chain. This ensures new models collect performance samples without compromising overall system availability.
Sticky-Session Pinning
Clients may request consistency by passing a preferredModelDbId parameter. When provided, the router moves that specific model to the front of the candidate chain, injecting it if absent. This sticky-session mechanism ensures that multi-turn conversations or specific workloads remain pinned to a consistent backend, falling back automatically if the model becomes disabled.
Capability Filtering and Context Validation
Before attempting a provider, the router applies fast-path filters to eliminate unsuitable candidates. These checks ensure functional compatibility and prevent guaranteed failures.
Feature Requirements
The router validates specialized capabilities through boolean flags:
requireVision: Ensures the model supports image inputsrequireTools: Confirms function-calling capabilitiesrequireStructured: Verifies JSON schema or structured output support
Context Window and Token Budgets
The fitsContextWindowStrict function calculates the request's total token footprint (input plus a capped output reserve) and validates it against the model's advertised context window. This check applies a safety factor for "margin-deferred" models to prevent overflow errors.
Additionally, the router enforces per-minute token budgets (tpm_limit) to avoid predictable 413 Payload Too Large errors. Administrators may also specify skipPlatforms as a Set of provider names to exclude specific backends from individual requests.
API Key Selection and Rate Limit Enforcement
Once a model passes capability filters, the router must select a usable credential.
The selectKeyForModel Function
For each surviving candidate, the router calls selectKeyForModel (implemented in server/src/services/select-key.ts). This function:
- Retrieves all API keys associated with the model's platform
- Applies rate-limit and concurrency checks via
canMakeRequest,canUseTokens, andcanUseProviderMinutefromserver/src/services/ratelimit.js - Honors key-level quotas, cooldown periods, and proxy overrides defined in
server/src/services/provider-quota.ts - Returns the first viable key or marks the model as exhausted
The provider-quota module distinguishes between free and paid tiers to calculate available headroom before attempting key selection.
Exhaustion Handling and Diagnostic Reporting
When no models or keys satisfy the request constraints, the router throws a RouteError. The error message is constructed by summarizeExhaustion, which groups diagnostic information into specific categories: rate-limiting, insufficient quota, oversized prompts, or absent API keys. This diagnostic grouping helps operators identify whether they need to add more keys, adjust quotas, or wait for rate-limit windows to reset.
Practical Routing Examples
The routeRequest function in server/src/services/router.ts serves as the primary entry point. Below are common configuration patterns.
Automatic Provider Selection
For standard requests with no special constraints, call routeRequest with an estimated token count:
import { routeRequest } from '@/server/src/services/router';
const estimatedTokens = 1200;
const result = routeRequest(estimatedTokens);
console.log({
provider: result.provider.name,
modelId: result.modelId,
apiKey: result.apiKey,
endpointScope: result.endpointScope,
});
Capability Filtering and Provider Exclusion
To require vision support while excluding a specific provider:
import { routeRequest } from '@/server/src/services/router';
const result = routeRequest(
800, // estimated tokens
undefined, // keyBlacklist
undefined, // explicit model id
true, // requireVision
false, // requireTools
undefined, // requireStructured
undefined, // sticky session id
new Set(['groq']) // skipPlatforms
);
Sticky Session with Fallback
To maintain session persistence using a previously selected model:
import { routeRequest } from '@/server/src/services/router';
const stickyModelId = 42; // DB id of the previously-used model
const result = routeRequest(
1000, // estimated tokens
undefined,
stickyModelId, // preferredModelDbId
false,
false
);
Summary
- Candidate construction uses
getActiveChainandorderChainto rank models by configurable strategies defined inshared/types.tsand scored viaserver/src/services/scoring.js. - Exploration logic ensures new models receive traffic for performance sampling, while sticky-session pinning respects client preferences via
preferredModelDbId. - Capability validation filters by vision, tools, structured output, context window fit, and token budgets before attempting key selection.
- Key selection via
selectKeyForModelenforces rate limits throughserver/src/services/ratelimit.jsand quota limits throughserver/src/services/provider-quota.ts. - Failure diagnostics aggregate exhaustion reasons through
summarizeExhaustionto guide operational responses.
Frequently Asked Questions
How does FreeLLMAPI balance load across multiple providers?
FreeLLMAPI applies the configured routing strategy—bandit-based, balanced, or manual priority—via the orderChain function. Bandit-based strategies use historical performance data from server/src/services/scoring.js to weight models, while balanced mode distributes requests evenly. The system dynamically reorders the candidate chain for each request based on real-time capacity and performance metrics.
What happens when no providers can handle a request?
When all candidates fail validation or key selection, the router throws a RouteError constructed by summarizeExhaustion. This error categorizes the failure reason—such as rate-limiting, exhausted quotas, or oversized context windows—and suggests specific remediation steps like adding new API keys or waiting for rate-limit resets.
Can I force specific models for certain request types?
Yes. You can enforce sticky sessions by passing a preferredModelDbId to routeRequest, which moves that model to the front of the candidate chain. Additionally, you can filter by capabilities (requireVision, requireTools, requireStructured) or exclude entire providers using the skipPlatforms parameter to guarantee specific routing behavior.
How does the router handle new models with no historical data?
Through the exploration step in bandit-based routing mode. When enabled, the router randomly promotes unmeasured models—those lacking reliability or speed samples—to the front of the chain with a certain probability. This ensures new integrations collect performance data without requiring manual traffic allocation or compromising system reliability.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →