# How FreeLLMAPI Handles Smart Routing for LLM Requests: Anatomy of the Routing Engine

> Discover how FreeLLMAPI routes LLM requests via a multi-stage pipeline. Learn how it ranks models, filters by capabilities, and selects API keys while managing rate limits and quotas.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-09-02

---

**FreeLLMAPI routes LLM requests through a multi-stage decision pipeline that ranks candidate models by configurable strategy, filters by capability requirements, and selects viable API keys while respecting rate limits and quotas.**

FreeLLMAPI is an open-source aggregation layer that unifies multiple large language model providers behind a single interface. Its approach to **smart routing for LLM requests** dynamically selects the optimal provider, model, and API key for every prompt, balancing reliability, latency, and capacity constraints. The core logic resides in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), supported by specialized modules for scoring, rate limiting, and quota management.

## Building and Sorting the Candidate Chain

The routing process begins by constructing a prioritized list of eligible models. The system first invokes `getActiveChain` to retrieve the current model catalog from the SQLite database, filtering out disabled entries. This produces the initial **candidate chain** containing all potentially available models.

Once populated, the chain is ordered via `orderChain(chain, strategy)`. The **routing strategy** determines the sort order and can be configured per deployment:

- **Bandit-based**: Dynamically balances exploration and exploitation using performance history
- **Balanced**: Distributes load evenly across capable providers  
- **Manual priority**: Respects static administrator-defined rankings

According to [`server/src/services/scoring.js`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.js), the bandit-based strategy relies on scoring functions including `speedScore`, `intelligenceScore`, and `combineScore` to rank models by historical performance.

## Exploration and Session Persistence

When operating in bandit-based mode with exploration enabled, the router injects randomness to optimize long-term performance while honoring client preferences.

### Exploration for Unmeasured Models

If the strategy is bandit-based and **exploration** is enabled, the router identifies *unmeasured* models—those lacking sufficient reliability or speed data—and randomly bumps one to the front of the chain. This ensures new models collect performance samples without compromising overall system availability.

### Sticky-Session Pinning

Clients may request consistency by passing a `preferredModelDbId` parameter. When provided, the router moves that specific model to the front of the candidate chain, injecting it if absent. This **sticky-session** mechanism ensures that multi-turn conversations or specific workloads remain pinned to a consistent backend, falling back automatically if the model becomes disabled.

## Capability Filtering and Context Validation

Before attempting a provider, the router applies **fast-path filters** to eliminate unsuitable candidates. These checks ensure functional compatibility and prevent guaranteed failures.

### Feature Requirements

The router validates specialized capabilities through boolean flags:

- `requireVision`: Ensures the model supports image inputs
- `requireTools`: Confirms function-calling capabilities  
- `requireStructured`: Verifies JSON schema or structured output support

### Context Window and Token Budgets

The `fitsContextWindowStrict` function calculates the request's total token footprint (input plus a capped output reserve) and validates it against the model's advertised context window. This check applies a safety factor for "margin-deferred" models to prevent overflow errors.

Additionally, the router enforces per-minute token budgets (`tpm_limit`) to avoid predictable `413 Payload Too Large` errors. Administrators may also specify `skipPlatforms` as a `Set` of provider names to exclude specific backends from individual requests.

## API Key Selection and Rate Limit Enforcement

Once a model passes capability filters, the router must select a usable credential.

### The selectKeyForModel Function

For each surviving candidate, the router calls `selectKeyForModel` (implemented in [`server/src/services/select-key.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/select-key.ts)). This function:

1. Retrieves all API keys associated with the model's platform
2. Applies rate-limit and concurrency checks via `canMakeRequest`, `canUseTokens`, and `canUseProviderMinute` from [`server/src/services/ratelimit.js`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.js)
3. Honors key-level quotas, cooldown periods, and proxy overrides defined in [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts)
4. Returns the first viable key or marks the model as exhausted

The provider-quota module distinguishes between free and paid tiers to calculate available headroom before attempting key selection.

## Exhaustion Handling and Diagnostic Reporting

When no models or keys satisfy the request constraints, the router throws a `RouteError`. The error message is constructed by `summarizeExhaustion`, which groups diagnostic information into specific categories: rate-limiting, insufficient quota, oversized prompts, or absent API keys. This diagnostic grouping helps operators identify whether they need to add more keys, adjust quotas, or wait for rate-limit windows to reset.

## Practical Routing Examples

The `routeRequest` function in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) serves as the primary entry point. Below are common configuration patterns.

### Automatic Provider Selection

For standard requests with no special constraints, call `routeRequest` with an estimated token count:

```typescript
import { routeRequest } from '@/server/src/services/router';

const estimatedTokens = 1200;
const result = routeRequest(estimatedTokens);

console.log({
  provider: result.provider.name,
  modelId: result.modelId,
  apiKey: result.apiKey,
  endpointScope: result.endpointScope,
});

```

### Capability Filtering and Provider Exclusion

To require vision support while excluding a specific provider:

```typescript
import { routeRequest } from '@/server/src/services/router';

const result = routeRequest(
  800,                     // estimated tokens
  undefined,               // keyBlacklist
  undefined,               // explicit model id
  true,                    // requireVision
  false,                   // requireTools
  undefined,               // requireStructured
  undefined,               // sticky session id
  new Set(['groq'])        // skipPlatforms
);

```

### Sticky Session with Fallback

To maintain session persistence using a previously selected model:

```typescript
import { routeRequest } from '@/server/src/services/router';

const stickyModelId = 42; // DB id of the previously-used model
const result = routeRequest(
  1000,                    // estimated tokens
  undefined,
  stickyModelId,           // preferredModelDbId
  false,
  false
);

```

## Summary

- **Candidate construction** uses `getActiveChain` and `orderChain` to rank models by configurable strategies defined in [`shared/types.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/shared/types.ts) and scored via [`server/src/services/scoring.js`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.js).
- **Exploration logic** ensures new models receive traffic for performance sampling, while **sticky-session pinning** respects client preferences via `preferredModelDbId`.
- **Capability validation** filters by vision, tools, structured output, context window fit, and token budgets before attempting key selection.
- **Key selection** via `selectKeyForModel` enforces rate limits through [`server/src/services/ratelimit.js`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.js) and quota limits through [`server/src/services/provider-quota.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/provider-quota.ts).
- **Failure diagnostics** aggregate exhaustion reasons through `summarizeExhaustion` to guide operational responses.

## Frequently Asked Questions

### How does FreeLLMAPI balance load across multiple providers?

FreeLLMAPI applies the configured **routing strategy**—bandit-based, balanced, or manual priority—via the `orderChain` function. Bandit-based strategies use historical performance data from [`server/src/services/scoring.js`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/scoring.js) to weight models, while balanced mode distributes requests evenly. The system dynamically reorders the candidate chain for each request based on real-time capacity and performance metrics.

### What happens when no providers can handle a request?

When all candidates fail validation or key selection, the router throws a `RouteError` constructed by `summarizeExhaustion`. This error categorizes the failure reason—such as rate-limiting, exhausted quotas, or oversized context windows—and suggests specific remediation steps like adding new API keys or waiting for rate-limit resets.

### Can I force specific models for certain request types?

Yes. You can enforce **sticky sessions** by passing a `preferredModelDbId` to `routeRequest`, which moves that model to the front of the candidate chain. Additionally, you can filter by capabilities (`requireVision`, `requireTools`, `requireStructured`) or exclude entire providers using the `skipPlatforms` parameter to guarantee specific routing behavior.

### How does the router handle new models with no historical data?

Through the **exploration step** in bandit-based routing mode. When enabled, the router randomly promotes *unmeasured* models—those lacking reliability or speed samples—to the front of the chain with a certain probability. This ensures new integrations collect performance data without requiring manual traffic allocation or compromising system reliability.