What Does the Router Service Do in FreeLLMAPI? Core Responsibilities Explained

The Router Service in FreeLLMAPI is the core decision-making component that selects models, manages API keys, handles automatic failover, and enforces rate limits to deliver reliable free-tier LLM aggregation.

The Router Service sits at the heart of the FreeLLMAPI open-source project, transforming scattered free-tier LLM offerings into a single, dependable OpenAI-compatible endpoint. Understanding what the Router Service does in FreeLLMAPI reveals how the system achieves high availability without requiring paid API credits.

Model Selection and Key Management

Selecting the Optimal Model for Each Request

For every incoming request, the Router Service scans the catalog of free-tier models and selects the highest-priority candidate that meets two strict criteria: the model must have a healthy API key and must remain under all of its rate-limit caps. This selection process is implemented in server/src/services/router.ts.

Once a model/key pair is chosen, the router decrypts the stored key in memory and forwards the request to the provider's SDK. This decryption happens at request time—keys are never persisted in plaintext, and decrypted values exist only transiently during the request lifecycle.

import { routeRequest } from '../services/router.js';

router.post('/v1/chat/completions', async (req, res, next) => {
  // The router decides which model/key to use and forwards the request.
  const result = await routeRequest(req.body);
  res.json(result);
});

The routeRequest function encapsulates the entire decision pipeline: evaluation, selection, decryption, and dispatch.

Automatic Failover and Retry Logic

Handling Provider Failures Gracefully

When a provider returns a 429 (rate limited), 5xx (server error), or times out, the Router Service does not fail the request. Instead, it:

  • Puts the failed key on a short cooldown period
  • Retries the next model in the fallback chain
  • Continues this process up to approximately 20 attempts
  • Respects a wall-clock retry budget to prevent unbounded delays

This failover mechanism ensures that transient provider issues do not propagate to end users.

Rate-Limit Tracking and Enforcement

Per-Key Quota Management

The Router Service consults the Rate-limit ledger (server/src/services/ratelimit.ts) before every selection. This ledger maintains per-key counters for:

Counter Meaning
RPM Requests Per Minute
RPD Requests Per Day
TPM Tokens Per Minute
TPD Tokens Per Day

These counters are stored in-memory with SQLite backing, providing fast quota checks that prevent the router from selecting a key that would exceed its provider-imposed limits.

Health Monitoring Integration

Key Status and Availability

A background Health service (server/src/services/health.ts) continuously probes API keys and assigns one of four statuses:

  • healthy — available for selection
  • rate_limited — temporarily throttled by provider
  • invalid — key rejected during probe
  • error — unexpected failure during check

The Router Service only considers healthy keys when making selections. This health data propagates automatically, ensuring the router adapts to changing provider conditions without manual intervention.

Strategy-Driven Routing Decisions

Six Built-In Routing Strategies

The FreeLLMAPI Router Service implements six configurable strategies that weight different optimization goals:

Strategy Optimization Focus
priority Strict priority list order
balanced Even distribution across healthy keys
smartest Highest capability models
fastest Lowest latency providers
reliable Lowest recent error rate
custom User-defined weighting

Scores are calculated using a Thompson-sampling bandit algorithm, allowing the router to explore new options while exploiting known-good performers.

Sticky Sessions and Context Preservation

Multi-Turn Conversation Handling

For conversational use cases, the Router Service attempts to maintain the same model for up to 30 minutes within a session. This reduces "hallucination spikes"—inconsistencies that can occur when switching models mid-conversation due to differing training data and behavior patterns.

Client Integration: Using the Router Service

Clients interact with the Router Service through a unified OpenAI-compatible endpoint without awareness of underlying providers.

cURL Example


# Let the router automatically select the best available model

curl -X POST http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-<your-token>" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "auto",
        "messages": [{"role":"user","content":"Explain quantum computing"}]
      }'

OpenAI SDK Example

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",
  apiKey: "freellmapi-<your-token>",
});

const resp = await client.chat.completions.create({
  model: "auto",  // Delegates selection to Router Service
  messages: [{ role: "user", content: "What does the router do?" }],
});

Setting model: "auto" triggers the full Router Service pipeline—strategy evaluation, health filtering, rate-limit checking, and potential failover.

Key Implementation Files

File Role
server/src/services/router.ts Core routing logic: selection, decryption, failover
server/src/services/ratelimit.ts In-memory quota counters with SQLite persistence
server/src/services/health.ts Background health probes and status management
server/src/providers/*.ts Provider adapters implementing common Provider interface
server/src/routes/proxy.ts HTTP entry point forwarding to routeRequest

These files collectively implement the Router Service architecture described in docs/architecture.md.

Summary

The Router Service in FreeLLMAPI fulfills seven critical responsibilities:

  • Model selection — chooses highest-priority healthy, under-quota models
  • Key decryption — handles secure in-memory credential management
  • Automatic failover — retries across ~20 alternatives with bounded budgets
  • Rate-limit enforcement — consults per-key RPM/RPD/TPM/TPD counters
  • Health awareness — integrates with background probe results
  • Strategy optimization — applies Thompson-sampling bandit scoring across six strategies
  • Session stickiness — preserves model continuity for multi-turn conversations

Clients receive a single, reliable OpenAI-compatible endpoint while the Router Service orchestrates the complexity of free-tier aggregation behind the scenes.

Frequently Asked Questions

How does the Router Service handle rate limits from multiple providers?

The Router Service queries the Rate-limit ledger (server/src/services/ratelimit.ts) before each selection. This service tracks per-key RPM, RPD, TPM, and TPD counters in real-time. Only keys with remaining quota headroom are considered eligible, ensuring requests never trigger provider-side throttling under normal operation.

What happens when all available keys are rate limited or unhealthy?

The Router Service implements bounded retry logic with approximately 20 attempts and a wall-clock timeout budget. If all keys exhaust their quota or fail health checks within this budget, the request eventually returns an error. However, the combination of multiple providers, automatic failover, and health-driven cooldowns makes total exhaustion rare in practice.

Can I force the Router Service to use a specific model instead of auto-selection?

Yes. While model: "auto" triggers the full Router Service pipeline, clients may specify exact model identifiers (e.g., model: "gpt-3.5-turbo") to bypass automatic selection. The router will still validate that the specified model has healthy keys and available quota before proceeding.

How does the Router Service maintain conversation context across requests?

The Router Service implements sticky sessions that attempt to reuse the same model for up to 30 minutes within a conversation identifier. This is implemented in server/src/services/router.ts to reduce "hallucination spikes" caused by switching between models with different training behaviors mid-conversation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →