What Does the Router Service Do in FreeLLMAPI? Core Responsibilities Explained
The Router Service in FreeLLMAPI is the core decision-making component that selects models, manages API keys, handles automatic failover, and enforces rate limits to deliver reliable free-tier LLM aggregation.
The Router Service sits at the heart of the FreeLLMAPI open-source project, transforming scattered free-tier LLM offerings into a single, dependable OpenAI-compatible endpoint. Understanding what the Router Service does in FreeLLMAPI reveals how the system achieves high availability without requiring paid API credits.
Model Selection and Key Management
Selecting the Optimal Model for Each Request
For every incoming request, the Router Service scans the catalog of free-tier models and selects the highest-priority candidate that meets two strict criteria: the model must have a healthy API key and must remain under all of its rate-limit caps. This selection process is implemented in server/src/services/router.ts.
Once a model/key pair is chosen, the router decrypts the stored key in memory and forwards the request to the provider's SDK. This decryption happens at request time—keys are never persisted in plaintext, and decrypted values exist only transiently during the request lifecycle.
import { routeRequest } from '../services/router.js';
router.post('/v1/chat/completions', async (req, res, next) => {
// The router decides which model/key to use and forwards the request.
const result = await routeRequest(req.body);
res.json(result);
});
The routeRequest function encapsulates the entire decision pipeline: evaluation, selection, decryption, and dispatch.
Automatic Failover and Retry Logic
Handling Provider Failures Gracefully
When a provider returns a 429 (rate limited), 5xx (server error), or times out, the Router Service does not fail the request. Instead, it:
- Puts the failed key on a short cooldown period
- Retries the next model in the fallback chain
- Continues this process up to approximately 20 attempts
- Respects a wall-clock retry budget to prevent unbounded delays
This failover mechanism ensures that transient provider issues do not propagate to end users.
Rate-Limit Tracking and Enforcement
Per-Key Quota Management
The Router Service consults the Rate-limit ledger (server/src/services/ratelimit.ts) before every selection. This ledger maintains per-key counters for:
| Counter | Meaning |
|---|---|
| RPM | Requests Per Minute |
| RPD | Requests Per Day |
| TPM | Tokens Per Minute |
| TPD | Tokens Per Day |
These counters are stored in-memory with SQLite backing, providing fast quota checks that prevent the router from selecting a key that would exceed its provider-imposed limits.
Health Monitoring Integration
Key Status and Availability
A background Health service (server/src/services/health.ts) continuously probes API keys and assigns one of four statuses:
healthy— available for selectionrate_limited— temporarily throttled by providerinvalid— key rejected during probeerror— unexpected failure during check
The Router Service only considers healthy keys when making selections. This health data propagates automatically, ensuring the router adapts to changing provider conditions without manual intervention.
Strategy-Driven Routing Decisions
Six Built-In Routing Strategies
The FreeLLMAPI Router Service implements six configurable strategies that weight different optimization goals:
| Strategy | Optimization Focus |
|---|---|
priority |
Strict priority list order |
balanced |
Even distribution across healthy keys |
smartest |
Highest capability models |
fastest |
Lowest latency providers |
reliable |
Lowest recent error rate |
custom |
User-defined weighting |
Scores are calculated using a Thompson-sampling bandit algorithm, allowing the router to explore new options while exploiting known-good performers.
Sticky Sessions and Context Preservation
Multi-Turn Conversation Handling
For conversational use cases, the Router Service attempts to maintain the same model for up to 30 minutes within a session. This reduces "hallucination spikes"—inconsistencies that can occur when switching models mid-conversation due to differing training data and behavior patterns.
Client Integration: Using the Router Service
Clients interact with the Router Service through a unified OpenAI-compatible endpoint without awareness of underlying providers.
cURL Example
# Let the router automatically select the best available model
curl -X POST http://localhost:3001/v1/chat/completions \
-H "Authorization: Bearer freellmapi-<your-token>" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [{"role":"user","content":"Explain quantum computing"}]
}'
OpenAI SDK Example
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "http://localhost:3001/v1",
apiKey: "freellmapi-<your-token>",
});
const resp = await client.chat.completions.create({
model: "auto", // Delegates selection to Router Service
messages: [{ role: "user", content: "What does the router do?" }],
});
Setting model: "auto" triggers the full Router Service pipeline—strategy evaluation, health filtering, rate-limit checking, and potential failover.
Key Implementation Files
| File | Role |
|---|---|
server/src/services/router.ts |
Core routing logic: selection, decryption, failover |
server/src/services/ratelimit.ts |
In-memory quota counters with SQLite persistence |
server/src/services/health.ts |
Background health probes and status management |
server/src/providers/*.ts |
Provider adapters implementing common Provider interface |
server/src/routes/proxy.ts |
HTTP entry point forwarding to routeRequest |
These files collectively implement the Router Service architecture described in docs/architecture.md.
Summary
The Router Service in FreeLLMAPI fulfills seven critical responsibilities:
- Model selection — chooses highest-priority healthy, under-quota models
- Key decryption — handles secure in-memory credential management
- Automatic failover — retries across ~20 alternatives with bounded budgets
- Rate-limit enforcement — consults per-key RPM/RPD/TPM/TPD counters
- Health awareness — integrates with background probe results
- Strategy optimization — applies Thompson-sampling bandit scoring across six strategies
- Session stickiness — preserves model continuity for multi-turn conversations
Clients receive a single, reliable OpenAI-compatible endpoint while the Router Service orchestrates the complexity of free-tier aggregation behind the scenes.
Frequently Asked Questions
How does the Router Service handle rate limits from multiple providers?
The Router Service queries the Rate-limit ledger (server/src/services/ratelimit.ts) before each selection. This service tracks per-key RPM, RPD, TPM, and TPD counters in real-time. Only keys with remaining quota headroom are considered eligible, ensuring requests never trigger provider-side throttling under normal operation.
What happens when all available keys are rate limited or unhealthy?
The Router Service implements bounded retry logic with approximately 20 attempts and a wall-clock timeout budget. If all keys exhaust their quota or fail health checks within this budget, the request eventually returns an error. However, the combination of multiple providers, automatic failover, and health-driven cooldowns makes total exhaustion rare in practice.
Can I force the Router Service to use a specific model instead of auto-selection?
Yes. While model: "auto" triggers the full Router Service pipeline, clients may specify exact model identifiers (e.g., model: "gpt-3.5-turbo") to bypass automatic selection. The router will still validate that the specified model has healthy keys and available quota before proceeding.
How does the Router Service maintain conversation context across requests?
The Router Service implements sticky sessions that attempt to reuse the same model for up to 30 minutes within a conversation identifier. This is implemented in server/src/services/router.ts to reduce "hallucination spikes" caused by switching between models with different training behaviors mid-conversation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →