How Sticky Session Routing Keeps Conversations on One Model for 30 Minutes

FreeLLMAPI uses an in‑process stickySessionMap with a 30‑minute time‑to‑live (TTL) to cache the last model that served a conversation, ensuring subsequent requests in the same session route to the same backend model.

Sticky session routing prevents jarring model switches mid‑conversation in multi‑provider LLM gateways. In the tashfeenahmed/freellmapi repository, this mechanism locks a conversation to a single model for up to 30 minutes of inactivity, preserving context and preventing hallucinations caused by switching backends between turns.

How Sticky Session Routing Works

The implementation relies on a lightweight, in‑memory cache that associates a conversation fingerprint with the database ID of the last successful model.

Session Key Generation

Every incoming request that contains an assistant turn triggers the generation of a unique session identifier. In server/src/routes/proxy.ts (lines 33–44), the getSessionKey() function constructs this key by hashing the first user message—or the x‑session‑id header if provided—combined with an optional strategy key (e.g., the routing strategy being used). The function returns a SHA‑1 hash of this payload, ensuring a consistent lookup key across distributed requests.

const key = getSessionKey(messages, sessionIdHeader, strategyKey);

The Sticky Session Cache

A global Map named stickySessionMap stores the mapping between session keys and model assignments. Defined in server/src/routes/proxy.ts (lines 151–153), this map holds objects structured as {modelDbId: number; lastUsed: number}. The lastUsed timestamp is critical: it is compared against the constant STICKY_TTL_MS (set to 30 * 60 * 1000 milliseconds) to determine if the entry is still valid.

Reading and Writing Sticky Entries

Two primary functions manage the cache lifecycle:

  • getStickyModel() (lines 46–60): Queries the map for the generated key. If the entry is missing or the lastUsed age exceeds the 30‑minute TTL, the entry is deleted and undefined is returned, forcing normal auto‑routing.
  • setStickyModel() (lines 63–85): Called after a successful provider response, this function writes the modelDbId to the map with a fresh lastUsed timestamp, extending the sticky window for another 30 minutes.

Implementation in the FreeLLMAPI Source Code

The sticky logic is tightly integrated into the request lifecycle. When handling a chat completion request in server/src/routes/responses.ts (lines 690–702), the router first invokes getStickyModel(). If a valid sticky entry exists, that modelDbId is passed as the preferredModel argument to routeRequest(), forcing the routing chain to attempt that model first.

// Check for an existing sticky assignment
const stickyId = getStickyModel(messages, req.headers['x-session-id'], 'auto');

if (stickyId) {
  // Force the router to try the sticky model first
  preferredModel = stickyId;
}

// ... provider call succeeds ...

// Pin the model for the next turn
setStickyModel(messages, chosenModelId, req.headers['x-session-id'], 'auto');

TTL and Automatic Cleanup

The 30‑minute TTL is enforced by comparing Date.now() against the stored lastUsed timestamp on every read operation. Additionally, setStickyModel() implements a periodic cleanup to prevent unbounded memory growth:

  1. If the map exceeds 500 entries, the function first scans for and removes any expired entries.
  2. If the size remains above 1000 entries after expiration cleanup, it evicts the oldest entries (by lastUsed) until the map shrinks to 1000.

This dual mechanism ensures that inactive sessions automatically expire while active sessions can persist beyond the initial 30 minutes as long as the conversation continues, because each successful turn refreshes the lastUsed timestamp.

Practical Usage Example

Clients opt into sticky sessions by including a consistent identifier in the request headers. The following flow demonstrates a multi‑turn conversation remaining pinned to model ID 42:

// Turn 1: New session
POST /v1/chat/completions
Headers: { "x-session-id": "conv-123" }
Body: { "messages": [{"role": "user", "content": "Hello"}] }

// Server generates key, finds no sticky entry, routes normally
// After success: setStickyModel caches modelDbId: 42

// Turn 2: Within 30 minutes
POST /v1/chat/completions
Headers: { "x-session-id": "conv-123" }

// Server finds key, sees lastUsed < 30 min ago, returns modelDbId: 42
// RouteRequest prioritizes model 42

Summary

  • 30‑minute TTL: The STICKY_TTL_MS constant defines a 30‑minute inactivity window; stale entries are purged on read.
  • SHA‑1 session keys: Keys are derived from the first user message or x‑session‑id header plus strategy salt, ensuring unique session isolation.
  • Automatic memory management: The cache self‑cleans at 500 and 1000 entries to prevent leaks in high‑traffic deployments.
  • Graceful fallback: If the sticky model is disabled or the TTL expires, the router falls back to standard load‑balancing logic.

Frequently Asked Questions

What happens when the 30‑minute sticky session TTL expires?

Once the lastUsed timestamp is older than STICKY_TTL_MS (30 minutes), getStickyModel() deletes the entry and returns undefined. The router then treats the request as a new session, applying standard auto‑routing logic to select the best available model rather than forcing a specific backend.

How does FreeLLMAPI identify a conversation session?

The system generates a session key by SHA‑1 hashing either the content of the first user message in the messages array or the value of the x‑session‑id HTTP header, concatenated with an optional strategy identifier. This allows clients to maintain continuity across requests by sending a consistent x‑session‑id header.

What occurs if the sticky model becomes unavailable?

If getStickyModel() returns a valid ID but that model is disabled or returns an error during the routing phase, the routeRequest() chain treats the failure as a standard provider error and falls back to the next available model in the priority list. The sticky entry itself is only updated or confirmed after a successful provider response.

How does the system prevent memory leaks in the sticky session cache?

setStickyModel() incorporates a two‑tier eviction policy: when the map size exceeds 500, it proactively deletes expired entries; if the size surpasses 1000, it removes the oldest entries (by lastUsed) until the count drops to 1000. This bounds memory usage regardless of traffic volume.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →