# How Sticky Session Routing Keeps Conversations on One Model for 30 Minutes

> Learn how FreeLLMAPI uses sticky session routing with a 30-minute TTL to ensure conversations consistently connect to the same backend model for uninterrupted AI interaction.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-08-31

---

**FreeLLMAPI uses an in‑process `stickySessionMap` with a 30‑minute time‑to‑live (TTL) to cache the last model that served a conversation, ensuring subsequent requests in the same session route to the same backend model.**

**Sticky session routing** prevents jarring model switches mid‑conversation in multi‑provider LLM gateways. In the `tashfeenahmed/freellmapi` repository, this mechanism locks a conversation to a single model for up to 30 minutes of inactivity, preserving context and preventing hallucinations caused by switching backends between turns.

## How Sticky Session Routing Works

The implementation relies on a lightweight, in‑memory cache that associates a conversation fingerprint with the database ID of the last successful model.

### Session Key Generation

Every incoming request that contains an assistant turn triggers the generation of a unique session identifier. In [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) (lines 33–44), the `getSessionKey()` function constructs this key by hashing the first user message—or the `x‑session‑id` header if provided—combined with an optional strategy key (e.g., the routing strategy being used). The function returns a **SHA‑1 hash** of this payload, ensuring a consistent lookup key across distributed requests.

```typescript
const key = getSessionKey(messages, sessionIdHeader, strategyKey);

```

### The Sticky Session Cache

A global `Map` named **`stickySessionMap`** stores the mapping between session keys and model assignments. Defined in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) (lines 151–153), this map holds objects structured as `{modelDbId: number; lastUsed: number}`. The **`lastUsed`** timestamp is critical: it is compared against the constant `STICKY_TTL_MS` (set to `30 * 60 * 1000` milliseconds) to determine if the entry is still valid.

### Reading and Writing Sticky Entries

Two primary functions manage the cache lifecycle:

- **`getStickyModel()`** (lines 46–60): Queries the map for the generated key. If the entry is missing or the `lastUsed` age exceeds the 30‑minute TTL, the entry is deleted and `undefined` is returned, forcing normal auto‑routing.
- **`setStickyModel()`** (lines 63–85): Called after a successful provider response, this function writes the `modelDbId` to the map with a fresh `lastUsed` timestamp, extending the sticky window for another 30 minutes.

## Implementation in the FreeLLMAPI Source Code

The sticky logic is tightly integrated into the request lifecycle. When handling a chat completion request in [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) (lines 690–702), the router first invokes `getStickyModel()`. If a valid sticky entry exists, that `modelDbId` is passed as the `preferredModel` argument to `routeRequest()`, forcing the routing chain to attempt that model first.

```typescript
// Check for an existing sticky assignment
const stickyId = getStickyModel(messages, req.headers['x-session-id'], 'auto');

if (stickyId) {
  // Force the router to try the sticky model first
  preferredModel = stickyId;
}

// ... provider call succeeds ...

// Pin the model for the next turn
setStickyModel(messages, chosenModelId, req.headers['x-session-id'], 'auto');

```

## TTL and Automatic Cleanup

The **30‑minute TTL** is enforced by comparing `Date.now()` against the stored `lastUsed` timestamp on every read operation. Additionally, `setStickyModel()` implements a periodic cleanup to prevent unbounded memory growth:

1. If the map exceeds **500 entries**, the function first scans for and removes any expired entries.
2. If the size remains above **1000 entries** after expiration cleanup, it evicts the oldest entries (by `lastUsed`) until the map shrinks to 1000.

This dual mechanism ensures that inactive sessions automatically expire while active sessions can persist beyond the initial 30 minutes as long as the conversation continues, because each successful turn refreshes the `lastUsed` timestamp.

## Practical Usage Example

Clients opt into sticky sessions by including a consistent identifier in the request headers. The following flow demonstrates a multi‑turn conversation remaining pinned to model ID `42`:

```typescript
// Turn 1: New session
POST /v1/chat/completions
Headers: { "x-session-id": "conv-123" }
Body: { "messages": [{"role": "user", "content": "Hello"}] }

// Server generates key, finds no sticky entry, routes normally
// After success: setStickyModel caches modelDbId: 42

// Turn 2: Within 30 minutes
POST /v1/chat/completions
Headers: { "x-session-id": "conv-123" }

// Server finds key, sees lastUsed < 30 min ago, returns modelDbId: 42
// RouteRequest prioritizes model 42

```

## Summary

- **30‑minute TTL**: The `STICKY_TTL_MS` constant defines a 30‑minute inactivity window; stale entries are purged on read.
- **SHA‑1 session keys**: Keys are derived from the first user message or `x‑session‑id` header plus strategy salt, ensuring unique session isolation.
- **Automatic memory management**: The cache self‑cleans at 500 and 1000 entries to prevent leaks in high‑traffic deployments.
- **Graceful fallback**: If the sticky model is disabled or the TTL expires, the router falls back to standard load‑balancing logic.

## Frequently Asked Questions

### What happens when the 30‑minute sticky session TTL expires?

Once the `lastUsed` timestamp is older than `STICKY_TTL_MS` (30 minutes), `getStickyModel()` deletes the entry and returns `undefined`. The router then treats the request as a new session, applying standard auto‑routing logic to select the best available model rather than forcing a specific backend.

### How does FreeLLMAPI identify a conversation session?

The system generates a session key by SHA‑1 hashing either the content of the first user message in the messages array or the value of the `x‑session‑id` HTTP header, concatenated with an optional strategy identifier. This allows clients to maintain continuity across requests by sending a consistent `x‑session‑id` header.

### What occurs if the sticky model becomes unavailable?

If `getStickyModel()` returns a valid ID but that model is disabled or returns an error during the routing phase, the `routeRequest()` chain treats the failure as a standard provider error and falls back to the next available model in the priority list. The sticky entry itself is only updated or confirmed after a successful provider response.

### How does the system prevent memory leaks in the sticky session cache?

`setStickyModel()` incorporates a two‑tier eviction policy: when the map size exceeds 500, it proactively deletes expired entries; if the size surpasses 1000, it removes the oldest entries (by `lastUsed`) until the count drops to 1000. This bounds memory usage regardless of traffic volume.