# How FreeLLMAPI Manages Sticky Sessions for Consistent Conversations: Complete Technical Guide

> Discover how FreeLLMAPI ensures consistent conversations with sticky sessions. Learn about its in-memory map, 30-minute TTL, and automatic cleanup in this technical guide.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-30

---

**FreeLLMAPI keeps multi-turn conversations anchored to the same LLM model using an in-memory sticky-session map with 30-minute TTL and automatic cleanup.**

FreeLLMAPI is an open-source API router that intelligently distributes requests across multiple language models. One critical feature is its **sticky session** mechanism—ensuring that once a conversation starts with a particular model, subsequent turns route to that same model rather than switching mid-conversation (which can cause jarring behavior or hallucinations). This article explains exactly how the implementation works, with full source references from `tashfeenahmed/freellmapi`.

## The Core Sticky Session Architecture

The sticky session system lives primarily in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts). It uses a simple but effective design: a JavaScript `Map` that associates conversation identifiers with the model that handled them.

### In-Memory Store and TTL Configuration

The foundational data structure is declared at the top of the proxy route file:

```ts
const stickySessionMap = new Map<string, { modelDbId: number; lastUsed: number }>();
const STICKY_TTL_MS = 30 * 60 * 1000; // 30 min session TTL

```

- **`modelDbId`**: The database identifier of the LLM model used
- **`lastUsed`**: Unix timestamp for TTL enforcement and LRU eviction
- **TTL**: Sessions expire after **30 minutes of inactivity**

## How Session Keys Are Generated

FreeLLMAPI creates deterministic session keys through `getSessionKey` (lines 33‑44 in [`proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/proxy.ts)). The function supports two identification strategies:

### Strategy 1: Explicit Session ID Header

If the client sends an `X-Session-Id` header, that value forms the key—optionally combined with a strategy key for multi-tenant scenarios:

```ts
function getSessionKey(messages, sessionIdHeader?, strategyKey?) {
  if (sessionIdHeader) {
    return strategyKey 
      ? `hdr:${sessionIdHeader}::${strategyKey}` 
      : `hdr:${sessionIdHeader}`;
  }
  // ... fallback to content hash

```

### Strategy 2: Content-Based Hash

Without a header, FreeLLMAPI hashes the first user message using SHA-1:

```ts
  const firstUser = messages.find(m => m.role === 'user');
  if (!firstUser) return '';
  const text = contentToString(firstUser.content ?? '');
  if (!text) return '';
  const payload = strategyKey ? `${text}::${strategyKey}` : text;
  return crypto.createHash('sha1').update(payload).digest('hex');
}

```

This allows stateless clients to still benefit from sticky sessions without managing session IDs.

## Reading Sticky Session Entries

The `getStickyModel` function (lines 46‑61) retrieves the stored model for a conversation:

```ts
export function getStickyModel(messages, sessionIdHeader?, strategyKey?) {
  const hasAssistant = messages.some(m => m.role === 'assistant');
  if (!hasAssistant) return undefined;
  
  const key = getSessionKey(messages, sessionIdHeader, strategyKey);
  if (!key) return undefined;
  
  const entry = stickySessionMap.get(key);
  if (!entry) return undefined;
  
  if (Date.now() - entry.lastUsed > STICKY_TTL_MS) {
    stickySessionMap.delete(key);
    return undefined;
  }
  return entry.modelDbId;
}

```

**Critical behavior**: Sticky sessions **only activate after the first assistant response**. The `hasAssistant` check prevents premature binding—ensuring the system doesn't lock to a model until a conversation actually begins.

## Recording and Updating Sessions

After a successful model response, `setStickyModel` (lines 63‑74) persists the binding:

```ts
export function setStickyModel(messages, modelDbId, sessionIdHeader?, strategyKey?) {
  const key = getSessionKey(messages, sessionIdHeader, strategyKey);
  if (!key) return;
  
  stickySessionMap.set(key, { modelDbId, lastUsed: Date.now() });
  // Periodic cleanup triggered here...
}

```

The `lastUsed` timestamp refreshes on every interaction, extending the session's lifetime.

## Automatic Memory Management

FreeLLMAPI implements two-layer cleanup to prevent unbounded memory growth (lines 68‑84):

| Trigger | Action |
|--------|--------|
| Map exceeds 500 entries | Remove all entries older than 30-minute TTL |
| Still exceeds 1000 entries | Evict oldest entries by `lastUsed` (LRU) |

This bounded approach handles high-throughput deployments without manual intervention.

## Sticky Sessions in the Request Flow

The sticky mechanism integrates at two critical points in the routing pipeline.

### Setting the Sticky Model After Response

In the main proxy handler, successful completions trigger `setStickyModel` to record the model for future turns.

### Prioritizing Sticky Models During Routing

In [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) (lines 695‑702), the routing logic consults sticky sessions before auto-routing:

```ts
const sticky = getStickyModel(messages, sessionIdHeader, requestedModelLabel);
preferredModel = (sticky != null && groupChain.some(r => r.model_db_id === sticky))
                 ? sticky
                 : undefined;

```

The `groupChain.some()` verification ensures the sticky model is still available in the current routing group—gracefully falling back to auto-selection if the model was removed or disabled.

## Client Usage Example

Send requests with `X-Session-Id` to leverage explicit session tracking:

```ts
import fetch from 'node-fetch';

await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer YOUR_API_KEY',
    'Content-Type': 'application/json',
    'X-Session-Id': 'user-1234',   // Sticky session identifier
  },
  body: JSON.stringify({
    messages: [{ role: 'user', content: 'What is the weather today?' }],
    model: 'auto',                  // Router picks initial model
  }),
});

```

Subsequent requests with the same `X-Session-Id` will route to the same model.

## Implementation Files Reference

| File | Purpose |
|------|---------|
| [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) | Core sticky-session implementation: map, TTL, key generation, get/set functions, cleanup |
| [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) | Reads sticky model to override auto-routing |
| [`server/src/lib/inbound-chat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/inbound-chat.ts) | Uses `getStickyModel` for inbound request construction |
| [`server/src/__tests__/routes/proxy-stream-integrity.test.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/__tests__/routes/proxy-stream-integrity.test.ts) | Test coverage for sticky-session behavior |

## Summary

- **Sticky sessions** keep conversations on the same LLM model across multiple turns, preventing mid-conversation switches
- **30-minute TTL** with automatic expiration balances consistency with resource cleanup
- **Dual key strategies** support both explicit session IDs (`X-Session-Id` header) and content-based hashing for stateless clients
- **Activation threshold**: Sessions only stick after the first assistant response appears in the message history
- **Bounded memory**: 500-entry soft limit with LRU eviction at 1000 entries prevents memory leaks
- **Graceful degradation**: Sticky models are verified against available models before enforcement

## Frequently Asked Questions

### What happens if the sticky model becomes unavailable?

FreeLLMAPI validates the sticky model against the current `groupChain` before using it. If the model was removed or disabled, the router falls back to standard auto-routing—the conversation continues seamlessly with a new model selection.

### Can multiple conversations share the same session ID?

Yes, but only intentionally. Using the same `X-Session-Id` across different conversation threads binds them to the same model. For isolation, use unique session IDs per conversation thread.

### How does FreeLLMAPI handle sticky sessions across server restarts?

The current implementation uses an in-memory `Map` without persistence. Server restarts clear all sticky sessions. For distributed deployments, sessions are node-local—clients should implement sticky load balancing at the infrastructure layer or use deterministic content hashing.

### What's the difference between using `X-Session-Id` and content-based hashing?

`X-Session-Id` provides explicit control and works for any message content, including identical prompts from different users. Content-based hashing requires no client-side state management but collapses identical opening messages into the same session—appropriate when conversation context is fully contained in the message history.