How FreeLLMAPI Manages Sticky Sessions for Consistent Conversations: Complete Technical Guide

FreeLLMAPI keeps multi-turn conversations anchored to the same LLM model using an in-memory sticky-session map with 30-minute TTL and automatic cleanup.

FreeLLMAPI is an open-source API router that intelligently distributes requests across multiple language models. One critical feature is its sticky session mechanism—ensuring that once a conversation starts with a particular model, subsequent turns route to that same model rather than switching mid-conversation (which can cause jarring behavior or hallucinations). This article explains exactly how the implementation works, with full source references from tashfeenahmed/freellmapi.

The Core Sticky Session Architecture

The sticky session system lives primarily in server/src/routes/proxy.ts. It uses a simple but effective design: a JavaScript Map that associates conversation identifiers with the model that handled them.

In-Memory Store and TTL Configuration

The foundational data structure is declared at the top of the proxy route file:

const stickySessionMap = new Map<string, { modelDbId: number; lastUsed: number }>();
const STICKY_TTL_MS = 30 * 60 * 1000; // 30 min session TTL
  • modelDbId: The database identifier of the LLM model used
  • lastUsed: Unix timestamp for TTL enforcement and LRU eviction
  • TTL: Sessions expire after 30 minutes of inactivity

How Session Keys Are Generated

FreeLLMAPI creates deterministic session keys through getSessionKey (lines 33‑44 in proxy.ts). The function supports two identification strategies:

Strategy 1: Explicit Session ID Header

If the client sends an X-Session-Id header, that value forms the key—optionally combined with a strategy key for multi-tenant scenarios:

function getSessionKey(messages, sessionIdHeader?, strategyKey?) {
  if (sessionIdHeader) {
    return strategyKey 
      ? `hdr:${sessionIdHeader}::${strategyKey}` 
      : `hdr:${sessionIdHeader}`;
  }
  // ... fallback to content hash

Strategy 2: Content-Based Hash

Without a header, FreeLLMAPI hashes the first user message using SHA-1:

  const firstUser = messages.find(m => m.role === 'user');
  if (!firstUser) return '';
  const text = contentToString(firstUser.content ?? '');
  if (!text) return '';
  const payload = strategyKey ? `${text}::${strategyKey}` : text;
  return crypto.createHash('sha1').update(payload).digest('hex');
}

This allows stateless clients to still benefit from sticky sessions without managing session IDs.

Reading Sticky Session Entries

The getStickyModel function (lines 46‑61) retrieves the stored model for a conversation:

export function getStickyModel(messages, sessionIdHeader?, strategyKey?) {
  const hasAssistant = messages.some(m => m.role === 'assistant');
  if (!hasAssistant) return undefined;
  
  const key = getSessionKey(messages, sessionIdHeader, strategyKey);
  if (!key) return undefined;
  
  const entry = stickySessionMap.get(key);
  if (!entry) return undefined;
  
  if (Date.now() - entry.lastUsed > STICKY_TTL_MS) {
    stickySessionMap.delete(key);
    return undefined;
  }
  return entry.modelDbId;
}

Critical behavior: Sticky sessions only activate after the first assistant response. The hasAssistant check prevents premature binding—ensuring the system doesn't lock to a model until a conversation actually begins.

Recording and Updating Sessions

After a successful model response, setStickyModel (lines 63‑74) persists the binding:

export function setStickyModel(messages, modelDbId, sessionIdHeader?, strategyKey?) {
  const key = getSessionKey(messages, sessionIdHeader, strategyKey);
  if (!key) return;
  
  stickySessionMap.set(key, { modelDbId, lastUsed: Date.now() });
  // Periodic cleanup triggered here...
}

The lastUsed timestamp refreshes on every interaction, extending the session's lifetime.

Automatic Memory Management

FreeLLMAPI implements two-layer cleanup to prevent unbounded memory growth (lines 68‑84):

Trigger Action
Map exceeds 500 entries Remove all entries older than 30-minute TTL
Still exceeds 1000 entries Evict oldest entries by lastUsed (LRU)

This bounded approach handles high-throughput deployments without manual intervention.

Sticky Sessions in the Request Flow

The sticky mechanism integrates at two critical points in the routing pipeline.

Setting the Sticky Model After Response

In the main proxy handler, successful completions trigger setStickyModel to record the model for future turns.

Prioritizing Sticky Models During Routing

In server/src/routes/responses.ts (lines 695‑702), the routing logic consults sticky sessions before auto-routing:

const sticky = getStickyModel(messages, sessionIdHeader, requestedModelLabel);
preferredModel = (sticky != null && groupChain.some(r => r.model_db_id === sticky))
                 ? sticky
                 : undefined;

The groupChain.some() verification ensures the sticky model is still available in the current routing group—gracefully falling back to auto-selection if the model was removed or disabled.

Client Usage Example

Send requests with X-Session-Id to leverage explicit session tracking:

import fetch from 'node-fetch';

await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer YOUR_API_KEY',
    'Content-Type': 'application/json',
    'X-Session-Id': 'user-1234',   // Sticky session identifier
  },
  body: JSON.stringify({
    messages: [{ role: 'user', content: 'What is the weather today?' }],
    model: 'auto',                  // Router picks initial model
  }),
});

Subsequent requests with the same X-Session-Id will route to the same model.

Implementation Files Reference

File Purpose
server/src/routes/proxy.ts Core sticky-session implementation: map, TTL, key generation, get/set functions, cleanup
server/src/routes/responses.ts Reads sticky model to override auto-routing
server/src/lib/inbound-chat.ts Uses getStickyModel for inbound request construction
server/src/__tests__/routes/proxy-stream-integrity.test.ts Test coverage for sticky-session behavior

Summary

  • Sticky sessions keep conversations on the same LLM model across multiple turns, preventing mid-conversation switches
  • 30-minute TTL with automatic expiration balances consistency with resource cleanup
  • Dual key strategies support both explicit session IDs (X-Session-Id header) and content-based hashing for stateless clients
  • Activation threshold: Sessions only stick after the first assistant response appears in the message history
  • Bounded memory: 500-entry soft limit with LRU eviction at 1000 entries prevents memory leaks
  • Graceful degradation: Sticky models are verified against available models before enforcement

Frequently Asked Questions

What happens if the sticky model becomes unavailable?

FreeLLMAPI validates the sticky model against the current groupChain before using it. If the model was removed or disabled, the router falls back to standard auto-routing—the conversation continues seamlessly with a new model selection.

Can multiple conversations share the same session ID?

Yes, but only intentionally. Using the same X-Session-Id across different conversation threads binds them to the same model. For isolation, use unique session IDs per conversation thread.

How does FreeLLMAPI handle sticky sessions across server restarts?

The current implementation uses an in-memory Map without persistence. Server restarts clear all sticky sessions. For distributed deployments, sessions are node-local—clients should implement sticky load balancing at the infrastructure layer or use deterministic content hashing.

What's the difference between using X-Session-Id and content-based hashing?

X-Session-Id provides explicit control and works for any message content, including identical prompts from different users. Content-based hashing requires no client-side state management but collapses identical opening messages into the same session—appropriate when conversation context is fully contained in the message history.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →