Context Handoff in FreeLLMAPI: How It Preserves Conversation Continuity Across Model Switches

Context handoff is a mechanism that automatically preserves conversation continuity when a user’s session is routed to a different LLM model part-way through a multi-turn chat by injecting a system message containing a summary of recent dialogue.

FreeLLMAPI, an open-source unified API gateway for large language models, implements context handoff to solve the problem of context loss during model failover or routing changes. When the system detects that the selectedModelKey differs from the lastModelKey stored for an active session, it seamlessly transfers conversation state to the new model. This ensures that even when sticky sessions break due to quota limits or provider failures, the user experience remains coherent and uninterrupted.

How Context Handoff Works

The handoff mechanism operates in server/src/services/context-handoff.ts through a six-step process that intercepts requests before they reach the downstream provider.

Detection Logic

The system determines whether to trigger a handoff in the maybeInjectContextHandoff function. It compares the newly selected model key against the lastModelKey stored in the session metadata:

  • If lastModelKey exists and differs from selectedModelKey, a handoff is required
  • If the keys match or no prior model exists, the request passes through unchanged

This check occurs within the request pipeline defined in server/src/routes/proxy.ts, where the service imports HANDOFF_MAX_TOKENS to account for the additional context overhead.

Message Injection

When triggered, the system constructs a system-role message containing:

  1. An explanation of the model switch – Declaring that the model is taking over from another instance (e.g., "You are taking over an ongoing conversation from another model...")
  2. A concise summary – Up to approximately 6,000 characters of recent user-assistant turns (typically the last 12 trimmed messages)

The injected message is strategically placed after any existing system prompts to preserve provider-specific prompt ordering. The buildSummary function handles the truncation and formatting of historical messages to fit within token constraints.

Token Accounting

To prevent context window overflow, FreeLLMAPI reserves approximately 1,500 tokens for the handoff message via the exported constant HANDOFF_MAX_TOKENS. The routing layer adds this buffer to its token-budget calculations before forwarding requests to the provider, ensuring that the injected summary does not unexpectedly truncate the actual user query.

Configuration Options

Control the handoff behavior using the environment variable FREELLMAPI_CONTEXT_HANDOFF:

Value Effect
off Disables context handoff entirely (default behavior)
on_model_switch Enables handoff injection when the router changes models

Enable the feature before starting the server:

export FREELLMAPI_CONTEXT_HANDOFF=on_model_switch
npm run start

When disabled, the system performs "hard" model switches where the new model receives no historical context from previous turns in the session.

Implementation Details

The core logic resides in server/src/services/context-handoff.ts, which exports three primary functions used by the request router:

  • getContextHandoffMode – Reads the FREELLMAPI_CONTEXT_HANDOFF environment variable to determine operational mode
  • recordIncomingMessages – Stores up to 12 recent trimmed messages per session key with TTL updates, maintaining the conversation history buffer
  • recordSuccessfulModel – Persists the lastModelKey after a successful response, establishing the baseline for future handoff comparisons

The proxy.ts route handler orchestrates these calls: it invokes recordIncomingMessages to capture the current turn, checks maybeInjectContextHandoff before forwarding, and calls recordSuccessfulModel upon receiving a response to update the session state.

Code Example: Triggering a Handoff

Consider a session sess-123 that previously received responses from gpt-4o. When you request claude-2 for the next turn:

POST http://localhost:3001/v1/chat/completions
Content-Type: application/json
Authorization: Bearer <YOUR_UNIFIED_KEY>

{
  "model": "claude-2",
  "messages": [
    { "role": "user", "content": "What did we discuss about the project timeline?" }
  ],
  "session_key": "sess-123"
}

FreeLLMAPI executes the following internally:

  1. Retrieves the stored lastModelKey (gpt-4o) using recordIncomingMessages
  2. Detects the mismatch with selectedModelKey (claude-2)
  3. Calls maybeInjectContextHandoff to generate a system message:
FreeLLMAPI context handoff:
You are taking over an ongoing conversation from another model (gpt-4o → claude-2).
Continue the user's task using the conversation context already provided in this request.
Do not restart the task, re-ask already answered setup questions, or discard prior tool results.
Respect the user's latest message as the highest-priority instruction.

Recent session summary:
User: Can you outline the project timeline?
Assistant: The project will run from Q1 to Q4, with milestones...
  1. Inserts this message after any existing system prompts in the request payload
  2. Accounts for the extra ~1,500 tokens via HANDOFF_MAX_TOKENS before routing to Claude
  3. Updates the session's lastModelKey to claude-2 via recordSuccessfulModel upon response

Summary

  • Context handoff prevents conversation fragmentation when FreeLLMAPI switches models mid-session by injecting a summary system message
  • The mechanism activates only when FREELLMAPI_CONTEXT_HANDOFF is set to on_model_switch and the selectedModelKey differs from lastModelKey
  • The system preserves up to 12 recent messages (approximately 6,000 characters) and reserves 1,500 tokens via HANDOFF_MAX_TOKENS to accommodate the handoff message
  • All logic is implemented in server/src/services/context-handoff.ts and integrated into the routing layer at server/src/routes/proxy.ts

Frequently Asked Questions

What happens if context handoff is disabled?

When FREELLMAPI_CONTEXT_HANDOFF is set to off (the default), the router performs hard model switches. The new model receives only the messages explicitly included in the current request payload, with no access to the conversation history stored from previous turns in that session. This may cause the model to ask clarifying questions that were already answered earlier in the conversation.

How much conversation history does the handoff message include?

The handoff summary includes up to 12 recent user-assistant message pairs, trimmed to fit within approximately 6,000 characters. This history is generated by the buildSummary function in server/src/services/context-handoff.ts and formatted as a readable dialogue transcript within the injected system message.

Does context handoff work with all LLM providers?

Yes, the handoff mechanism is provider-agnostic. Because FreeLLMAPI injects the summary as a standard system-role message conforming to the OpenAI-compatible chat completions format, any provider supporting system messages—including Anthropic, OpenAI, and Google models—can interpret the context handoff correctly.

How does token accounting prevent context window errors?

Before routing, the system imports HANDOFF_MAX_TOKENS (approximately 1,500 tokens) from server/src/services/context-handoff.ts and adds this value to the token count check in server/src/routes/proxy.ts. This ensures the provider's context window limit is not exceeded when the summary message is appended to the existing conversation payload.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →