Context Handoff in FreeLLMAPI: How It Preserves Conversation Continuity Across Model Switches
Context handoff is a mechanism that automatically preserves conversation continuity when a user’s session is routed to a different LLM model part-way through a multi-turn chat by injecting a system message containing a summary of recent dialogue.
FreeLLMAPI, an open-source unified API gateway for large language models, implements context handoff to solve the problem of context loss during model failover or routing changes. When the system detects that the selectedModelKey differs from the lastModelKey stored for an active session, it seamlessly transfers conversation state to the new model. This ensures that even when sticky sessions break due to quota limits or provider failures, the user experience remains coherent and uninterrupted.
How Context Handoff Works
The handoff mechanism operates in server/src/services/context-handoff.ts through a six-step process that intercepts requests before they reach the downstream provider.
Detection Logic
The system determines whether to trigger a handoff in the maybeInjectContextHandoff function. It compares the newly selected model key against the lastModelKey stored in the session metadata:
- If
lastModelKeyexists and differs fromselectedModelKey, a handoff is required - If the keys match or no prior model exists, the request passes through unchanged
This check occurs within the request pipeline defined in server/src/routes/proxy.ts, where the service imports HANDOFF_MAX_TOKENS to account for the additional context overhead.
Message Injection
When triggered, the system constructs a system-role message containing:
- An explanation of the model switch – Declaring that the model is taking over from another instance (e.g., "You are taking over an ongoing conversation from another model...")
- A concise summary – Up to approximately 6,000 characters of recent user-assistant turns (typically the last 12 trimmed messages)
The injected message is strategically placed after any existing system prompts to preserve provider-specific prompt ordering. The buildSummary function handles the truncation and formatting of historical messages to fit within token constraints.
Token Accounting
To prevent context window overflow, FreeLLMAPI reserves approximately 1,500 tokens for the handoff message via the exported constant HANDOFF_MAX_TOKENS. The routing layer adds this buffer to its token-budget calculations before forwarding requests to the provider, ensuring that the injected summary does not unexpectedly truncate the actual user query.
Configuration Options
Control the handoff behavior using the environment variable FREELLMAPI_CONTEXT_HANDOFF:
| Value | Effect |
|---|---|
off |
Disables context handoff entirely (default behavior) |
on_model_switch |
Enables handoff injection when the router changes models |
Enable the feature before starting the server:
export FREELLMAPI_CONTEXT_HANDOFF=on_model_switch
npm run start
When disabled, the system performs "hard" model switches where the new model receives no historical context from previous turns in the session.
Implementation Details
The core logic resides in server/src/services/context-handoff.ts, which exports three primary functions used by the request router:
getContextHandoffMode– Reads theFREELLMAPI_CONTEXT_HANDOFFenvironment variable to determine operational moderecordIncomingMessages– Stores up to 12 recent trimmed messages per session key with TTL updates, maintaining the conversation history bufferrecordSuccessfulModel– Persists thelastModelKeyafter a successful response, establishing the baseline for future handoff comparisons
The proxy.ts route handler orchestrates these calls: it invokes recordIncomingMessages to capture the current turn, checks maybeInjectContextHandoff before forwarding, and calls recordSuccessfulModel upon receiving a response to update the session state.
Code Example: Triggering a Handoff
Consider a session sess-123 that previously received responses from gpt-4o. When you request claude-2 for the next turn:
POST http://localhost:3001/v1/chat/completions
Content-Type: application/json
Authorization: Bearer <YOUR_UNIFIED_KEY>
{
"model": "claude-2",
"messages": [
{ "role": "user", "content": "What did we discuss about the project timeline?" }
],
"session_key": "sess-123"
}
FreeLLMAPI executes the following internally:
- Retrieves the stored
lastModelKey(gpt-4o) usingrecordIncomingMessages - Detects the mismatch with
selectedModelKey(claude-2) - Calls
maybeInjectContextHandoffto generate a system message:
FreeLLMAPI context handoff:
You are taking over an ongoing conversation from another model (gpt-4o → claude-2).
Continue the user's task using the conversation context already provided in this request.
Do not restart the task, re-ask already answered setup questions, or discard prior tool results.
Respect the user's latest message as the highest-priority instruction.
Recent session summary:
User: Can you outline the project timeline?
Assistant: The project will run from Q1 to Q4, with milestones...
- Inserts this message after any existing system prompts in the request payload
- Accounts for the extra ~1,500 tokens via
HANDOFF_MAX_TOKENSbefore routing to Claude - Updates the session's
lastModelKeytoclaude-2viarecordSuccessfulModelupon response
Summary
- Context handoff prevents conversation fragmentation when FreeLLMAPI switches models mid-session by injecting a summary system message
- The mechanism activates only when
FREELLMAPI_CONTEXT_HANDOFFis set toon_model_switchand theselectedModelKeydiffers fromlastModelKey - The system preserves up to 12 recent messages (approximately 6,000 characters) and reserves 1,500 tokens via
HANDOFF_MAX_TOKENSto accommodate the handoff message - All logic is implemented in
server/src/services/context-handoff.tsand integrated into the routing layer atserver/src/routes/proxy.ts
Frequently Asked Questions
What happens if context handoff is disabled?
When FREELLMAPI_CONTEXT_HANDOFF is set to off (the default), the router performs hard model switches. The new model receives only the messages explicitly included in the current request payload, with no access to the conversation history stored from previous turns in that session. This may cause the model to ask clarifying questions that were already answered earlier in the conversation.
How much conversation history does the handoff message include?
The handoff summary includes up to 12 recent user-assistant message pairs, trimmed to fit within approximately 6,000 characters. This history is generated by the buildSummary function in server/src/services/context-handoff.ts and formatted as a readable dialogue transcript within the injected system message.
Does context handoff work with all LLM providers?
Yes, the handoff mechanism is provider-agnostic. Because FreeLLMAPI injects the summary as a standard system-role message conforming to the OpenAI-compatible chat completions format, any provider supporting system messages—including Anthropic, OpenAI, and Google models—can interpret the context handoff correctly.
How does token accounting prevent context window errors?
Before routing, the system imports HANDOFF_MAX_TOKENS (approximately 1,500 tokens) from server/src/services/context-handoff.ts and adds this value to the token count check in server/src/routes/proxy.ts. This ensures the provider's context window limit is not exceeded when the summary message is appended to the existing conversation payload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →