# Context Handoff in FreeLLMAPI: How It Preserves Conversation Continuity Across Model Switches

> Discover context handoff in FreeLLMAPI. This feature seamlessly preserves conversation continuity across LLM model switches by summarizing recent dialogue, ensuring a smooth user experience.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: deep-dive
- Published: 2026-09-01

---

**Context handoff is a mechanism that automatically preserves conversation continuity when a user’s session is routed to a different LLM model part-way through a multi-turn chat by injecting a system message containing a summary of recent dialogue.**

FreeLLMAPI, an open-source unified API gateway for large language models, implements **context handoff** to solve the problem of context loss during model failover or routing changes. When the system detects that the `selectedModelKey` differs from the `lastModelKey` stored for an active session, it seamlessly transfers conversation state to the new model. This ensures that even when sticky sessions break due to quota limits or provider failures, the user experience remains coherent and uninterrupted.

## How Context Handoff Works

The handoff mechanism operates in [`server/src/services/context-handoff.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/context-handoff.ts) through a six-step process that intercepts requests before they reach the downstream provider.

### Detection Logic

The system determines whether to trigger a handoff in the `maybeInjectContextHandoff` function. It compares the newly selected model key against the `lastModelKey` stored in the session metadata:

- If `lastModelKey` exists **and** differs from `selectedModelKey`, a handoff is required
- If the keys match or no prior model exists, the request passes through unchanged

This check occurs within the request pipeline defined in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts), where the service imports `HANDOFF_MAX_TOKENS` to account for the additional context overhead.

### Message Injection

When triggered, the system constructs a **system-role message** containing:

1. **An explanation of the model switch** – Declaring that the model is taking over from another instance (e.g., "You are taking over an ongoing conversation from another model...")
2. **A concise summary** – Up to approximately 6,000 characters of recent user-assistant turns (typically the last 12 trimmed messages)

The injected message is strategically placed **after any existing system prompts** to preserve provider-specific prompt ordering. The `buildSummary` function handles the truncation and formatting of historical messages to fit within token constraints.

### Token Accounting

To prevent context window overflow, FreeLLMAPI reserves approximately **1,500 tokens** for the handoff message via the exported constant `HANDOFF_MAX_TOKENS`. The routing layer adds this buffer to its token-budget calculations before forwarding requests to the provider, ensuring that the injected summary does not unexpectedly truncate the actual user query.

## Configuration Options

Control the handoff behavior using the environment variable `FREELLMAPI_CONTEXT_HANDOFF`:

| Value | Effect |
|-------|--------|
| `off` | Disables context handoff entirely (default behavior) |
| `on_model_switch` | Enables handoff injection when the router changes models |

Enable the feature before starting the server:

```bash
export FREELLMAPI_CONTEXT_HANDOFF=on_model_switch
npm run start

```

When disabled, the system performs "hard" model switches where the new model receives no historical context from previous turns in the session.

## Implementation Details

The core logic resides in [`server/src/services/context-handoff.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/context-handoff.ts), which exports three primary functions used by the request router:

- **`getContextHandoffMode`** – Reads the `FREELLMAPI_CONTEXT_HANDOFF` environment variable to determine operational mode
- **`recordIncomingMessages`** – Stores up to 12 recent trimmed messages per session key with TTL updates, maintaining the conversation history buffer
- **`recordSuccessfulModel`** – Persists the `lastModelKey` after a successful response, establishing the baseline for future handoff comparisons

The [`proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/proxy.ts) route handler orchestrates these calls: it invokes `recordIncomingMessages` to capture the current turn, checks `maybeInjectContextHandoff` before forwarding, and calls `recordSuccessfulModel` upon receiving a response to update the session state.

## Code Example: Triggering a Handoff

Consider a session `sess-123` that previously received responses from `gpt-4o`. When you request `claude-2` for the next turn:

```json
POST http://localhost:3001/v1/chat/completions
Content-Type: application/json
Authorization: Bearer <YOUR_UNIFIED_KEY>

{
  "model": "claude-2",
  "messages": [
    { "role": "user", "content": "What did we discuss about the project timeline?" }
  ],
  "session_key": "sess-123"
}

```

FreeLLMAPI executes the following internally:

1. Retrieves the stored `lastModelKey` (`gpt-4o`) using `recordIncomingMessages`
2. Detects the mismatch with `selectedModelKey` (`claude-2`)
3. Calls `maybeInjectContextHandoff` to generate a system message:

```text
FreeLLMAPI context handoff:
You are taking over an ongoing conversation from another model (gpt-4o → claude-2).
Continue the user's task using the conversation context already provided in this request.
Do not restart the task, re-ask already answered setup questions, or discard prior tool results.
Respect the user's latest message as the highest-priority instruction.

Recent session summary:
User: Can you outline the project timeline?
Assistant: The project will run from Q1 to Q4, with milestones...

```

4. Inserts this message after any existing system prompts in the request payload
5. Accounts for the extra ~1,500 tokens via `HANDOFF_MAX_TOKENS` before routing to Claude
6. Updates the session's `lastModelKey` to `claude-2` via `recordSuccessfulModel` upon response

## Summary

- **Context handoff** prevents conversation fragmentation when FreeLLMAPI switches models mid-session by injecting a summary system message
- The mechanism activates only when `FREELLMAPI_CONTEXT_HANDOFF` is set to `on_model_switch` and the `selectedModelKey` differs from `lastModelKey`
- The system preserves up to 12 recent messages (approximately 6,000 characters) and reserves 1,500 tokens via `HANDOFF_MAX_TOKENS` to accommodate the handoff message
- All logic is implemented in [`server/src/services/context-handoff.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/context-handoff.ts) and integrated into the routing layer at [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts)

## Frequently Asked Questions

### What happens if context handoff is disabled?

When `FREELLMAPI_CONTEXT_HANDOFF` is set to `off` (the default), the router performs hard model switches. The new model receives only the messages explicitly included in the current request payload, with no access to the conversation history stored from previous turns in that session. This may cause the model to ask clarifying questions that were already answered earlier in the conversation.

### How much conversation history does the handoff message include?

The handoff summary includes up to 12 recent user-assistant message pairs, trimmed to fit within approximately 6,000 characters. This history is generated by the `buildSummary` function in [`server/src/services/context-handoff.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/context-handoff.ts) and formatted as a readable dialogue transcript within the injected system message.

### Does context handoff work with all LLM providers?

Yes, the handoff mechanism is provider-agnostic. Because FreeLLMAPI injects the summary as a standard system-role message conforming to the OpenAI-compatible chat completions format, any provider supporting system messages—including Anthropic, OpenAI, and Google models—can interpret the context handoff correctly.

### How does token accounting prevent context window errors?

Before routing, the system imports `HANDOFF_MAX_TOKENS` (approximately 1,500 tokens) from [`server/src/services/context-handoff.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/context-handoff.ts) and adds this value to the token count check in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts). This ensures the provider's context window limit is not exceeded when the summary message is appended to the existing conversation payload.