How OmniRoute Reasoning Replay Works for Multi-Turn Conversations
OmniRoute automatically captures and replays reasoning_content from thinking-mode models using a hybrid in-memory and SQLite cache, ensuring compliant multi-turn conversations without client-side complexity.
OmniRoute is an open-source AI gateway that normalizes interactions with "thinking-mode" models like DeepSeek V4, Kimi K2, and Xiaomi MiMo. Its reasoning replay mechanism solves the critical requirement that reasoning chains must be echoed back in subsequent API calls, handling the complexity transparently so developers can build multi-turn chat applications without managing state manually.
The Provider Contract Problem
Thinking-mode models such as DeepSeek V4, Kimi K2, and Xiaomi MiMo return a reasoning_content field in the first assistant turn. These providers require that identical content be sent back on every subsequent turn. If the client omits this field, the upstream API returns a 400 error with the message: "The reasoning_content in the thinking mode must be passed back to the API".
Detecting Replay Requirements
The detection logic resides in open-sse/services/reasoningCache.ts (lines 78‑108) within the requiresReasoningReplay function. This validator checks three distinct sources to determine if a request needs replay handling:
Explicit Interleaved Fields
If the provider signals interleavedField === "reasoning_content", the cache activates unconditionally for that request.
DeepSeek V4 Pattern Matching
The isDeepSeekReasoningModel helper validates the model name against the V4-specific pattern, triggering replay for compatible DeepSeek endpoints.
Provider and Model Sets
A hard-coded REASONING_REPLAY_PROVIDERS set (including deepseek, kimi-coding, and xiaomi-mimo) and regex patterns (REASONING_REPLAY_MODEL_PATTERNS) catch additional model contracts that require reasoning round-trips.
Capturing Reasoning on the First Turn
When the assistant generates a response, open-sse/handlers/chatCore.ts invokes cacheReasoningFromAssistantMessage (lines 4324‑4328). The implementation:
- Verifies the message role is
assistant - Extracts
reasoning_content(with a fallback to the legacyreasoningfield) - Generates a cache key using either the
tool_call.id(for function calling) or a compositerequestId:messageIndex(for plain assistant messages) - Stores the entry via
cacheReasoning/cacheReasoningByKey, which writes to both the in-memoryMapand SQLite persistence layer
The entire capture block wraps in a non-blocking try...catch to ensure cache failures never interrupt the chat flow.
Hybrid Cache Architecture
OmniRoute uses a dual-layer storage system to balance speed with durability.
In-Memory Hot Path
The memoryCache maintains a Map<string, MemoryCacheEntry> with a default TTL of 2 hours (TTL_MS). Each entry stores the reasoning string, provider, model, and timestamp. When the map exceeds 200 entries (MAX_MEMORY_ENTRIES), the evictOldest routine removes the stalest item. The purgeExpiredMemory function runs on each lookup to prune dead entries.
SQLite Persistence Layer
The module src/lib/db/reasoningCache.ts provides CRUD helpers including setReasoningCache, getReasoningCache, and cleanupExpiredReasoning. This SQLite backing ensures reasoning survives process restarts and enables cache warming across deployments.
Re-injecting Reasoning on Subsequent Turns
When a new request arrives, translators (such as the Claude helper in open-sse/translator/helpers/claudeHelper.ts) call lookupReasoning(toolCallId) in open-sse/services/reasoningCache.ts. The lookup follows a strict priority:
- Memory check – Fast path; increments
hitscounter on success - Database fallback – Calls
getReasoningCacheto read from SQLite, promotes the entry back into memory, then returns the value - Cache miss – Returns
nulland incrementsmisses
On a hit, the translator re-hydrates the reasoning_content into the outgoing request payload, satisfying the provider contract transparently.
Observability and Metrics
The cache service exposes detailed telemetry through getReasoningCacheServiceStats. It tracks:
hits– Successful memory or DB retrievalsmisses– Failed lookupsreplays– Explicit replay events (incremented viarecordReplay)
Operators can view aggregate statistics via the dashboard or programmatically through the /api/cache/reasoning endpoint.
Practical Implementation Examples
Manual Capture (Advanced Use)
While the pipeline handles this automatically, you can manually cache reasoning:
import { cacheReasoning } from '@omniroute/open-sse/services/reasoningCache';
const toolCallId = 'tool_123';
const provider = 'deepseek';
const model = 'deepseek-v4-pro';
const reasoning = 'Step-by-step analysis shows...';
cacheReasoning(toolCallId, provider, model, reasoning);
Manual Lookup in Custom Translators
import { lookupReasoning, recordReplay } from '@omniroute/open-sse/services/reasoningCache';
function injectReasoningIfNeeded(toolCallId: string, requestBody: any) {
const reasoning = lookupReasoning(toolCallId);
if (reasoning) {
requestBody.reasoning_content = reasoning;
recordReplay(); // Track for metrics
}
}
Monitoring Cache Performance
import { getReasoningCacheServiceStats } from '@omniroute/open-sse/services/reasoningCache';
const stats = getReasoningCacheServiceStats();
console.log('Replay rate:', stats.replayRate);
console.log('Hit ratio:', stats.hits / (stats.hits + stats.misses));
Summary
- OmniRoute's reasoning replay is transparent to clients—no extra headers or parameters required
- The hybrid cache combines in-memory speed (2‑hour TTL, 200-entry limit) with SQLite persistence
- Detection covers DeepSeek V4, Kimi K2, Xiaomi MiMo, and other thinking-mode models via provider sets and regex patterns
- Capturing occurs automatically in
chatCore.tsusing tool-call IDs or message indices as keys - Re-injection happens via
lookupReasoning, with automatic fallback from memory to database
Frequently Asked Questions
What happens if reasoning_content is not replayed?
The upstream provider returns a 400 Bad Request with the error message "The reasoning_content in the thinking mode must be passed back to the API", terminating the conversation stream.
How long does OmniRoute retain cached reasoning?
Entries persist for 2 hours (TTL_MS) in memory, with unlimited duration in the SQLite backing store until explicit cleanup. The eviction policy removes the oldest entries when the in-memory cache exceeds 200 items.
Does reasoning replay support function calling?
Yes. When tool calls are present, OmniRoute uses each tool_call.id as the cache key. For plain assistant messages without tools, it falls back to a composite key of requestId:messageIndex, ensuring coverage for both interaction patterns.
How can I verify the cache is working correctly?
Query the /api/cache/reasoning endpoint or import getReasoningCacheServiceStats from @omniroute/open-sse/services/reasoningCache to inspect hit rates, miss counts, and active replay statistics in real time.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →