How OmniRoute ReasoningCache Replays reasoning_content for Multi-Turn Conversations with Strict Providers
OmniRoute's reasoningCache implements a hybrid in-memory and SQLite caching layer that automatically captures and replays reasoning_content on every subsequent turn for strict providers like DeepSeek V4, Kimi Coding, and Xiaomi MiMo, preventing 400 errors during multi-turn conversations.
OmniRoute is an open-source AI routing platform that normalizes interactions across diverse LLM providers. When handling strict providers that require the exact reasoning_content generated on previous turns to be echoed back unchanged, the reasoningCache service ensures seamless multi-turn conversations by persisting and replaying this critical state.
Why Strict Providers Require reasoning_content Replay
Strict providers enforce a special contract: any reasoning_content generated during an assistant turn must be sent back unchanged on every subsequent turn. If the content is missing or modified, the upstream API returns a 400 error with the message "The reasoning_content in the thinking mode must be passed back to the API."
This requirement applies to providers such as DeepSeek V4, Kimi Coding, and Xiaomi MiMo, where the reasoning trace forms part of the conversational state that the model expects to maintain across turns.
Detecting When Replay Is Required
The requiresReasoningReplay() function in open-sse/services/reasoningCache.ts (lines 79-109) determines whether a provider/model combination necessitates reasoning replay. The detection logic:
- Normalizes provider and model strings to handle case variations
- Checks an explicit
interleavedFieldflag indicating interleaved reasoning support - Falls back to a whitelist of providers (
REASONING_REPLAY_PROVIDERS) and model regex patterns (REASONING_REPLAY_MODEL_PATTERNS)
// Detection happens before request processing
if (requiresReasoningReplay(providerId, modelId)) {
// Enable reasoning capture for this conversation
}
Capturing reasoning_content from Assistant Messages
When an assistant message contains reasoning_content (or the legacy reasoning field), the cacheReasoningFromAssistantMessage() function (lines 45-86) extracts and stores the text.
The function handles two storage strategies:
- Tool-call keys: When the assistant message contains tool calls, the reasoning is stored under each tool-call ID present in the message
- Assistant-message fallback: If no tool calls exist, the system generates a cache key using
buildAssistantMessageCacheKeybased on the request ID and message index
import { cacheReasoningFromAssistantMessage } from '@omniroute/open-sse/services/reasoningCache';
// Capture reasoning from provider response
cacheReasoningFromAssistantMessage(assistantMessage, providerId, modelId, {
requestId: requestId,
messageIndex: turnIndex,
});
Hybrid Storage Architecture
The reasoningCache implements a dual-layer storage system combining volatile memory with durable SQLite persistence.
In-Memory Layer
- Capacity: 200 entries maximum
- Size cap: 10 KB per entry
- TTL: 2 hours expiration
- Structure: In-process
Mapobject for O(1) lookups
SQLite Persistence
The cacheReasoningByKey() function writes entries to both memory and SQLite via setReasoningCache() in src/lib/db/reasoningCache.ts, ensuring reasoning_content survives process restarts.
// Writes to both Map and SQLite
cacheReasoningByKey(key, reasoningContent, metadata);
Replaying Cached Content on Subsequent Turns
The lookupReasoning() function (lines 88-136) retrieves stored reasoning_content when preparing requests for strict providers. The lookup process:
- Checks the in-memory cache first for hot-path performance
- Falls back to SQLite (
getReasoningCache) if not found in memory - Promotes SQLite hits back into memory for faster subsequent lookups
- Tracks metrics: increments
hitsormissescounters accordingly
When the cache successfully provides reasoning_content, recordReplay() increments the replays metric for observability.
import { lookupReasoning, recordReplay } from '@omniroute/open-sse/services/reasoningCache';
const cachedReasoning = lookupReasoning(toolCallId);
if (cachedReasoning) {
// Inject into request body as required by strict provider
requestBody.reasoning_content = cachedReasoning;
recordReplay(); // Track successful replay
}
Cache Statistics and Maintenance
The service exposes detailed telemetry through getReasoningCacheServiceStats() (lines 148-186), aggregating:
- Memory utilization (current size vs. capacity)
- Database entry count
- Hit/miss/replay ratios
- Per-provider and per-model breakdowns
Automatic Cleanup
The module initializes an auto-cleanup timer via startAutoCleanup() on import. This background process:
- Removes expired entries from the memory cache
- Executes
cleanupExpiredReasoningin SQLite - Runs on a configurable interval (default: 30 minutes)
import { getReasoningCacheServiceStats, clearReasoningCacheAll } from '@omniroute/open-sse/services/reasoningCache';
// Monitor cache performance
const stats = await getReasoningCacheServiceStats();
console.log('Hit rate:', (stats.hits / (stats.hits + stats.misses) * 100).toFixed(1) + '%');
// Clear all entries (useful for admin operations)
const removed = clearReasoningCacheAll();
Key Implementation Files
| File | Role |
|---|---|
open-sse/services/reasoningCache.ts |
Core hybrid cache implementation including detection, capture, lookup, and cleanup logic |
src/lib/db/reasoningCache.ts |
SQLite schema and CRUD operations (setReasoningCache, getReasoningCache, cleanupExpiredReasoning) |
open-sse/handlers/chatCore.ts |
Request pipeline integration where cache lookups occur before provider dispatch |
open-sse/translator/helpers/claudeHelper.ts |
Example implementation showing cache integration for Claude-style tool calling |
Summary
- Strict providers require identical
reasoning_contentto be replayed on every subsequent turn or they return 400 errors - OmniRoute's reasoningCache uses
requiresReasoningReplay()to detect when this behavior is necessary based on provider/model whitelists - Hybrid storage combines a bounded in-memory Map (200 entries, 10KB limit, 2h TTL) with SQLite persistence for durability
- Capture occurs via
cacheReasoningFromAssistantMessage(), which keys storage by tool-call ID or assistant message index - Replay happens through
lookupReasoning(), which checks memory first, falls back to SQLite, and promotes hits for performance - Automatic maintenance includes hit/miss tracking, replay metrics, and scheduled cleanup of expired entries
Frequently Asked Questions
What happens if reasoning_content is not replayed to strict providers?
The upstream API returns a 400 Bad Request error with the message "The reasoning_content in the thinking mode must be passed back to the API." This terminates the conversation thread, requiring the client to resubmit the request with the correct reasoning_content included.
How does OmniRoute handle reasoning_content for messages without tool calls?
When an assistant message lacks tool calls, the system falls back to an assistant-message cache key constructed from the requestId and messageIndex parameters via buildAssistantMessageCacheKey(). This ensures reasoning traces are still recoverable for conversational turns that don't involve function calling.
What are the memory limits of the reasoningCache?
The in-memory layer enforces three constraints: 200 entries maximum, 10 KB size cap per entry, and a 2-hour TTL (time-to-live). Exceeding the entry count triggers eviction, while entries exceeding 10 KB or lasting longer than 2 hours are automatically removed during cleanup cycles.
Can I manually clear the reasoningCache for debugging?
Yes. The clearReasoningCacheAll() function removes all entries from both memory and SQLite persistence. Additionally, the public API endpoint at src/app/api/cache/reasoning/route.ts exposes administrative endpoints for inspection and manual cache clearing via HTTP requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →