How Reasoning Replay Works in OmniRoute: A Complete Technical Guide
Reasoning replay in OmniRoute captures the reasoning_content from thinking-mode LLM responses, stores it server-side, and re-injects it into subsequent requests to prevent HTTP 400 errors from providers that require complete conversation history.
The reasoning replay feature powers OmniRoute's compatibility with strict thinking-mode providers like DeepSeek, Kimi, and Qwen. This article examines the implementation across capture, storage, and replay layers based on the diegosouzapw/OmniRoute source code.
Why Reasoning Replay Is Necessary
Thinking-mode LLMs return structured responses containing both content (the final answer) and reasoning_content (the model's chain-of-thought). Standard clients often discard the reasoning field to save tokens, but providers like DeepSeek V4 and Kimi K2 require it to be present in every subsequent turn. When missing, these providers return HTTP 400 errors.
OmniRoute solves this by caching reasoning server-side and transparently restoring it, eliminating the burden on clients while maintaining strict provider compatibility.
Capture Layer: Intercepting Assistant Responses
Assistant responses are intercepted in the chat handler at two critical points in open-sse/handlers/chatCore.ts (around lines 4093 and 4380). The cacheReasoningFromAssistantMessage() function extracts reasoning content and immediately writes it to both storage layers.
// Capture flow inside the chat handler
import { setReasoningCache } from '@/src/lib/db/reasoningCache';
// After an assistant turn with tool calls:
setReasoningCache(
toolCallId, // unique ID from the tool call
provider, // e.g., 'deepseek'
model, // e.g., 'deepseek-chat'
reasoningContent, // the raw reasoning string
);
The capture operation is synchronous with the response stream, ensuring no reasoning is lost even under high load.
Storage Architecture: Dual-Layer Caching
OmniRoute implements a hot-path cache + persistent storage pattern for reasoning data.
In-Memory LRU Cache
The open-sse/services/reasoningCache.ts module maintains a Map<string, ReasoningCacheEntry> with these characteristics:
- Capacity: 2,000 entries maximum
- Eviction: Oldest entries removed first when limit reached
- TTL: 2 hours (
DEFAULT_TTL_MS) - Lookup complexity: O(1)
This cache survives only for the process lifetime and handles the vast majority of lookups without touching disk.
SQLite Persistence
Misses fall back to the reasoning_cache table defined in src/lib/db/migrations/033_create_reasoning_cache.sql:
| Column | Purpose |
|---|---|
tool_call_id |
Primary lookup key |
provider |
Source provider name |
model |
Specific model identifier |
reasoning |
Full reasoning content |
char_count |
Length for quick stats |
created_at |
Unix epoch timestamp |
expires_at |
TTL enforcement via EXPIRES_AT_EPOCH_SQL |
The SQLite layer ensures crash resistance and powers the dashboard visualization. The src/lib/db/reasoningCache.ts module encapsulates all CRUD operations against this table.
Replay Layer: Injecting Missing Reasoning
When a client sends a new turn, the translator in open-sse/translator/index.ts evaluates whether replay is required.
Provider/Model Detection
The requiresReasoningReplay(provider, model) function returns true for:
Providers (case-insensitive exact match):
deepseekopencode-gosiliconflownebiusdeepinfrasambanovafireworkstogetherkimi-codingkimi-coding-apikeyxiaomi-mimo
Models (case-insensitive regex):
/deepseek-r1/i/deepseek-reasoner/i/deepseek-chat/i
This logic resides in open-sse/services/reasoningCache.ts and is evaluated on every incoming request.
Replay Execution
// Replay logic in the translator
import { getReasoningFromCache } from '@/open-sse/services/reasoningCache';
async function maybeReplay(message: AssistantMessage) {
if (!message.reasoning_content && message.tool_calls?.length) {
const cached = await getReasoningFromCache(message.tool_calls[0].id);
if (cached) {
message.reasoning_content = cached.reasoning;
// Replay is now recorded for observability
}
}
}
The translator mutates the message in-place before dispatch to the upstream provider, making the operation transparent to both client and provider.
Observability and Operations
OmniRoute exposes reasoning cache metrics through a dedicated API endpoint.
Cache Status Endpoint
GET /api/v1/cache/reasoning (implemented in src/app/api/v1/cache/reasoning/route.ts) returns:
- Current cache size (in-memory)
- Hit/miss ratios
- Expiration statistics
- TTL distribution
# Inspect cache status
curl http://localhost:20128/api/v1/cache/reasoning
Dashboard Integration
The SQLite table is directly queried by dashboard widgets for historical analysis, including:
- Reasoning volume by provider
- Cache efficiency trends
- TTL expiration patterns
Key Design Decisions
| Decision | Rationale |
|---|---|
| Tool-call ID as primary key | Guarantees uniqueness across multi-turn conversations |
| Dual storage layers | Sub-millisecond lookups for hot data, durability for recovery |
| 2-hour TTL | Balances provider requirements with storage constraints |
| Provider whitelist | Prevents unnecessary overhead for compatible models |
| In-place message mutation | Maintains OpenAI-compatible API surface |
File Reference Map
| File | Responsibility |
|---|---|
open-sse/handlers/chatCore.ts |
Response interception and capture triggers |
open-sse/translator/index.ts |
Replay decision and injection |
open-sse/services/reasoningCache.ts |
In-memory LRU and provider detection |
src/lib/db/reasoningCache.ts |
SQLite persistence layer |
src/lib/db/migrations/033_create_reasoning_cache.sql |
Schema definition |
src/app/api/v1/cache/reasoning/route.ts |
HTTP metrics endpoint |
docs/routing/REASONING_REPLAY.md |
Design documentation |
Summary
- Reasoning replay prevents HTTP 400 errors by preserving
reasoning_contentacross conversation turns - Two-tier storage combines sub-millisecond in-memory lookups with SQLite durability
- Automatic provider detection enables the feature only for strict thinking-mode models
- Transparent operation requires no client changes—OmniRoute handles capture and replay internally
- Full observability via REST API and dashboard integration supports production monitoring
Frequently Asked Questions
How does OmniRoute know which providers require reasoning replay?
The requiresReasoningReplay() function in open-sse/services/reasoningCache.ts maintains explicit whitelists. It checks the provider name against 11 known strict providers and the model name against three DeepSeek patterns. This list is hardcoded and updated as new thinking-mode providers emerge.
What happens if the cache entry expires before the next turn?
Expired entries are treated as cache misses. The translator proceeds without injecting reasoning, which may cause the upstream provider to return a 400 error. The 2-hour TTL balances this risk against storage constraints; most conversational contexts complete well within this window.
Can I manually clear or inspect the reasoning cache?
Yes. The GET /api/v1/cache/reasoning endpoint exposes current statistics, and the SQLite database at reasoning_cache can be queried directly for detailed inspection. There is no built-in manual eviction API—entries rely on TTL expiration or LRU eviction when the 2,000-entry limit is reached.
Does reasoning replay affect response latency?
The hot-path in-memory cache provides O(1) lookups with negligible overhead. SQLite fallback adds milliseconds only on cold starts or cache pressure scenarios. The capture operation is pipelined with response streaming and does not block client delivery.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →