How to Configure and Utilize the Reasoning Replay Feature in OmniRoute
TLDR: OmniRoute's reasoning replay feature automatically caches reasoning_content from thinking-mode LLMs and re-injects it on subsequent turns, preventing HTTP 400 errors from providers that require original reasoning in conversation history.
OmniRoute's reasoning replay feature solves a critical compatibility problem with modern thinking-mode LLMs like DeepSeek R1, Kimi K2, and Qwen-Thinking. This guide explains how the feature works, which providers support it, and how to monitor and manage the caching layer according to the diegosouzapw/OmniRoute source code.
What Is Reasoning Replay?
Thinking-mode providers require the original reasoning_content field to be present in every subsequent message of a conversation. Most clients strip this field, causing strict upstream providers to return HTTP 400 errors. OmniRoute transparently solves this by:
- Capturing reasoning on assistant turns and storing it server-side
- Replaying the cached reasoning into future requests when needed
- Persisting data across restarts with SQLite while maintaining fast in-memory lookup
Supported Providers and Models
The requiresReasoningReplay(provider, model) function in open-sse/services/reasoningCache.ts determines when replay is enabled. It matches against hardcoded lists:
Provider whitelist (case-insensitive):
deepseekopencode-gosiliconflownebiusdeepinfrasambanovafireworkstogetherkimi-codingkimi-coding-apikeyxiaomi-mimo
Model regex patterns (case-insensitive):
/deepseek-r1/i/deepseek-reasoner/i/deepseek-chat/i
No configuration is required—these checks run automatically on every request.
Architecture: Capture, Cache, and Replay
Capture Layer
In open-sse/handlers/chatCore.ts (lines 4093 and 4380), assistant responses are intercepted. The cacheReasoningFromAssistantMessage() function extracts reasoning and writes it to two locations:
// Automatic capture during chat handling (no manual setup needed)
import { setReasoningCache } from '@/src/lib/db/reasoningCache';
setReasoningCache(
toolCallId, // unique ID from the tool call
provider, // e.g., 'deepseek'
model, // e.g., 'deepseek-chat'
reasoningContent, // the raw reasoning string from the LLM
);
Dual-Layer Caching
| Layer | Location | Capacity | Persistence |
|---|---|---|---|
| Hot-path cache | open-sse/services/reasoningCache.ts |
2,000 entries (LRU by creation) | Process only |
| Persistent store | src/lib/db/reasoningCache.ts |
Unlimited (within SQLite limits) | Survives restarts |
The in-memory Map<string, ReasoningCacheEntry> provides O(1) lookups. Misses fall back to SQLite. Entries expire after 2 hours (DEFAULT_TTL_MS) with oldest-first eviction.
Replay Layer
The translator in open-sse/translator/index.ts checks requiresReasoningReplay() before dispatch. When true, it retrieves missing reasoning by tool-call ID:
// Automatic replay during request translation
import { getReasoningFromCache } from '@/open-sse/services/reasoningCache';
async function maybeReplay(message: AssistantMessage) {
if (!message.reasoning_content && message.tool_calls?.length) {
const cached = await getReasoningFromCache(message.tool_calls[0].id);
if (cached) {
message.reasoning_content = cached.reasoning;
}
}
}
Database Schema and Migration
The SQLite table is created by migration src/lib/db/migrations/033_create_reasoning_cache.sql:
| Column | Type | Purpose |
|---|---|---|
tool_call_id |
TEXT PRIMARY KEY |
Unique identifier for lookup |
provider |
TEXT |
Source provider name |
model |
TEXT |
Model identifier |
reasoning |
TEXT |
Full reasoning content |
char_count |
INTEGER |
Length for stats |
created_at |
INTEGER |
Unix epoch |
expires_at |
INTEGER |
Unix epoch (normalized via EXPIRES_AT_EPOCH_SQL) |
Monitoring Cache Status
OmniRoute exposes cache metrics via REST API:
# Query cache statistics
curl http://localhost:20128/api/v1/cache/reasoning
The endpoint is implemented in src/app/api/v1/cache/reasoning/route.ts and returns:
- Current in-memory cache size
- Hit/miss ratios
- Expiration statistics
Dashboard widgets read directly from SQLite via the DB domain module for inspection and manual eviction when needed.
Configuration: What You Can and Cannot Change
| Aspect | Configurable? | Notes |
|---|---|---|
| Provider/model list | No | Hardcoded in open-sse/services/reasoningCache.ts |
| Cache capacity (2,000) | No | Compile-time constant |
| TTL (2 hours) | No | DEFAULT_TTL_MS constant |
| SQLite path | Yes | Standard OmniRoute database configuration |
To add new providers or adjust limits, modify the source in open-sse/services/reasoningCache.ts and rebuild.
Key Files Reference
| File | Purpose |
|---|---|
open-sse/services/reasoningCache.ts |
In-memory LRU logic and provider detection |
src/lib/db/reasoningCache.ts |
SQLite persistence layer |
open-sse/handlers/chatCore.ts |
Capture hooks (lines 4093, 4380) |
open-sse/translator/index.ts |
Replay injection logic |
src/lib/db/migrations/033_create_reasoning_cache.sql |
Schema definition |
src/app/api/v1/cache/reasoning/route.ts |
Status API endpoint |
docs/routing/REASONING_REPLAY.md |
Design documentation |
Summary
- Reasoning replay eliminates HTTP 400 errors with strict thinking-mode providers by server-side caching of
reasoning_content - Zero configuration required for supported providers—detection is automatic via
requiresReasoningReplay() - Dual-layer architecture balances speed (in-memory LRU, 2,000 entries) with durability (SQLite persistence)
- 2-hour TTL with automatic eviction prevents unbounded growth
- Monitor health via
GET /api/v1/cache/reasoningor dashboard SQLite access
Frequently Asked Questions
Does reasoning replay work with custom or self-hosted models?
Only if the provider name matches the hardcoded whitelist in open-sse/services/reasoningCache.ts. The list includes major providers like DeepSeek, Together, and Fireworks. Self-hosted deployments using these provider identifiers will trigger replay automatically.
What happens if the cache is full or an entry expires?
The LRU eviction removes oldest entries first when the 2,000-entry memory limit is reached. Expired entries (after 2 hours) are treated as cache misses. The system falls back to SQLite lookup; if absent there, the request proceeds without replayed reasoning, which may cause provider errors.
Can I disable reasoning replay for specific requests?
There is no per-request opt-out. Replay is determined solely by requiresReasoningReplay(provider, model). To disable, you would need to fork and modify open-sse/services/reasoningCache.ts to exclude your target provider or model pattern.
How do I verify that reasoning is being cached and replayed?
Query the status endpoint (/api/v1/cache/reasoning) to see hit/miss statistics. For detailed inspection, connect to the SQLite database and query the reasoning_cache table directly. Each successful replay is recorded with the original tool_call_id for traceability.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →