How Cache-Aware Context Maintenance Works with Prefix Cache in Reasonix

Cache-aware context maintenance in Reasonix works by constructing a byte-stable system-prompt prefix from immutable standing instructions (REASONIX.md, tool schemas, and memory blocks) that DeepSeek's prefix cache reuses across turns, reducing token costs by roughly 80% while maintaining conversation state through append-only tail messages.

The esengine/DeepSeek-Reasonix repository implements a sophisticated cache-aware context maintenance strategy that leverages DeepSeek's automatic prefix caching to minimize token costs. By treating the system prompt as an immutable byte sequence that persists across conversation turns, Reasonix achieves cache-hit rates exceeding 90% while ensuring that only new user messages trigger additional billing. This architecture separates stable context from mutable conversation history through a carefully managed prefix cache.

The Immutable Prefix Architecture

Reasonix constructs every LLM request around a byte-stable system-prompt prefix consisting of three immutable components. These elements are loaded once at program start and never mutated mid-session, ensuring that DeepSeek's caching layer can reuse the exact byte sequence across multiple turns.

The prefix comprises:

  • Base prompt: The core system instructions defining the agent's role, sourced from REASONIX.md and other standing instruction files
  • Tool schema: JSON definitions of available tools, generated from the internal/tool package and injected during initial composition
  • Memory block: Project-wide documentation and global guidance managed by internal/memory/Set (see internal/memory/memory.go)

According to the Reasonix source code, these components are intentionally separated from mutable conversation data. The memory.Set type handles the memory block component, while standing instruction files remain version-controlled to guarantee byte-for-byte consistency across session restarts.

Constructing the Cache-Stable Prefix

The construction process relies on two primary operations defined in internal/memory/memory.go: Load and Compose.

Loading Immutable Context

At program initialization, memory.Load discovers all standing-instruction files and the auto-memory index, producing a memory.Set structure. This function executes once per session, ensuring that the underlying bytes remain stable throughout the conversation lifecycle.

The Compose Function

The Compose function merges the base prompt with the memory block to generate the final cached prefix:

// Compose folds the memory block onto the base system prompt …
func Compose(base string, s *Set) string {
    block := s.Block()
    if block == "" {
        return base
    }
    if strings.TrimSpace(base) == "" {
        return block
    }
    return strings.TrimRight(base, "\n") + "\n\n" + block
}

The string returned by Compose becomes the cache-stable prefix that DeepSeek retains in its internal cache. This prefix is prepended to every request without modification.

Append-Only Turn Management

Once the prefix is established, Reasonix employs an append-only strategy for conversation turns. For each user interaction, the controller adds only the new user message or synthetic response after the cached prefix.

This approach ensures that:

  1. The byte sequence of the prefix remains identical across all turns
  2. DeepSeek's cache hit rate exceeds 90%
  3. Token billing is reduced to approximately one-fifth of naive implementations

Handling Mutable Data

When documents are edited via memory.Set.WriteDoc or new facts are remembered, changes are persisted to disk without touching the active prefix. These edits are queued as turn-tail notes and incorporated only when a new session initializes, preserving the current cache warmth.

Cache Invalidation Scenarios

The prefix cache is deliberately invalidated only in controlled scenarios, as implemented in desktop/session_prompt.go and enforced by CI checks.

Invalidation occurs when:

  • System-prompt upgrades: Modifications to REASONIX.md or standing-instruction files alter the base bytes, forcing a cold start
  • Provider configuration changes: Switching models or altering provider-level settings triggers a cache reset with a logged warning
  • Explicit invalidation: Direct mutations to the prefix through deliberate API calls (detected in session_prompt.go) bypass the cache

CI jobs labeled cache-impact enforce this discipline to prevent unintentional cache breaks during routine development.

Cache-Aware Compaction Strategy

For long-running sessions, Reasonix implements cache-aware compaction (documented in docs/research/cache-aware-compaction-design.md). This strategy periodically removes stale or duplicated tail content while preserving the immutable prefix.

The compaction algorithm ensures that:

  • Session history remains manageable without growing indefinitely
  • The cached prefix remains warm throughout the session lifecycle
  • Only mutable tail content is affected by cleanup operations

Implementation Example

The following example demonstrates building a prompt with a warm prefix:

package main

import (
    "fmt"
    "reasonix/internal/memory"
    "reasonix/internal/control"
)

func main() {
    // Load the immutable memory set (docs, global guidance, auto-memory)
    mem := memory.Load(memory.Options{CWD: "."})

    // Base system prompt – normally read from REASONIX.md
    base := "You are a helpful coding assistant."

    // Compose the prefix (base + memory block). This is the cached prefix.
    prefix := memory.Compose(base, mem)

    // Create a controller that will use this prefix for each turn
    ctrl := control.NewController(prefix)

    // Simulate a user turn
    userInput := "Explain how the prefix cache works."
    fullPrompt := ctrl.Compose(userInput) // returns prefix + userInput

    fmt.Println(fullPrompt)
}

Key implementation details:

  • memory.Load discovers standing docs that remain cache-stable
  • memory.Compose constructs the immutable prefix
  • controller.Compose appends new turns after the cached region

Summary

  • Reasonix builds requests around a byte-stable system-prompt prefix composed of REASONIX.md, tool schemas, and memory blocks
  • The Compose function in internal/memory/memory.go generates the immutable prefix that DeepSeek caches automatically
  • Append-only turn tails ensure cache-hit rates above 90% while reducing token costs by approximately 80%
  • Mutable data is queued as turn-tail notes and incorporated only during session restarts, preserving cache warmth
  • Cache invalidation is strictly controlled through CI checks (cache-impact labels) and explicit configuration changes
  • Cache-aware compaction manages long-running sessions by cleaning tail content without disturbing the cached prefix

Frequently Asked Questions

What is a prefix cache in the context of DeepSeek and Reasonix?

A prefix cache is DeepSeek's mechanism for storing the initial byte sequence of a conversation (the system prompt and early context) in GPU memory across multiple API calls. In Reasonix, this allows the LLM to process only new "tail" messages rather than re-processing the entire conversation history, significantly reducing latency and token costs.

How does Reasonix prevent cache invalidation during document edits?

Reasonix prevents invalidation by treating document edits as turn-tail notes rather than prefix modifications. When memory.Set.WriteDoc is called, changes are saved to disk immediately but are not incorporated into the active prefix. The current session continues using the cached prefix, and edits take effect only after a session restart when memory.Load reconstructs the prefix from the updated files.

What happens when the system prompt files are modified mid-session?

Modifying REASONIX.md or other standing-instruction files creates a byte mismatch in the prefix. According to the source code in desktop/session_prompt.go, this triggers an explicit cache invalidation, causing DeepSeek to treat the next request as a cold start. Subsequent turns miss the cache until a new stable prefix is established.

Can the prefix cache improve performance for long-running coding sessions?

Yes. Through cache-aware compaction (implemented per docs/research/cache-aware-compaction-design.md), Reasonix periodically cleans stale conversation history from the mutable tail while keeping the immutable prefix intact. This prevents context window exhaustion and maintains the performance benefits of the warm cache even during extended multi-hour coding sessions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →