# How Cache-Aware Context Maintenance Works with Prefix Cache in Reasonix

> Discover how Reasonix's cache-aware context maintenance uses prefix cache for efficient conversation state management. Reduce token costs by 80% with this innovative technique.

- Repository: [YHH/DeepSeek-Reasonix](https://github.com/esengine/DeepSeek-Reasonix)
- Tags: internals
- Published: 2026-08-08

---

**Cache-aware context maintenance in Reasonix works by constructing a byte-stable system-prompt prefix from immutable standing instructions (REASONIX.md, tool schemas, and memory blocks) that DeepSeek's prefix cache reuses across turns, reducing token costs by roughly 80% while maintaining conversation state through append-only tail messages.**

The `esengine/DeepSeek-Reasonix` repository implements a sophisticated **cache-aware context maintenance** strategy that leverages DeepSeek's automatic prefix caching to minimize token costs. By treating the system prompt as an immutable byte sequence that persists across conversation turns, Reasonix achieves cache-hit rates exceeding 90% while ensuring that only new user messages trigger additional billing. This architecture separates stable context from mutable conversation history through a carefully managed prefix cache.

## The Immutable Prefix Architecture

Reasonix constructs every LLM request around a **byte-stable system-prompt prefix** consisting of three immutable components. These elements are loaded once at program start and never mutated mid-session, ensuring that DeepSeek's caching layer can reuse the exact byte sequence across multiple turns.

The prefix comprises:

- **Base prompt**: The core system instructions defining the agent's role, sourced from [`REASONIX.md`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/REASONIX.md) and other standing instruction files
- **Tool schema**: JSON definitions of available tools, generated from the `internal/tool` package and injected during initial composition
- **Memory block**: Project-wide documentation and global guidance managed by `internal/memory/Set` (see [`internal/memory/memory.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/memory/memory.go))

According to the Reasonix source code, these components are intentionally separated from mutable conversation data. The `memory.Set` type handles the memory block component, while standing instruction files remain version-controlled to guarantee byte-for-byte consistency across session restarts.

## Constructing the Cache-Stable Prefix

The construction process relies on two primary operations defined in [`internal/memory/memory.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/memory/memory.go): `Load` and `Compose`.

### Loading Immutable Context

At program initialization, `memory.Load` discovers all standing-instruction files and the auto-memory index, producing a `memory.Set` structure. This function executes once per session, ensuring that the underlying bytes remain stable throughout the conversation lifecycle.

### The Compose Function

The `Compose` function merges the base prompt with the memory block to generate the final cached prefix:

```go
// Compose folds the memory block onto the base system prompt …
func Compose(base string, s *Set) string {
    block := s.Block()
    if block == "" {
        return base
    }
    if strings.TrimSpace(base) == "" {
        return block
    }
    return strings.TrimRight(base, "\n") + "\n\n" + block
}

```

The string returned by `Compose` becomes the **cache-stable prefix** that DeepSeek retains in its internal cache. This prefix is prepended to every request without modification.

## Append-Only Turn Management

Once the prefix is established, Reasonix employs an **append-only strategy** for conversation turns. For each user interaction, the controller adds only the new user message or synthetic response after the cached prefix.

This approach ensures that:

1. The byte sequence of the prefix remains identical across all turns
2. DeepSeek's cache hit rate exceeds 90%
3. Token billing is reduced to approximately one-fifth of naive implementations

### Handling Mutable Data

When documents are edited via `memory.Set.WriteDoc` or new facts are remembered, changes are persisted to disk **without touching the active prefix**. These edits are queued as turn-tail notes and incorporated only when a new session initializes, preserving the current cache warmth.

## Cache Invalidation Scenarios

The prefix cache is deliberately invalidated only in controlled scenarios, as implemented in [`desktop/session_prompt.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/desktop/session_prompt.go) and enforced by CI checks.

Invalidation occurs when:

- **System-prompt upgrades**: Modifications to [`REASONIX.md`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/REASONIX.md) or standing-instruction files alter the base bytes, forcing a cold start
- **Provider configuration changes**: Switching models or altering provider-level settings triggers a cache reset with a logged warning
- **Explicit invalidation**: Direct mutations to the prefix through deliberate API calls (detected in [`session_prompt.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/session_prompt.go)) bypass the cache

CI jobs labeled `cache-impact` enforce this discipline to prevent unintentional cache breaks during routine development.

## Cache-Aware Compaction Strategy

For long-running sessions, Reasonix implements **cache-aware compaction** (documented in [`docs/research/cache-aware-compaction-design.md`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/docs/research/cache-aware-compaction-design.md)). This strategy periodically removes stale or duplicated tail content while preserving the immutable prefix.

The compaction algorithm ensures that:

- Session history remains manageable without growing indefinitely
- The cached prefix remains warm throughout the session lifecycle
- Only mutable tail content is affected by cleanup operations

## Implementation Example

The following example demonstrates building a prompt with a warm prefix:

```go
package main

import (
    "fmt"
    "reasonix/internal/memory"
    "reasonix/internal/control"
)

func main() {
    // Load the immutable memory set (docs, global guidance, auto-memory)
    mem := memory.Load(memory.Options{CWD: "."})

    // Base system prompt – normally read from REASONIX.md
    base := "You are a helpful coding assistant."

    // Compose the prefix (base + memory block). This is the cached prefix.
    prefix := memory.Compose(base, mem)

    // Create a controller that will use this prefix for each turn
    ctrl := control.NewController(prefix)

    // Simulate a user turn
    userInput := "Explain how the prefix cache works."
    fullPrompt := ctrl.Compose(userInput) // returns prefix + userInput

    fmt.Println(fullPrompt)
}

```

Key implementation details:

- `memory.Load` discovers standing docs that remain cache-stable
- `memory.Compose` constructs the immutable prefix
- `controller.Compose` appends new turns after the cached region

## Summary

- Reasonix builds requests around a **byte-stable system-prompt prefix** composed of [`REASONIX.md`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/REASONIX.md), tool schemas, and memory blocks
- The `Compose` function in [`internal/memory/memory.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/internal/memory/memory.go) generates the immutable prefix that DeepSeek caches automatically
- **Append-only turn tails** ensure cache-hit rates above 90% while reducing token costs by approximately 80%
- Mutable data is queued as turn-tail notes and incorporated only during session restarts, preserving cache warmth
- Cache invalidation is strictly controlled through CI checks (`cache-impact` labels) and explicit configuration changes
- **Cache-aware compaction** manages long-running sessions by cleaning tail content without disturbing the cached prefix

## Frequently Asked Questions

### What is a prefix cache in the context of DeepSeek and Reasonix?

A prefix cache is DeepSeek's mechanism for storing the initial byte sequence of a conversation (the system prompt and early context) in GPU memory across multiple API calls. In Reasonix, this allows the LLM to process only new "tail" messages rather than re-processing the entire conversation history, significantly reducing latency and token costs.

### How does Reasonix prevent cache invalidation during document edits?

Reasonix prevents invalidation by treating document edits as **turn-tail notes** rather than prefix modifications. When `memory.Set.WriteDoc` is called, changes are saved to disk immediately but are not incorporated into the active prefix. The current session continues using the cached prefix, and edits take effect only after a session restart when `memory.Load` reconstructs the prefix from the updated files.

### What happens when the system prompt files are modified mid-session?

Modifying [`REASONIX.md`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/REASONIX.md) or other standing-instruction files creates a byte mismatch in the prefix. According to the source code in [`desktop/session_prompt.go`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/desktop/session_prompt.go), this triggers an explicit cache invalidation, causing DeepSeek to treat the next request as a cold start. Subsequent turns miss the cache until a new stable prefix is established.

### Can the prefix cache improve performance for long-running coding sessions?

Yes. Through **cache-aware compaction** (implemented per [`docs/research/cache-aware-compaction-design.md`](https://github.com/esengine/DeepSeek-Reasonix/blob/main/docs/research/cache-aware-compaction-design.md)), Reasonix periodically cleans stale conversation history from the mutable tail while keeping the immutable prefix intact. This prevents context window exhaustion and maintains the performance benefits of the warm cache even during extended multi-hour coding sessions.