How Context Engine v2 Handles Session Memory Retrieval in DeepSeek-Reasonix

The Context Engine v2 retrieves session memory through a three-tier architecture that persists conversation history in a SessionStore, indexes vector embeddings via a MemoryIndexer, and assembles relevant context windows using ANN search while enforcing LLM token limits.

The DeepSeek-Reasonix framework implements a sophisticated Context Engine v2 to manage conversational state across LLM interactions. This component orchestrates how session memory is stored, indexed, and retrieved to build contextually rich prompts. Understanding the retrieval mechanism requires examining the interaction between persistent storage, vector indexing, and dynamic context assembly.

The Three-Layer Retrieval Architecture

The session memory retrieval system separates concerns across three distinct layers to optimize for scalability and speed.

Session Store Layer

The Session Store persists raw turn-by-turn history including messages, embeddings, and timestamps. Implemented in reasonix/session_store.py, the SessionStore class provides a lightweight key-value interface backed by SQLite by default, with the flexibility to swap in Redis or PostgreSQL.

The SessionStore.append method handles write operations:


# reasonix/session_store.py

store = SessionStore(db_path="sessions.db")
store.append(session_id, role="user", content=user_message)

Memory Indexer Layer

The Memory Indexer manages vector embeddings and maintains an ANN (Approximate Nearest-Neighbour) index for relevance queries. Located in reasonix/memory_indexer.py, the MemoryIndexer class uses FAISS or Annoy to enable sub-millisecond similarity search.

The MemoryIndexer.upsert method adds new embeddings lazily after each turn:


# reasonix/memory_indexer.py

indexer = MemoryIndexer(persist_path="index.faiss")
indexer.upsert(session_id, new_message_embedding)

This design guarantees O(log N) query time without requiring full index rebuilds.

Retriever Layer

The Retriever orchestrates the final context assembly. The ContextEngineV2 class in reasonix/context_engine_v2.py implements the retrieve_session method, which queries the MemoryIndexer for relevant history and merges results with the current request while respecting the LLM's token budget.

Step-by-Step Session Memory Retrieval

The retrieval process follows a strict pipeline from persistence to prompt assembly.

  1. Persist the incoming turn

    Every user message is immediately stored via SessionStore.append, which triggers the embedding pipeline:

    # reasonix/session_store.py
    
    store = SessionStore(db_path="sessions.db")
    store.append(session_id, role="user", content=user_message)
  2. Update the vector index

    The MemoryIndexer.upsert method adds the new embedding to the FAISS index:

    # reasonix/memory_indexer.py
    
    indexer = MemoryIndexer(persist_path="index.faiss")
    indexer.upsert(session_id, new_message_embedding)
  3. Fetch relevant past turns

    The ContextEngineV2.retrieve_session method dynamically selects top-k nearest neighbors to fit within max_context_tokens:

    # reasonix/context_engine_v2.py
    
    relevant_turns = engine.retrieve_session(
        session_id=session_id,
        query_embedding=current_query_embedding,
        max_tokens=engine.max_context_tokens,
    )

    This method walks back through returned turn IDs, pulls raw messages from SessionStore, and assembles them chronologically.

  4. Assemble the final prompt

    The build_prompt method merges system instructions, retrieved context, and the current user message, truncating oldest turns if necessary:

    # reasonix/context_engine_v2.py
    
    prompt = engine.build_prompt(
        system_prompt=engine.system_prompt,
        context=relevant_turns,
        user_message=user_message,
    )
  5. Execute the LLM call

    The assembled prompt is passed to reasonix/llm_executor.py for inference.

Implementation Examples

Initializing the Context Engine

Configure the engine with storage paths and token limits:


# ----------------------------------------------------------------------

# Example 1 – Initialising the Context Engine

# ----------------------------------------------------------------------

from reasonix.context_engine_v2 import ContextEngineV2

engine = ContextEngineV2(
    store_path="sessions.db",
    index_path="index.faiss",
    system_prompt="You are a helpful assistant.",
    max_context_tokens=1500,
)

Handling a Complete Turn

Process user input and store assistant responses in a single workflow:


# ----------------------------------------------------------------------

# Example 2 – Handling a new user turn

# ----------------------------------------------------------------------

def handle_turn(session_id: str, user_message: str):
    # 1️⃣ Store the raw turn

    engine.store.append(session_id, role="user", content=user_message)

    # 2️⃣ Retrieve the enriched context

    prompt = engine.build_prompt_for(session_id, user_message)

    # 3️⃣ Send to the LLM (the executor is already wired)

    response = engine.llm_executor.run(prompt)

    # 4️⃣ Store the assistant’s reply

    engine.store.append(session_id, role="assistant", content=response)

    return response

Querying the Memory Index Directly

Access the vector index for debugging or custom retrieval logic:


# ----------------------------------------------------------------------

# Example 3 – Directly querying the memory index

# ----------------------------------------------------------------------

from reasonix.memory_indexer import MemoryIndexer
import numpy as np

indexer = MemoryIndexer(persist_path="index.faiss")
query = "What did I ask about my last order?"
embedding = engine.embedder.encode(query)          # → np.ndarray

nearest = indexer.query(embedding, k=5)            # returns turn IDs

print("Most relevant past turns:", nearest)

Key Design Advantages

  • Scalability: Separating raw storage (SessionStore) from the vector index (MemoryIndexer) allows the system to scale to millions of sessions without loading full histories into memory.
  • Speed: ANN search via FAISS or Annoy yields sub-millisecond retrieval latency, critical for real-time chat applications.
  • Flexibility: Abstracted interfaces enable swapping SQLite for PostgreSQL or Redis on the persistence layer, and FAISS for Annoy or HNSW on the vector layer.

Summary

  • The Context Engine v2 implements a three-tier architecture comprising SessionStore, MemoryIndexer, and ContextEngineV2 classes.
  • Session memory retrieval combines persistent key-value storage with vector similarity search to identify relevant historical turns.
  • The retrieve_session method in reasonix/context_engine_v2.py dynamically balances relevance against token budget constraints.
  • All components utilize pluggable backends, allowing deployment customization without modifying core retrieval logic.

Frequently Asked Questions

What storage backends does the Session Store support?

The SessionStore class in reasonix/session_store.py uses SQLite by default but exposes an abstraction interface that supports Redis and PostgreSQL. This allows operators to choose between lightweight local development setups and production-grade distributed caches.

How does the Memory Indexer handle embedding updates?

The MemoryIndexer.upsert method implements lazy refresh logic, adding new embeddings to the existing FAISS or Annoy index without requiring a full rebuild. This approach maintains O(log N) query complexity and ensures minimal latency during active conversation turns.

What determines how many past turns are retrieved?

The ContextEngineV2.retrieve_session method dynamically calculates the top-k nearest neighbors based on the max_context_tokens parameter. It queries the vector index for candidate turns, then filters and truncates the chronological assembly to fit within the LLM's context window.

Where does the final prompt assembly occur?

Final prompt construction occurs in reasonix/context_engine_v2.py via the build_prompt method. This function merges the system prompt, retrieved session context, and current user message, automatically truncating older turns if the total token count exceeds the configured limit.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →