How Session Memory Retrieval Works in Reasonix's Context Engine v2

Reasonix's Context Engine v2 retrieves session memory through a three-layer architecture: Session Store for persistence, Memory Indexer for vector embeddings, and a Retriever that assembles context windows respecting token budgets.

The session memory retrieval system in Reasonix is the backbone of its conversational AI capabilities. DeepSeek-Reasonix, an open-source framework for building memory-augmented LLM applications, implements this through a modular design that separates storage, indexing, and retrieval concerns. Understanding how these components interact is essential for anyone customizing or debugging the framework.

The Three-Layer Architecture

Reasonix splits session memory retrieval into three specialized components. This separation allows each layer to be optimized, scaled, or swapped independently.

Session Store: Raw Turn Persistence

The Session Store (reasonix/session_store.py) persists every conversational turn as structured data. By default, it uses SQLite, though Redis or PostgreSQL can be substituted.


# reasonix/session_store.py

store = SessionStore(db_path="sessions.db")
store.append(session_id, role="user", content=user_message)

The SessionStore.append method writes raw messages with timestamps and immediately notifies the Memory Indexer to embed the new turn.

The Memory Indexer (reasonix/memory_indexer.py) maintains vector embeddings for semantic search. It uses FAISS or Annoy for Approximate Nearest Neighbor (ANN) search, guaranteeing O(log N) query time regardless of history size.


# reasonix/memory_indexer.py

indexer = MemoryIndexer(persist_path="index.faiss")
indexer.upsert(session_id, new_message_embedding)

The MemoryIndexer.upsert method adds embeddings lazily—old embeddings remain untouched, minimizing write amplification.

ContextEngineV2: Orchestration and Token Budgeting

The Retriever layer lives in reasonix/context_engine_v2.py. Its retrieve_session method coordinates the other components:


# reasonix/context_engine_v2.py

relevant_turns = engine.retrieve_session(
    session_id=session_id,
    query_embedding=current_query_embedding,
    max_tokens=engine.max_context_tokens,
)

This method dynamically selects k nearest neighbors to fit within max_context_tokens, then reassembles them chronologically from the Session Store.

Step-by-Step Retrieval Flow

1. Persist and Embed the Incoming Turn

Every user message triggers storage and indexing:


# From session_store.py and memory_indexer.py

engine.store.append(session_id, role="user", content=user_message)

# MemoryIndexer automatically receives notification and embeds

2. Query for Relevant Historical Context

When building a prompt, the engine encodes the current query and searches:


# reasonix/context_engine_v2.py

nearest_ids = engine.indexer.query(query_embedding, k=dynamic_k)

The dynamic_k value is computed to maximize relevant context without exceeding max_context_tokens.

3. Assemble the Context Window

Retrieved turn IDs are mapped back to full messages, sorted chronologically, and truncated if necessary:


# reasonix/context_engine_v2.py

prompt = engine.build_prompt(
    system_prompt=engine.system_prompt,
    context=relevant_turns,  # Chronologically ordered

    user_message=user_message,
)

4. Execute and Store Response

After LLM generation, the assistant's reply is appended to maintain conversation continuity:

engine.store.append(session_id, role="assistant", content=response)

Why This Session Memory Design Performs

  • Scalability: Separating storage from indexing means millions of sessions can coexist without memory pressure—only active session indexes stay resident.
  • Speed: FAISS/Annoy ANN search delivers sub-millisecond retrieval even with millions of vectors.
  • Flexibility: Abstract interfaces allow swapping SQLite↔PostgreSQL↔Redis for persistence and FAISS↔Annoy↔HNSW for vectors without touching retrieval logic.

Complete Usage Example

from reasonix.context_engine_v2 import ContextEngineV2

# Initialize engine

engine = ContextEngineV2(
    store_path="sessions.db",
    index_path="index.faiss",
    system_prompt="You are a helpful assistant.",
    max_context_tokens=1500,
)

def handle_turn(session_id: str, user_message: str) -> str:
    # Store user turn

    engine.store.append(session_id, role="user", content=user_message)
    
    # Retrieve context and build prompt

    prompt = engine.build_prompt_for(session_id, user_message)
    
    # Execute LLM call

    response = engine.llm_executor.run(prompt)
    
    # Store assistant response for future context

    engine.store.append(session_id, role="assistant", content=response)
    
    return response

Direct Memory Index Querying

For debugging or custom retrieval pipelines:

from reasonix.memory_indexer import MemoryIndexer
import numpy as np

indexer = MemoryIndexer(persist_path="index.faiss")
query = "What did I ask about my last order?"
embedding = engine.embedder.encode(query)

nearest = indexer.query(embedding, k=5)  # Returns turn IDs

print("Most relevant past turns:", nearest)

Summary

  • Session memory retrieval in Reasonix uses three decoupled layers: SessionStore, MemoryIndexer, and ContextEngineV2.
  • reasonix/session_store.py handles raw persistence with pluggable backends.
  • reasonix/memory_indexer.py manages FAISS/Annoy vector indexes for semantic search.
  • reasonix/context_engine_v2.py orchestrates retrieval with token-aware context assembly.
  • The design prioritizes O(log N) query performance and horizontal scalability.

Frequently Asked Questions

How does Reasonix handle very long conversation histories?

Reasonix relies on the MemoryIndexer's ANN search to surface only semantically relevant turns rather than loading full histories. The retrieve_session method in reasonix/context_engine_v2.py dynamically calculates how many turns fit within max_context_tokens, typically selecting 5-20 most relevant turns regardless of total conversation length.

Can I replace FAISS with a different vector database?

Yes. The MemoryIndexer class in reasonix/memory_indexer.py is designed behind an interface abstraction. You can substitute FAISS with Annoy, HNSW, or cloud vector stores by implementing the same upsert and query method signatures. The ContextEngineV2 accepts any indexer instance through its constructor.

What happens if the session store and vector index get out of sync?

The SessionStore.append method triggers synchronous notification to MemoryIndexer. If embedding fails, the raw message is still stored, and retrieval gracefully degrades to exclude that turn from semantic search. A background reconciliation job (configurable in reasonix/config.py) can rebuild indexes from stored messages if corruption is detected.

How does token budgeting work when assembling prompts?

The build_prompt method in reasonix/context_engine_v2.py reserves tokens for the system prompt and current user message, then fills remaining budget with retrieved historical turns in chronological order. If the oldest retrieved turn would exceed the limit, it is truncated or excluded entirely—never the current request or system prompt.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →