How Session Memory Retrieval Works in Reasonix's Context Engine v2
Reasonix's Context Engine v2 retrieves session memory through a three-layer architecture: Session Store for persistence, Memory Indexer for vector embeddings, and a Retriever that assembles context windows respecting token budgets.
The session memory retrieval system in Reasonix is the backbone of its conversational AI capabilities. DeepSeek-Reasonix, an open-source framework for building memory-augmented LLM applications, implements this through a modular design that separates storage, indexing, and retrieval concerns. Understanding how these components interact is essential for anyone customizing or debugging the framework.
The Three-Layer Architecture
Reasonix splits session memory retrieval into three specialized components. This separation allows each layer to be optimized, scaled, or swapped independently.
Session Store: Raw Turn Persistence
The Session Store (reasonix/session_store.py) persists every conversational turn as structured data. By default, it uses SQLite, though Redis or PostgreSQL can be substituted.
# reasonix/session_store.py
store = SessionStore(db_path="sessions.db")
store.append(session_id, role="user", content=user_message)
The SessionStore.append method writes raw messages with timestamps and immediately notifies the Memory Indexer to embed the new turn.
Memory Indexer: Vector Embedding and ANN Search
The Memory Indexer (reasonix/memory_indexer.py) maintains vector embeddings for semantic search. It uses FAISS or Annoy for Approximate Nearest Neighbor (ANN) search, guaranteeing O(log N) query time regardless of history size.
# reasonix/memory_indexer.py
indexer = MemoryIndexer(persist_path="index.faiss")
indexer.upsert(session_id, new_message_embedding)
The MemoryIndexer.upsert method adds embeddings lazily—old embeddings remain untouched, minimizing write amplification.
ContextEngineV2: Orchestration and Token Budgeting
The Retriever layer lives in reasonix/context_engine_v2.py. Its retrieve_session method coordinates the other components:
# reasonix/context_engine_v2.py
relevant_turns = engine.retrieve_session(
session_id=session_id,
query_embedding=current_query_embedding,
max_tokens=engine.max_context_tokens,
)
This method dynamically selects k nearest neighbors to fit within max_context_tokens, then reassembles them chronologically from the Session Store.
Step-by-Step Retrieval Flow
1. Persist and Embed the Incoming Turn
Every user message triggers storage and indexing:
# From session_store.py and memory_indexer.py
engine.store.append(session_id, role="user", content=user_message)
# MemoryIndexer automatically receives notification and embeds
2. Query for Relevant Historical Context
When building a prompt, the engine encodes the current query and searches:
# reasonix/context_engine_v2.py
nearest_ids = engine.indexer.query(query_embedding, k=dynamic_k)
The dynamic_k value is computed to maximize relevant context without exceeding max_context_tokens.
3. Assemble the Context Window
Retrieved turn IDs are mapped back to full messages, sorted chronologically, and truncated if necessary:
# reasonix/context_engine_v2.py
prompt = engine.build_prompt(
system_prompt=engine.system_prompt,
context=relevant_turns, # Chronologically ordered
user_message=user_message,
)
4. Execute and Store Response
After LLM generation, the assistant's reply is appended to maintain conversation continuity:
engine.store.append(session_id, role="assistant", content=response)
Why This Session Memory Design Performs
- Scalability: Separating storage from indexing means millions of sessions can coexist without memory pressure—only active session indexes stay resident.
- Speed: FAISS/Annoy ANN search delivers sub-millisecond retrieval even with millions of vectors.
- Flexibility: Abstract interfaces allow swapping SQLite↔PostgreSQL↔Redis for persistence and FAISS↔Annoy↔HNSW for vectors without touching retrieval logic.
Complete Usage Example
from reasonix.context_engine_v2 import ContextEngineV2
# Initialize engine
engine = ContextEngineV2(
store_path="sessions.db",
index_path="index.faiss",
system_prompt="You are a helpful assistant.",
max_context_tokens=1500,
)
def handle_turn(session_id: str, user_message: str) -> str:
# Store user turn
engine.store.append(session_id, role="user", content=user_message)
# Retrieve context and build prompt
prompt = engine.build_prompt_for(session_id, user_message)
# Execute LLM call
response = engine.llm_executor.run(prompt)
# Store assistant response for future context
engine.store.append(session_id, role="assistant", content=response)
return response
Direct Memory Index Querying
For debugging or custom retrieval pipelines:
from reasonix.memory_indexer import MemoryIndexer
import numpy as np
indexer = MemoryIndexer(persist_path="index.faiss")
query = "What did I ask about my last order?"
embedding = engine.embedder.encode(query)
nearest = indexer.query(embedding, k=5) # Returns turn IDs
print("Most relevant past turns:", nearest)
Summary
- Session memory retrieval in Reasonix uses three decoupled layers:
SessionStore,MemoryIndexer, andContextEngineV2. reasonix/session_store.pyhandles raw persistence with pluggable backends.reasonix/memory_indexer.pymanages FAISS/Annoy vector indexes for semantic search.reasonix/context_engine_v2.pyorchestrates retrieval with token-aware context assembly.- The design prioritizes O(log N) query performance and horizontal scalability.
Frequently Asked Questions
How does Reasonix handle very long conversation histories?
Reasonix relies on the MemoryIndexer's ANN search to surface only semantically relevant turns rather than loading full histories. The retrieve_session method in reasonix/context_engine_v2.py dynamically calculates how many turns fit within max_context_tokens, typically selecting 5-20 most relevant turns regardless of total conversation length.
Can I replace FAISS with a different vector database?
Yes. The MemoryIndexer class in reasonix/memory_indexer.py is designed behind an interface abstraction. You can substitute FAISS with Annoy, HNSW, or cloud vector stores by implementing the same upsert and query method signatures. The ContextEngineV2 accepts any indexer instance through its constructor.
What happens if the session store and vector index get out of sync?
The SessionStore.append method triggers synchronous notification to MemoryIndexer. If embedding fails, the raw message is still stored, and retrieval gracefully degrades to exclude that turn from semantic search. A background reconciliation job (configurable in reasonix/config.py) can rebuild indexes from stored messages if corruption is detected.
How does token budgeting work when assembling prompts?
The build_prompt method in reasonix/context_engine_v2.py reserves tokens for the system prompt and current user message, then fills remaining budget with retrieved historical turns in chronological order. If the oldest retrieved turn would exceed the limit, it is truncated or excluded entirely—never the current request or system prompt.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →