Conversation Memory System Architecture in CodeWiki: A Technical Deep Dive

The conversation memory system architecture in CodeWiki uses a three-layer design built on AdalFlow's DataComponent pattern, combining DialogTurn objects, a CustomConversation container, and a Memory wrapper to maintain stateful LLM interactions without external database dependencies.

The conversation memory system architecture in CodeWiki enables persistent, context-aware dialogue between users and the LLM backend. Built on top of the AdalFlow framework, this lightweight implementation tracks multi-turn exchanges by storing the complete interaction history in memory for the duration of the session. The system centers around three core classes defined in api/rag.py that work together to capture, manage, and retrieve conversational context.

Core Components of the Conversation Memory System Architecture

DialogTurn: The Atomic Exchange Unit

The DialogTurn dataclass represents a single exchange in the conversation. Defined in api/rag.py, it encapsulates the three essential pieces of any dialogue turn:


# api/rag.py

@dataclass
class DialogTurn:
    id: str
    user_query: UserQuery
    assistant_response: AssistantResponse

Each DialogTurn receives a unique identifier upon creation, allowing the system to reference specific exchanges when building context windows or debugging conversation flow.

CustomConversation: Safe List Management

The CustomConversation class provides a robust container for DialogTurn objects. Unlike a standard Python list, this implementation includes defensive programming patterns to prevent runtime errors:


# api/rag.py

class CustomConversation:
    """Custom implementation of Conversation to fix the list assignment index out of range error"""
    def __init__(self):
        self.dialog_turns = []

    def append_dialog_turn(self, dialog_turn):
        """Safely append a dialog turn to the conversation"""
        if not hasattr(self, "dialog_turns"):
            self.dialog_turns = []
        self.dialog_turns.append(dialog_turn)

The append_dialog_turn method explicitly checks for the dialog_turns attribute before appending, eliminating the "list assignment index out of range" errors that can occur with dynamic attribute assignment in Python.

Memory: The AdalFlow DataComponent Interface

The Memory class serves as the primary interface between the conversation storage system and the AdalFlow framework. By extending adalflow.core.component.DataComponent, it integrates seamlessly with AdalFlow's pipeline architecture:


# api/rag.py

class Memory(adalflow.core.component.DataComponent):
    def __init__(self):
        super().__init__()
        self.current_conversation = CustomConversation()

    def call(self):
        all_diaglog_turns = {}
        # … iterate over current_conversation.dialog_turns, populate dict …

        return all_diaglog_turns

    def add_dialog_turn(self, user_query: str, assistant_response: str):
        # … create DialogTurn, ensure conversation container exists, append …

        return True

The call() method returns the dialogue history as a dictionary keyed by turn IDs, while add_dialog_turn handles the instantiation of new DialogTurn objects and their safe insertion into the conversation chain.

How the Conversation Memory System Processes Requests

The Request-Response Lifecycle

The conversation memory system architecture follows a strict lifecycle during each interaction:

  1. Request Reception: The streaming chat endpoint in api/stream_chat.py receives the user query
  2. History Retrieval: The endpoint accesses stored dialogue via request_rag.memory()
  3. Context Construction: The system iterates over memory entries to build the conversation history string
  4. Prompt Assembly: System prompt, conversation history, retrieved context, and current query are combined
  5. Response Generation: The LLM generates a response based on the full context
  6. Persistence: The server calls memory.add_dialog_turn() to store the new exchange for subsequent requests

Building the Prompt with Historical Context

The streaming endpoint constructs the conversation history using XML-like tags to delimit turns:


# Inside api/stream_chat.py (lines 48-58)

conversation_history = ""
for turn_id, turn in request_rag.memory().items():
    conversation_history += (
        f"<turn>\n<user>{turn.user_query.query_str}</user>\n"
        f"<assistant>{turn.assistant_response.response_str}</assistant>\n</turn>\n"
    )

This structured format ensures the LLM can distinguish between user inputs and assistant outputs across multiple turns, maintaining coherent multi-session context.

Working with the Conversation Memory System

Adding a Turn Manually

You can interact with the memory system directly for testing or custom implementations:

from api.rag import Memory

mem = Memory()
mem.add_dialog_turn(
    user_query="How does the embedder work?",
    assistant_response="The embedder converts text to a vector using the Ollama model."
)

After execution, mem() returns a dictionary containing the new turn indexed by its generated UUID.

Retrieving Full History

Access the complete conversation history for inspection or debugging:

history = mem()
for turn_id, turn in history.items():
    print(f"Turn {turn_id}:")
    print("User →", turn.user_query.query_str)
    print("Assistant →", turn.assistant_response.response_str)

This iteration pattern mirrors the approach used in the streaming endpoint to serialize history for LLM prompts.

Integration in the Streaming Endpoint

The production implementation demonstrates how to consume the memory system within an async streaming context:


# Inside api/stream_chat.py

conversation_history = ""
for turn_id, turn in request_rag.memory().items():
    conversation_history += (
        f"<turn>\n<user>{turn.user_query.query_str}</user>\n"
        f"<assistant>{turn.assistant_response.response_str}</assistant>\n</turn>\n"
    )

This construction feeds directly into the prompt template that the LLM receives, ensuring contextual continuity across turns.

Key Files in the Conversation Memory System Architecture

File Role
api/rag.py Defines DialogTurn, CustomConversation, and Memory classes; implements the core storage logic and AdalFlow integration.
api/stream_chat.py Consumes the Memory component to retrieve history and construct prompts for the streaming chat endpoint.
utils/logger.py Provides diagnostic logging used throughout the memory component for debugging conversation flow.
utils/localdb_manager.py Supports the RAG pipeline by fetching document context that is combined with stored conversation turns in the final prompt.

These files collectively implement the complete conversation memory lifecycle—from creation and safe storage to retrieval and prompt integration.

Summary

  • Three-layer architecture: The system uses DialogTurn for atomic storage, CustomConversation for safe list management, and Memory for AdalFlow integration.
  • AdalFlow compatibility: The Memory class extends DataComponent, enabling seamless pipeline integration and callable interface.
  • Session-based persistence: Conversation history persists in memory for the duration of the backend process, with no external database required.
  • Defensive programming: CustomConversation includes safety checks to prevent index out of range errors during dynamic attribute access.
  • Prompt integration: The streaming endpoint in api/stream_chat.py serializes history into XML-like tags for LLM context construction.

Frequently Asked Questions

How does the conversation memory system architecture handle concurrent user sessions?

Each user session maintains its own RAG component instance, which contains an isolated Memory object. Since the Memory class stores conversation data in instance attributes (self.current_conversation), concurrent requests operate on separate memory spaces without cross-contamination. The architecture assumes session state is managed at the application level above the memory component.

What is the difference between CustomConversation and a standard Python list?

While CustomConversation uses a Python list internally (self.dialog_turns), it provides the append_dialog_turn method that includes a safety check: if not hasattr(self, "dialog_turns"): self.dialog_turns = []. This defensive pattern prevents the "list assignment index out of range" errors that can occur when objects are deserialized or when attributes are accessed before initialization, making it more robust than a raw list for long-running conversation tracking.

How long does conversation history persist in CodeWiki?

The conversation history persists only for the duration of the backend process session. Since the Memory class stores data in memory (specifically in self.current_conversation.dialog_turns), the history is lost when the server restarts or the process terminates. There is no built-in persistence to disk, database, or external cache like Redis in the current implementation.

Can I extend the Memory component to use external storage?

Yes, because Memory extends adalflow.core.component.DataComponent, you can override the call() and add_dialog_turn() methods to integrate external storage. For example, you could modify add_dialog_turn to write to a PostgreSQL database or Redis cache, and update call() to retrieve history from that external store rather than self.current_conversation. The AdalFlow base class ensures the component remains compatible with the rest of the pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →