Memory Management Strategies in Qwen-Agent: A Deep Dive into RAG-Based Knowledge Handling
Qwen-Agent implements a two-layer memory system using Retrieval-Augmented Generation (RAG) with configurable token budgets, hybrid search strategies, and dynamic keyword generation to manage external knowledge efficiently.
The Qwen-Agent framework provides sophisticated memory management strategies designed to handle external knowledge sources during conversational AI interactions. At its core, the Memory class in qwen_agent/memory/memory.py orchestrates a flexible RAG pipeline that ingests, stores, and retrieves information from files and web pages. This system creates a unified interface for both short-term session awareness and long-term system knowledge.
Core Memory Architecture
The memory system distinguishes between two primary knowledge layers, each handled differently within the retrieval pipeline.
System-Level vs Session-Level Files
System-level files represent static knowledge bases supplied when the Memory object is instantiated. These files are initialized in self.system_files at the end of Memory.__init__ in qwen_agent/memory/memory.py (lines 78-80). They persist across the entire conversation lifecycle.
Session files are dynamic uploads extracted from individual chat messages. The method extract_files_from_messages inside get_rag_files (lines 46-48) identifies these temporary knowledge sources, allowing users to inject context-specific documents during active conversations.
RAG Configuration and Token Budgeting
The rag_cfg parameter controls retrieval behavior through the self.cfg dictionary initialized in Memory.__init__ (lines 55-61). Key configuration options include:
max_ref_token: Maximum tokens allowed for retrieved content (default 20,000 inqwen_agent/settings.py, lines 29-40)parser_page_size: Chunk size for document parsingrag_searchers: List of search strategies (default:['keyword_search', 'front_page_search'])rag_keygen_strategy: Keyword generation approach (e.g.,"SplitQueryThenGenKeyword"or"none")
The Retrieval Pipeline
The memory system executes a five-stage retrieval workflow within the _run method (lines 132-141).
File Collection and Filtering
get_rag_files merges system and session files, filtering by supported extensions defined in PARSER_SUPPORTED_FILE_TYPES. This ensures only parseable documents enter the retrieval pipeline.
Keyword Generation Strategies
When rag_keygen_strategy is not "none", the system dynamically loads a key-generation agent from qwen_agent.agents.keygen_strategies (lines 106-122). This agent rewrites user queries into JSON-encoded keyword sets optimized for retrieval, implementing strategies like SplitQueryThenGenKeyword for complex query decomposition.
Hybrid Search Implementation
The retrieval tool (wrapped in self.function_map['retrieval']) executes document parsing and hybrid search combining BM25 with vector-based methods. The DocParser in qwen_agent/tools/retrieval.py (lines 68-70) handles chunking according to parser_page_size while respecting max_ref_token limits.
Results return as Message objects with role=ASSISTANT and name='memory', containing JSON-encoded snippet lists.
Integration with Agent Framework
Higher-level agents automatically instantiate Memory objects to handle file operations without manual configuration.
Automatic Memory Instantiation
In qwen_agent/agents/fncall_agent.py (lines 56-72), FnCallAgent.__init__ automatically creates self.mem = Memory(...) when file handling capabilities are required. This integration allows function-calling agents to seamlessly access the RAG pipeline.
VirtualMemoryAgent Wrapper
The VirtualMemoryAgent class in qwen_agent/agents/virtual_memory_agent.py provides a specialized wrapper that exposes Memory functionality as a retrieval-only tool. This agent is demonstrated in examples/virtual_memory_qa.py, showing how to pre-load system files for persistent knowledge access.
Configuration and Customization
Basic usage requires only a model specification and file path:
from qwen_agent.memory import Memory
from qwen_agent.llm.schema import ContentItem, Message
mem = Memory(llm={'model': 'qwen-max'})
messages = [
Message('user', [
ContentItem(text='How does the algorithm work?'),
ContentItem(file='examples/resource/doc.pdf')
])
]
_, _, last = mem.run(messages, max_ref_token=4000, parser_page_size=500)
print(last[-1].content)
Customize retrieval behavior through the rag_cfg dictionary:
custom_cfg = {
'max_ref_token': 10000,
'parser_page_size': 300,
'rag_keygen_strategy': 'SplitQueryThenGenKeyword',
'rag_searchers': ['keyword_search']
}
mem = Memory(llm={'model': 'qwen-max'}, rag_cfg=custom_cfg)
For higher-level integration, use VirtualMemoryAgent with pre-loaded files:
from qwen_agent.agents import VirtualMemoryAgent
bot = VirtualMemoryAgent(
llm={'model': 'qwen-max'},
files=['examples/resource/doc.pdf']
)
response = bot.run(
[Message('user', [ContentItem(text='Summarize the PDF')])]
)
Summary
- Two-layer architecture: Qwen-Agent separates persistent system files from temporary session uploads, managed through
self.system_filesandget_rag_filesinqwen_agent/memory/memory.py. - Configurable RAG pipeline: Token budgets (
max_ref_token), chunk sizes (parser_page_size), and search strategies (rag_searchers) control retrieval precision and context window usage. - Dynamic keyword generation: The
rag_keygen_strategyparameter enables query optimization through strategies likeSplitQueryThenGenKeywordbefore hybrid search execution. - Seamless agent integration:
FnCallAgentandVirtualMemoryAgentautomatically instantiateMemoryobjects, exposing retrieval capabilities without manual pipeline configuration.
Frequently Asked Questions
How does Qwen-Agent handle file uploads during active conversations?
Session files are extracted dynamically from message content using extract_files_from_messages within get_rag_files (lines 46-48 of qwen_agent/memory/memory.py). These temporary files are merged with persistent system files before retrieval, allowing users to inject context-specific documents without restarting the conversation.
What is the default token budget for retrieved content, and how can it be modified?
The default maximum reference token count is 20,000 tokens, defined as DEFAULT_MAX_REF_TOKEN in qwen_agent/settings.py (lines 29-40). Override this limit by passing max_ref_token to Memory.run() or configuring rag_cfg={'max_ref_token': 10000} during Memory initialization.
Which search algorithms does the retrieval system use?
Qwen-Agent employs a hybrid search approach combining BM25 keyword search with vector-based methods. The default configuration uses ['keyword_search', 'front_page_search'] as specified in rag_searchers, though developers can customize this list through the rag_cfg parameter to include additional searchers from qwen_agent/tools/search_tools/.
How does keyword generation improve retrieval accuracy?
When rag_keygen_strategy is set to strategies like SplitQueryThenGenKeyword, the system dynamically loads a specialized agent from qwen_agent/agents/keygen_strategies (lines 106-122 of qwen_agent/memory/memory.py). This agent decomposes complex queries into optimized keyword sets before retrieval, improving the relevance of retrieved passages compared to raw query matching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →