# Memory Management Strategies in Qwen-Agent: A Deep Dive into RAG-Based Knowledge Handling

> Explore Qwen-Agent's memory management strategies. Learn how RAG, token budgets, and hybrid search optimize external knowledge handling for efficient agent performance.

- Repository: [Qwen/Qwen-Agent](https://github.com/qwenlm/Qwen-Agent)
- Tags: deep-dive
- Published: 2026-03-09

---

**Qwen-Agent implements a two-layer memory system using Retrieval-Augmented Generation (RAG) with configurable token budgets, hybrid search strategies, and dynamic keyword generation to manage external knowledge efficiently.**

The Qwen-Agent framework provides sophisticated memory management strategies designed to handle external knowledge sources during conversational AI interactions. At its core, the `Memory` class in [`qwen_agent/memory/memory.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/memory/memory.py) orchestrates a flexible RAG pipeline that ingests, stores, and retrieves information from files and web pages. This system creates a unified interface for both short-term session awareness and long-term system knowledge.

## Core Memory Architecture

The memory system distinguishes between two primary knowledge layers, each handled differently within the retrieval pipeline.

### System-Level vs Session-Level Files

**System-level files** represent static knowledge bases supplied when the `Memory` object is instantiated. These files are initialized in `self.system_files` at the end of `Memory.__init__` in [`qwen_agent/memory/memory.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/memory/memory.py) (lines 78-80). They persist across the entire conversation lifecycle.

**Session files** are dynamic uploads extracted from individual chat messages. The method `extract_files_from_messages` inside `get_rag_files` (lines 46-48) identifies these temporary knowledge sources, allowing users to inject context-specific documents during active conversations.

### RAG Configuration and Token Budgeting

The `rag_cfg` parameter controls retrieval behavior through the [`self.cfg`](https://github.com/QwenLM/Qwen-Agent/blob/main/self.cfg) dictionary initialized in `Memory.__init__` (lines 55-61). Key configuration options include:

- **`max_ref_token`**: Maximum tokens allowed for retrieved content (default 20,000 in [`qwen_agent/settings.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/settings.py), lines 29-40)
- **`parser_page_size`**: Chunk size for document parsing
- **`rag_searchers`**: List of search strategies (default: `['keyword_search', 'front_page_search']`)
- **`rag_keygen_strategy`**: Keyword generation approach (e.g., `"SplitQueryThenGenKeyword"` or `"none"`)

## The Retrieval Pipeline

The memory system executes a five-stage retrieval workflow within the `_run` method (lines 132-141).

### File Collection and Filtering

`get_rag_files` merges system and session files, filtering by supported extensions defined in `PARSER_SUPPORTED_FILE_TYPES`. This ensures only parseable documents enter the retrieval pipeline.

### Keyword Generation Strategies

When `rag_keygen_strategy` is not `"none"`, the system dynamically loads a key-generation agent from `qwen_agent.agents.keygen_strategies` (lines 106-122). This agent rewrites user queries into JSON-encoded keyword sets optimized for retrieval, implementing strategies like `SplitQueryThenGenKeyword` for complex query decomposition.

### Hybrid Search Implementation

The `retrieval` tool (wrapped in `self.function_map['retrieval']`) executes document parsing and hybrid search combining BM25 with vector-based methods. The `DocParser` in [`qwen_agent/tools/retrieval.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/tools/retrieval.py) (lines 68-70) handles chunking according to `parser_page_size` while respecting `max_ref_token` limits.

Results return as `Message` objects with `role=ASSISTANT` and `name='memory'`, containing JSON-encoded snippet lists.

## Integration with Agent Framework

Higher-level agents automatically instantiate `Memory` objects to handle file operations without manual configuration.

### Automatic Memory Instantiation

In [`qwen_agent/agents/fncall_agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/fncall_agent.py) (lines 56-72), `FnCallAgent.__init__` automatically creates `self.mem = Memory(...)` when file handling capabilities are required. This integration allows function-calling agents to seamlessly access the RAG pipeline.

### VirtualMemoryAgent Wrapper

The `VirtualMemoryAgent` class in [`qwen_agent/agents/virtual_memory_agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/virtual_memory_agent.py) provides a specialized wrapper that exposes `Memory` functionality as a retrieval-only tool. This agent is demonstrated in [`examples/virtual_memory_qa.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/examples/virtual_memory_qa.py), showing how to pre-load system files for persistent knowledge access.

## Configuration and Customization

Basic usage requires only a model specification and file path:

```python
from qwen_agent.memory import Memory
from qwen_agent.llm.schema import ContentItem, Message

mem = Memory(llm={'model': 'qwen-max'})

messages = [
    Message('user', [
        ContentItem(text='How does the algorithm work?'),
        ContentItem(file='examples/resource/doc.pdf')
    ])
]

_, _, last = mem.run(messages, max_ref_token=4000, parser_page_size=500)
print(last[-1].content)

```

Customize retrieval behavior through the `rag_cfg` dictionary:

```python
custom_cfg = {
    'max_ref_token': 10000,
    'parser_page_size': 300,
    'rag_keygen_strategy': 'SplitQueryThenGenKeyword',
    'rag_searchers': ['keyword_search']
}

mem = Memory(llm={'model': 'qwen-max'}, rag_cfg=custom_cfg)

```

For higher-level integration, use `VirtualMemoryAgent` with pre-loaded files:

```python
from qwen_agent.agents import VirtualMemoryAgent

bot = VirtualMemoryAgent(
    llm={'model': 'qwen-max'},
    files=['examples/resource/doc.pdf']
)

response = bot.run(
    [Message('user', [ContentItem(text='Summarize the PDF')])]
)

```

## Summary

- **Two-layer architecture**: Qwen-Agent separates persistent system files from temporary session uploads, managed through `self.system_files` and `get_rag_files` in [`qwen_agent/memory/memory.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/memory/memory.py).
- **Configurable RAG pipeline**: Token budgets (`max_ref_token`), chunk sizes (`parser_page_size`), and search strategies (`rag_searchers`) control retrieval precision and context window usage.
- **Dynamic keyword generation**: The `rag_keygen_strategy` parameter enables query optimization through strategies like `SplitQueryThenGenKeyword` before hybrid search execution.
- **Seamless agent integration**: `FnCallAgent` and `VirtualMemoryAgent` automatically instantiate `Memory` objects, exposing retrieval capabilities without manual pipeline configuration.

## Frequently Asked Questions

### How does Qwen-Agent handle file uploads during active conversations?

Session files are extracted dynamically from message content using `extract_files_from_messages` within `get_rag_files` (lines 46-48 of [`qwen_agent/memory/memory.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/memory/memory.py)). These temporary files are merged with persistent system files before retrieval, allowing users to inject context-specific documents without restarting the conversation.

### What is the default token budget for retrieved content, and how can it be modified?

The default maximum reference token count is 20,000 tokens, defined as `DEFAULT_MAX_REF_TOKEN` in [`qwen_agent/settings.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/settings.py) (lines 29-40). Override this limit by passing `max_ref_token` to `Memory.run()` or configuring `rag_cfg={'max_ref_token': 10000}` during `Memory` initialization.

### Which search algorithms does the retrieval system use?

Qwen-Agent employs a hybrid search approach combining BM25 keyword search with vector-based methods. The default configuration uses `['keyword_search', 'front_page_search']` as specified in `rag_searchers`, though developers can customize this list through the `rag_cfg` parameter to include additional searchers from `qwen_agent/tools/search_tools/`.

### How does keyword generation improve retrieval accuracy?

When `rag_keygen_strategy` is set to strategies like `SplitQueryThenGenKeyword`, the system dynamically loads a specialized agent from `qwen_agent/agents/keygen_strategies` (lines 106-122 of [`qwen_agent/memory/memory.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/memory/memory.py)). This agent decomposes complex queries into optimized keyword sets before retrieval, improving the relevance of retrieved passages compared to raw query matching.