# How memU's Three-Layer Memory Architecture (Category/Item/Resource) Works Internally

> Explore memU's three-layer memory architecture Category Item Resource internally. Discover how it separates raw assets semantic units and organizational clusters for efficient LLM vector search and recall.

- Repository: [NevaMind AI/memU](https://github.com/nevamind-ai/memu)
- Tags: internals
- Published: 2026-02-19

---

**memU implements a hierarchical three-layer memory architecture that separates raw multimodal assets (Resources) from semantic memory units (MemoryItems) and organizational clusters (MemoryCategories), enabling efficient vector search and contextual recall for LLM applications.**

The NevaMind-AI/memU repository provides an open-source memory system for AI agents that organizes knowledge using a sophisticated three-layer memory architecture. This design separates ingestion artifacts from semantic content and categorical organization, allowing the system to handle multimodal inputs while maintaining fast retrieval through vector search and in-memory caching.

## The Three Memory Layers: Resource, MemoryItem, and MemoryCategory

memU stores knowledge in three hierarchical layers that map directly onto concrete data-model classes and repository implementations in the `src/memu/database/` directory.

### Resource Layer: Raw Multimodal Assets

The **Resource** layer represents the raw artifact that is ingested—whether a URL, local file path, or multimodal content.

- **Core model**: `Resource` class defined in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) (lines 68-74)
- **Repository**: `ResourceRepo` SQLite implementation in [`src/memu/database/sqlite/repositories/resource_repo.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/sqlite/repositories/resource_repo.py)
- **Stored fields**: URL, local file path, modality (video, text, audio), optional caption, and an embedding of the artifact itself

### MemoryItem Layer: Semantic Memory Units

The **MemoryItem** layer contains single extracted memory units—facts, skills, behaviors, or tool calls extracted from Resources.

- **Core model**: `MemoryItem` class defined in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) (lines 76-94)
- **Repository**: Protocol defined in [`src/memu/database/repositories/memory_item.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/repositories/memory_item.py); SQLite concrete class in [`src/memu/database/sqlite/repositories/memory_item_repo.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/sqlite/repositories/memory_item_repo.py)
- **Stored fields**: 
  - `resource_id` (foreign key to Resource)
  - `memory_type` enum: `profile`, `event`, `knowledge`, `behavior`, `skill`, or `tool`
  - `summary` text and its embedding vector
  - `extra` dict for reinforcement counters, reference IDs, and extensible metadata

### MemoryCategory Layer: Organizational Clusters

The **MemoryCategory** layer provides higher-level buckets that group related items—such as "travel" or "coding-tips."

- **Core model**: `MemoryCategory` class defined in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) (lines 96-100)
- **Repository**: `MemoryCategoryRepo` in [`src/memu/database/sqlite/repositories/memory_category_repo.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/sqlite/repositories/memory_category_repo.py)
- **Stored fields**: Name, description, embedding of the category description, running summary of contained items, and timestamps

### CategoryItem: The Link Table

The **CategoryItem** table provides the many-to-many relationship between `MemoryItem` and `MemoryCategory`.

- **Core model**: `CategoryItem` defined in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) (lines 103-106)
- **Repository**: `CategoryItemRepo` in [`src/memu/database/sqlite/repositories/category_item_repo.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/sqlite/repositories/category_item_repo.py)
- **Function**: Records `item_id`, `category_id`, and scope fields to link memory units with their organizational clusters

## The Memorization Workflow: From Ingestion to Persistence

The memorization process in [`src/memu/app/memorize.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/memorize.py) orchestrates the three layers through a five-step pipeline:

1. **Ingest and preprocess** – `MemorizeMixin._memorize_ingest_resource` fetches the remote file via `fs.fetch` and runs modality-specific preprocessing (video frame extraction, transcription, etc.), producing raw text and a caption.

2. **Extract items** – `MemorizeMixin._memorize_extract_items` sends preprocessed text to LLM prompts that emit structured tuples containing `memory_type`, `summary`, and `categories` via `_generate_structured_entries`.

3. **Create the Resource** – `_create_resource_with_caption` calls `store.resource_repo.create_resource`, which inserts a row into the `resources` table and caches the `Resource` object in `DatabaseState.resources` (defined in [`src/memu/database/state.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/state.py)).

4. **Persist MemoryItems** – `_persist_memory_items` executes:
   - Embeds each summary via the embedding client
   - Invokes `store.memory_item_repo.create_item` to insert into `memory_items` table and cache in `DatabaseState.items`
   - Maps each category name to a `MemoryCategory` (creating if absent) via `store.memory_category_repo.get_or_create_category`
   - Links items to categories using `store.category_item_repo.link_item_category`

5. **Update category summaries** – `_update_category_summaries` builds a prompt including new item summaries (or `[ref:xxx]` shortcuts) and writes back a refreshed `summary` field to the `MemoryCategory` row.

All repositories share a **single in-memory cache** (`DatabaseState`) that holds dictionaries for resources, items, categories, and a list for relations. This cache is populated on first access (`load_existing`) and kept in sync with the database on every create/update/delete operation, enabling fast read-paths for retrieval.

## Retrieval Architecture and Vector Search

When a user query arrives (e.g., via the OpenAI wrapper), the retrieval flow in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py) executes:

1. **Vector search** – `MemoryService.retrieve` calls `self._get_memory_items`, which performs vector search on `MemoryItem` embeddings via `vector_search_items` to obtain the top-k most relevant items.

2. **Category resolution** – The system resolves associated categories via `store.memory_category_repo.categories` to optionally enrich responses with category summaries.

3. **Prompt injection** – The OpenAI wrapper ([`src/memu/client/openai_wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/client/openai_wrapper.py)) transparently injects recalled memories into the system prompt via `MemuChatCompletions._inject_memories`, allowing the LLM to answer with the user's personal context.

## Key Design Patterns in memU's Architecture

**Separation of concerns** – Resources are immutable raw assets, Items are the semantic extracts, and Categories are semantic clusters. This distinction allows the system to handle multimodal ingestion while maintaining clean semantic retrieval.

**Scope-aware caching** – All repository constructors receive `scope_fields`; the base class (`SQLiteRepoBase`) extracts those fields from rows and stores them in the model's `extra` dict. This enables multi-tenant isolation via `user_id`, `agent_id`, or `session_id` without schema changes.

**Extensible "extra" fields** – Both `MemoryItem` and `Resource` expose `extra: dict[str, Any]` to store reinforcement counters, reference IDs, or tool-call metadata without requiring database migrations.

**Reference handling** – When `enable_item_references` is true, `_persist_item_references` parses `[ref:xxx]` tags from updated category summaries, maps them back to the underlying `MemoryItem.id`, and stores a short `ref_id` in the item's `extra` field. This enables later lookup via `list_items_by_ref_ids`.

## Code Examples

### Example 1: Memorize a YouTube video (multimodal)

```python
from memu.app.service import MemoryService
from memu.app.settings import MemUConfig

service = MemoryService(MemUConfig())          # initialise with default SQLite DB

response = await service.memorize(
    resource_url="https://youtu.be/dQw4w9WgXcQ",
    modality="video",
    user={"user_id": "alice"},
)
print(response["items"])   # list of MemoryItem dicts with summaries & categories

```

### Example 2: Retrieve memories about a user-specific topic

```python
from memu.client import wrap_openai
from openai import OpenAI
from memu.app.service import MemoryService

service = MemoryService(MemUConfig())
openai_client = OpenAI()
wrapped = wrap_openai(openai_client, service, user_id="alice")

# The wrapper injects the top-5 relevant memories automatically

chat = wrapped.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "user", "content": "What did I learn about Python last week?"}
    ],
)
print(chat.choices[0].message.content)

```

### Example 3: Direct low-level DB access (read-only)

```python
from memu.database.factory import create_sqlite_database

db = create_sqlite_database("sqlite:///memu.db")

# List all categories

for cat_id, cat in db.memory_category_repo.list_categories().items():
    print(cat.name, "=>", cat.summary)

```

## Summary

- **memU's three-layer memory architecture** separates raw multimodal assets (Resources) from semantic extracts (MemoryItems) and organizational clusters (MemoryCategories), enabling clean data flow from ingestion to retrieval.
- **Concrete implementations** reside in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) for data structures and `src/memu/database/sqlite/repositories/` for persistence logic, with a shared `DatabaseState` cache for performance.
- **The memorization pipeline** in [`src/memu/app/memorize.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/memorize.py) orchestrates five distinct phases: ingestion, extraction, resource creation, item persistence with category linking, and category summary updates.
- **Retrieval leverages vector search** on `MemoryItem` embeddings via [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py), with optional OpenAI wrapper integration in [`src/memu/client/openai_wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/client/openai_wrapper.py) for automatic memory injection.
- **Scope-aware design** supports multi-tenancy through `scope_fields` (user_id, agent_id, session_id) stored in extensible `extra` dictionaries without schema migrations.

## Frequently Asked Questions

### How does memU handle different data types like video and audio within the Resource layer?

The Resource layer in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) stores a `modality` field that identifies whether the raw asset is video, text, audio, or other formats. During the memorization workflow, `MemorizeMixin._memorize_ingest_resource` in [`src/memu/app/memorize.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/memorize.py) runs modality-specific preprocessing—such as video frame extraction or audio transcription—before the content reaches the MemoryItem extraction phase.

### What is the relationship between MemoryItems and MemoryCategories in memU's architecture?

MemoryItems and MemoryCategories share a many-to-many relationship mediated by the `CategoryItem` link table defined in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py) (lines 103-106). When `_persist_memory_items` executes during memorization, it maps each extracted item to its named categories via `store.memory_category_repo.get_or_create_category`, then persists the associations through `store.category_item_repo.link_item_category`. This allows a single memory unit to belong to multiple organizational clusters simultaneously.

### How does memU ensure fast retrieval performance when searching through large memory stores?

memU employs a dual strategy of vector indexing and in-memory caching. The `DatabaseState` class in [`src/memu/database/state.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/state.py) maintains dictionaries for resources, items, and categories that are populated on first access and synchronized on every write operation. For semantic search, `MemoryService.retrieve` in [`src/memu/app/retrieve.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/retrieve.py) performs vector similarity searches on `MemoryItem` embeddings via `vector_search_items`, avoiding full table scans even as the dataset scales.

### Can memU support multiple users or agents without data leakage between contexts?

Yes, memU implements scope-aware multi-tenancy through the `scope_fields` mechanism. All repository constructors in `src/memu/database/sqlite/repositories/` receive scope parameters (such as `user_id`, `agent_id`, or `session_id`), which the base `SQLiteRepoBase` class extracts from database rows and stores in each model's `extra` dictionary. This ensures that queries in `ResourceRepo`, `MemoryItemRepo`, and `MemoryCategoryRepo` automatically filter by scope, preventing cross-tenant data access without requiring separate database instances.