How to Create Custom Memory Categories with Specific Extraction Prompts in MemU

You create custom memory categories by extending CategoryConfig in MemorizeConfig and define custom extraction prompts by overriding memory_type_prompts, which the MemorizeMixin merges with built-in templates during the memorization workflow.

MemU is an open-source memory management system that organizes extracted information into structured categories. When the default memory types—such as preferences, goals, or knowledge—don't fit your domain, you can define custom categories with specialized extraction prompts that guide the LLM to capture specific information. This configuration-driven approach requires no changes to the core library, only extensions to the settings classes defined in src/memu/app/settings.py.

Understanding Memory Categories and Extraction Prompts

MemU stores extracted information in memory categories (e.g., "preferences", "goals") defined at service startup in the memory-category table. During memorization, the system processes resources through distinct memory types: profile, event, knowledge, behavior, skill, and tool. For each type, MemU builds a prompt that instructs the LLM how to extract and structure information.

The MemorizeMixin class orchestrates this workflow in src/memu/app/memorize.py. When memorize() is called, it first ensures categories are initialized via _initialize_categories, then generates structured entries using _build_memory_type_prompt to construct the LLM instructions. Customization happens entirely through configuration objects, not by subclassing the mixin.

Step 1: Define a Custom Memory Category

To add a new category, extend the memory_categories list in your MemorizeConfig subclass. Each category requires a CategoryConfig object specifying the name, description, and optional parameters like target_length for auto-summarization.


# src/memu/app/settings.py

from memu.app.settings import MemorizeConfig, CategoryConfig

class MyConfig(MemorizeConfig):
    memory_categories: list[CategoryConfig] = MemorizeConfig.memory_categories + [
        CategoryConfig(
            name="travel_plans",
            description="User's upcoming trips and itineraries",
            target_length=300,  # Max tokens for auto-summary

        ),
    ]

When the service starts, MemorizeMixin._initialize_categories executes the following sequence:

  1. Builds embedding text using _category_embedding_text (format: "{name}: {description}")
  2. Calls the embedding LLM via self._get_llm_client("embedding") to generate a vector
  3. Persists the record through MemoryCategoryRepo.get_or_create_category in src/memu/database/repositories/memory_category_repo.py

The category is now available for assignment during memory extraction.

Step 2: Override Extraction Prompts for Specific Memory Types

To ensure the LLM extracts information relevant to your new category, override the prompt for the appropriate memory type using the memory_type_prompts dictionary in MemorizeConfig. You can provide either a plain string or a structured CustomPrompt object.

Using CustomPrompt for Complex Instructions

For granular control over prompt sections, use CustomPrompt with PromptBlock objects. The ordinal field controls the ordering of blocks when merged with default templates.


# src/memu/app/settings.py

from memu.app.settings import MemorizeConfig, CustomPrompt, PromptBlock

class MyConfig(MemorizeConfig):
    memory_type_prompts: dict[str, str | CustomPrompt] = {
        "behavior": CustomPrompt(
            root={
                "objective": PromptBlock(
                    ordinal=10,
                    prompt=(
                        "You are a travel-behavior extractor. "
                        "From the conversation extract ONLY recurring travel-related habits, "
                        "e.g. packing routines, preferred airlines, or typical trip planning steps."
                    ),
                ),
                "output": PromptBlock(
                    ordinal=90,
                    prompt=(
                        "# Response Format (JSON):\n"

                        "{{\n"
                        '    "memories_items": [\n'
                        "        {{\n"
                        '            "content": "extracted sentence",\n'
                        '            "categories": ["travel_plans"]\n'
                        "        }}\n"
                        "    ]\n"
                        "}}"
                    ),
                ),
            }
        ),
    }

The key "behavior" must match the literal value in the MemoryType enum defined in src/memu/database/models.py. The MemorizeMixin._resolve_custom_prompt method merges your override with the default templates from src/memu/prompts/memory_type/behavior.py, and _build_memory_type_prompt (lines 665-680 of src/memu/app/memorize.py) formats the final prompt with resource text and category lists.

Using Simple String Overrides

For straightforward prompt replacements, provide a plain string:

memory_type_prompts = {
    "behavior": (
        "Extract ONLY recurring travel habits and write each as a short sentence. "
        "Assign them to the category 'travel_plans'."
    )
}

Step 3: Verify the Integration

After restarting the service with your custom configuration, invoke the memorization workflow to verify the category and prompt are active:

await memu.memorize(
    resource_url="https://example.com/chat.txt",
    modality="conversation",
    user={"user_id": "alice"},
)

During execution:

  • _generate_structured_entries calls _build_memory_type_prompt, which pulls your custom behavior prompt from MemorizeConfig.memory_type_prompts
  • The LLM receives instructions referencing the "travel_plans" category
  • Extracted items with "travel_plans" in their categories list are linked to the MemoryCategory row via store.memory_category_repo.link_item_category

Complete Working Example

This standalone configuration demonstrates both custom category definition and prompt override:


# examples/custom_category.py

from memu.app.service import MemU
from memu.app.settings import MemorizeConfig, CategoryConfig, CustomPrompt, PromptBlock

class MyConfig(MemorizeConfig):
    # 1. Define the custom category

    memory_categories = MemorizeConfig.memory_categories + [
        CategoryConfig(
            name="travel_plans",
            description="User's trips, itineraries and travel habits",
            target_length=250,
        )
    ]

    # 2. Override the behavior extraction prompt

    memory_type_prompts = {
        "behavior": CustomPrompt(
            root={
                "objective": PromptBlock(
                    ordinal=10,
                    prompt=(
                        "You are a travel-behavior extractor. "
                        "From the conversation extract ONLY recurring travel-related patterns "
                        "(packing routine, preferred airlines, etc.)."
                    ),
                ),
                "output": PromptBlock(
                    ordinal=90,
                    prompt=(
                        "# Response Format (JSON):\n"

                        "{{\n"
                        '    "memories_items": [\n'
                        "        {{\n"
                        '            "content": "extracted sentence",\n'
                        '            "categories": ["travel_plans"]\n'
                        "        }}\n"
                        "    ]\n"
                        "}}"
                    ),
                ),
            }
        ),
    }

# Initialize with custom config

memu = MemU(settings=MyConfig())

# Execute memorization

await memu.memorize(
    resource_url="https://example.com/conversation.txt",
    modality="conversation",
    user={"user_id": "alice"},
)

Running this creates a MemoryCategory record named travel_plans and applies the custom extraction logic to the behavior memory type.

Summary

  • Configuration Location: Define categories in MemorizeConfig.memory_categories and prompts in MemorizeConfig.memory_type_prompts within src/memu/app/settings.py.
  • Persistence: Categories are embedded and stored via MemoryCategoryRepo.get_or_create_category during initialization.
  • Prompt Resolution: The system merges custom prompts with defaults in _build_memory_type_prompt, supporting both CustomPrompt objects and plain strings.
  • Memory Types: Valid keys for memory_type_prompts correspond to the MemoryType enum values: profile, event, knowledge, behavior, skill, and tool.
  • Zero Core Changes: All customization happens through configuration; no modifications to src/memu/app/memorize.py or other core files are required.

Frequently Asked Questions

What memory types support custom extraction prompts?

All six memory types defined in src/memu/database/models.py support custom prompts: profile, event, knowledge, behavior, skill, and tool. Use these exact strings as keys in the memory_type_prompts dictionary. The default templates for each type reside in src/memu/prompts/memory_type/<type>.py.

Can I use a simple string instead of the CustomPrompt class?

Yes. The memory_type_prompts dictionary accepts values of type str | CustomPrompt. A plain string replaces the entire prompt template, while a CustomPrompt object merges specific blocks (like objective or output) with default templates using ordinal-based ordering defined in DEFAULT_MEMORY_CUSTOM_PROMPT_ORDINAL.

How does MemU generate embeddings for custom categories?

During initialization, MemorizeMixin._initialize_categories calls _category_embedding_text to format the string "{name}: {description}". This text is passed to the embedding LLM via self._get_llm_client("embedding"), and the resulting vector is stored alongside the category metadata in the database through MemoryCategoryRepo.get_or_create_category.

Do I need to restart the service after adding new categories?

Yes. Categories are initialized once at service startup when MemorizeMixin._ensure_categories_ready triggers _initialize_categories. New CategoryConfig entries are not picked up dynamically; you must restart the service to persist new categories to the database and make them available for memory extraction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →