How memU Reduces Token Costs While Maintaining Proactive Memory Capabilities

memU reduces token costs by instrumenting every LLM call with granular usage tracking, caching expensive tool outputs, summarizing long-form content before extraction, and filtering memories based on salience scores—all while preserving proactive surfacing of relevant facts through intent-aware extraction rules.

The open-source memU framework from NevaMind-AI implements a cost-aware memory architecture that helps large language model (LLM) applications reduce token costs while maintaining proactive memory capabilities. Unlike reactive systems that only retrieve context after explicit queries, memU anticipates user needs without exploding API budgets. By embedding token accounting directly into the memory lifecycle, the system creates a feedback loop where every operation carries a measurable cost that informs subsequent decisions.

Tool-Call Caching and Token Accounting

MemU treats every external tool invocation as a billable event that must be tracked and cached. In src/memu/database/models.py, the ToolCallResult model declares a token_cost: int field with a default of -1, establishing the schema for recording exact token expenditure per invocation【models.py – token_cost field】.

When a tool executes, the add_tool_call helper in src/memu/utils/tool.py validates the memory type, generates a deterministic call_hash for deduplication, and persists the result inside the associated MemoryItem.extra["tool_calls"] array【tool.py – add_tool_call】. This caching prevents redundant API calls for identical inputs.

The system exposes budget intelligence through get_tool_statistics, which aggregates the last N calls to compute an average token cost (avg_token_cost). This metric flows back into the workflow layer, enabling runtime decisions to skip expensive operations when accumulated costs approach limits【tool.py – avg_token_cost calculation】.

Granular LLM Usage Extraction

Every LLM provider response passes through a thin wrapper that extracts detailed token metadata. The _extract_usage_from_raw_response method in src/memu/llm/wrapper.py parses input_tokens, output_tokens, total_tokens, cached_input_tokens, and reasoning_tokens from the raw API response【wrapper.py – usage extraction】.

This data populates the LLMUsage object, which the workflow logs and the tool-statistics routine consumes. Because costs are captured per call, memU can enforce hard budgets—raising exceptions or falling back to cheaper models when cumulative usage exceeds thresholds.

Selective Summarization and Proactive Guards

To prevent long resources from inflating prompt sizes, memU segments incoming content and runs a single-sentence summarizer on each chunk. The _summarize_segment function in src/memu/app/memorize.py processes segments with minimal token overhead, dramatically shrinking downstream context windows【memorize.py – segment summarisation】.

After summarization, extraction prompts include a proactive guard that prevents unnecessary memory generation. In src/memu/prompts/memory_type/profile.py, the prompt explicitly instructs the model: “Only extract content that the user shows strong proactive intent for; otherwise ignore follow-up questions”【profile prompt – proactive guard】. This constraint stops the system from generating speculative facts that would require additional LLM calls to verify or store, keeping token usage low while still surfacing memories when users hint at future needs.

Salience-Aware Retrieval Pruning

When querying the memory store, memU calculates a salience_score that combines similarity, reinforcement history, and recency decay. The implementation in src/memu/database/inmemory/vector.py applies the formula: similarity × log(reinforcement) × recency_decay【vector.py – salience_score】.

By returning only the top-k highest-scoring candidates, the system ensures that subsequent LLM calls receive a small, high-quality context set rather than an exhaustive dump of loosely related memories. This early pruning significantly reduces the token count of generation prompts.

Implementation Examples

The following patterns demonstrate how to interact with memU's cost-aware infrastructure:


# Retrieve tool-type memory and inspect token statistics

from memu.client.openai_wrapper import MemuClient

client = MemuClient(api_key="…")
mem = client.get_memory(memory_type="tool", query="file_reader usage")
stats = mem.stats  # contains avg_token_cost, success_rate, etc.

print(f"Avg token cost per file_reader call: {stats['avg_token_cost']}")

# Use the proactive memory extractor with automatic summarization

from memu.app.memorize import MemorizeMixin

class MyAgent(MemorizeMixin):
    ...

agent = MyAgent(...)

# The extractor surfaces "travel plans" only if the user shows proactive intent

response = await agent.memorize(
    resource_url="https://example.com/chatlog.txt",
    modality="conversation",
    user={"name": "Alice"}
)
print(response["memories"])

# Enforce a hard token budget on LLM calls

from memu.llm.wrapper import LLMWrapper

wrapper = LLMWrapper(client=some_llm_client, token_budget=500)

# Raises BudgetExceededError if cumulative usage exceeds 500 tokens

await wrapper.chat("Summarise the last 10 messages", system_prompt="You are concise.")

Summary

  • Tool-call caching stores expensive API results with deterministic hashes and tracks per-call token_cost in ToolCallResult objects.
  • Granular usage extraction captures input_tokens, output_tokens, and cached counts from every LLM response via LLMWrapper.
  • Selective summarization reduces long inputs via _summarize_segment before extraction, while proactive guards in profile prompts prevent unnecessary memory generation.
  • Salience scoring filters low-value memories early using similarity × log(reinforcement) × recency_decay, limiting context window size.

Frequently Asked Questions

How does memU track token costs for individual tool calls?

MemU stores token costs in the token_cost field of the ToolCallResult model defined in src/memu/database/models.py. The add_tool_call utility in src/memu/utils/tool.py writes these values into the extra["tool_calls"] array of each MemoryItem, while get_tool_statistics aggregates historical data to compute average costs for budgeting decisions.

What prevents memU from generating unnecessary proactive memories?

The extraction prompts in src/memu/prompts/memory_type/profile.py contain explicit proactive guards that instruct the model to extract content only when the user demonstrates strong proactive intent. This prevents the system from expanding the memory graph with speculative facts that would trigger additional expensive LLM calls.

How does salience scoring reduce token consumption?

The salience_score calculation in src/memu/database/inmemory/vector.py ranks candidates using similarity × log(reinforcement) × recency_decay. By retrieving only the top-k highest-scoring memories, memU ensures that downstream LLM prompts contain only the most relevant context, avoiding the token bloat associated with large, unfiltered retrieval sets.

Can memU enforce hard token budgets during execution?

Yes. The LLMWrapper class in src/memu/llm/wrapper.py accepts a token_budget parameter. As the wrapper extracts usage data from each response via _extract_usage_from_raw_response, it maintains cumulative counters and raises exceptions when the budget is exceeded, allowing workflows to skip expensive reasoning steps or switch to cheaper models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →