# How memU Reduces Token Costs While Maintaining Proactive Memory Capabilities

> Discover how memU slashes token costs for LLMs. Learn its strategies for granular tracking caching summarization and salient memory filtering to cut expenses without losing proactive capabilities.

- Repository: [NevaMind AI/memU](https://github.com/nevamind-ai/memu)
- Tags: performance
- Published: 2026-02-19

---

**memU reduces token costs by instrumenting every LLM call with granular usage tracking, caching expensive tool outputs, summarizing long-form content before extraction, and filtering memories based on salience scores—all while preserving proactive surfacing of relevant facts through intent-aware extraction rules.**

The open-source memU framework from NevaMind-AI implements a cost-aware memory architecture that helps large language model (LLM) applications reduce token costs while maintaining proactive memory capabilities. Unlike reactive systems that only retrieve context after explicit queries, memU anticipates user needs without exploding API budgets. By embedding token accounting directly into the memory lifecycle, the system creates a feedback loop where every operation carries a measurable cost that informs subsequent decisions.

## Tool-Call Caching and Token Accounting

MemU treats every external tool invocation as a billable event that must be tracked and cached. In [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py), the `ToolCallResult` model declares a `token_cost: int` field with a default of `-1`, establishing the schema for recording exact token expenditure per invocation【[models.py – token_cost field](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py#L51)】.

When a tool executes, the `add_tool_call` helper in [`src/memu/utils/tool.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/utils/tool.py) validates the memory type, generates a deterministic `call_hash` for deduplication, and persists the result inside the associated `MemoryItem.extra["tool_calls"]` array【[tool.py – add_tool_call](https://github.com/NevaMind-AI/memU/blob/main/src/memu/utils/tool.py#L36-L53)】. This caching prevents redundant API calls for identical inputs.

The system exposes budget intelligence through `get_tool_statistics`, which aggregates the last *N* calls to compute an **average token cost** (`avg_token_cost`). This metric flows back into the workflow layer, enabling runtime decisions to skip expensive operations when accumulated costs approach limits【[tool.py – avg_token_cost calculation](https://github.com/NevaMind-AI/memU/blob/main/src/memu/utils/tool.py#L90-L92)】.

## Granular LLM Usage Extraction

Every LLM provider response passes through a thin wrapper that extracts detailed token metadata. The `_extract_usage_from_raw_response` method in [`src/memu/llm/wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/wrapper.py) parses `input_tokens`, `output_tokens`, `total_tokens`, `cached_input_tokens`, and `reasoning_tokens` from the raw API response【[wrapper.py – usage extraction](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/wrapper.py#L420-L426)】.

This data populates the `LLMUsage` object, which the workflow logs and the tool-statistics routine consumes. Because costs are captured **per call**, memU can enforce hard budgets—raising exceptions or falling back to cheaper models when cumulative usage exceeds thresholds.

## Selective Summarization and Proactive Guards

To prevent long resources from inflating prompt sizes, memU segments incoming content and runs a **single-sentence summarizer** on each chunk. The `_summarize_segment` function in [`src/memu/app/memorize.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/memorize.py) processes segments with minimal token overhead, dramatically shrinking downstream context windows【[memorize.py – segment summarisation](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/memorize.py#L26-L33)】.

After summarization, extraction prompts include a **proactive guard** that prevents unnecessary memory generation. In [`src/memu/prompts/memory_type/profile.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/memory_type/profile.py), the prompt explicitly instructs the model: “Only extract content that the user shows strong proactive intent for; otherwise ignore follow-up questions”【[profile prompt – proactive guard](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/memory_type/profile.py#L88)】. This constraint stops the system from generating speculative facts that would require additional LLM calls to verify or store, keeping token usage low while still surfacing memories when users hint at future needs.

## Salience-Aware Retrieval Pruning

When querying the memory store, memU calculates a `salience_score` that combines similarity, reinforcement history, and recency decay. The implementation in [`src/memu/database/inmemory/vector.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/inmemory/vector.py) applies the formula: `similarity × log(reinforcement) × recency_decay`【[vector.py – salience_score](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/inmemory/vector.py#L22-L31)】.

By returning only the top-*k* highest-scoring candidates, the system ensures that subsequent LLM calls receive a **small, high-quality context set** rather than an exhaustive dump of loosely related memories. This early pruning significantly reduces the token count of generation prompts.

## Implementation Examples

The following patterns demonstrate how to interact with memU's cost-aware infrastructure:

```python

# Retrieve tool-type memory and inspect token statistics

from memu.client.openai_wrapper import MemuClient

client = MemuClient(api_key="…")
mem = client.get_memory(memory_type="tool", query="file_reader usage")
stats = mem.stats  # contains avg_token_cost, success_rate, etc.

print(f"Avg token cost per file_reader call: {stats['avg_token_cost']}")

```

```python

# Use the proactive memory extractor with automatic summarization

from memu.app.memorize import MemorizeMixin

class MyAgent(MemorizeMixin):
    ...

agent = MyAgent(...)

# The extractor surfaces "travel plans" only if the user shows proactive intent

response = await agent.memorize(
    resource_url="https://example.com/chatlog.txt",
    modality="conversation",
    user={"name": "Alice"}
)
print(response["memories"])

```

```python

# Enforce a hard token budget on LLM calls

from memu.llm.wrapper import LLMWrapper

wrapper = LLMWrapper(client=some_llm_client, token_budget=500)

# Raises BudgetExceededError if cumulative usage exceeds 500 tokens

await wrapper.chat("Summarise the last 10 messages", system_prompt="You are concise.")

```

## Summary

- **Tool-call caching** stores expensive API results with deterministic hashes and tracks per-call `token_cost` in `ToolCallResult` objects.
- **Granular usage extraction** captures `input_tokens`, `output_tokens`, and cached counts from every LLM response via `LLMWrapper`.
- **Selective summarization** reduces long inputs via `_summarize_segment` before extraction, while proactive guards in profile prompts prevent unnecessary memory generation.
- **Salience scoring** filters low-value memories early using `similarity × log(reinforcement) × recency_decay`, limiting context window size.

## Frequently Asked Questions

### How does memU track token costs for individual tool calls?

MemU stores token costs in the `token_cost` field of the `ToolCallResult` model defined in [`src/memu/database/models.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/models.py). The `add_tool_call` utility in [`src/memu/utils/tool.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/utils/tool.py) writes these values into the `extra["tool_calls"]` array of each `MemoryItem`, while `get_tool_statistics` aggregates historical data to compute average costs for budgeting decisions.

### What prevents memU from generating unnecessary proactive memories?

The extraction prompts in [`src/memu/prompts/memory_type/profile.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/prompts/memory_type/profile.py) contain explicit proactive guards that instruct the model to extract content only when the user demonstrates strong proactive intent. This prevents the system from expanding the memory graph with speculative facts that would trigger additional expensive LLM calls.

### How does salience scoring reduce token consumption?

The `salience_score` calculation in [`src/memu/database/inmemory/vector.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/database/inmemory/vector.py) ranks candidates using `similarity × log(reinforcement) × recency_decay`. By retrieving only the top-*k* highest-scoring memories, memU ensures that downstream LLM prompts contain only the most relevant context, avoiding the token bloat associated with large, unfiltered retrieval sets.

### Can memU enforce hard token budgets during execution?

Yes. The `LLMWrapper` class in [`src/memu/llm/wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/wrapper.py) accepts a `token_budget` parameter. As the wrapper extracts usage data from each response via `_extract_usage_from_raw_response`, it maintains cumulative counters and raises exceptions when the budget is exceeded, allowing workflows to skip expensive reasoning steps or switch to cheaper models.