# How Nemori's Predict-Calibrate Learning Mechanism Works

> Understand Nemori's predict-calibrate learning mechanism. Discover how Nemori continuously predicts and corrects conversations using semantic memories to extract novel knowledge and update its vector store.

- Repository: [Nemori AI/nemori](https://github.com/nemori-ai/nemori)
- Tags: deep-dive
- Published: 2026-03-08

---

**Nemori's predict-calibrate learning mechanism implements a continuous prediction-correction loop that generates anticipated conversation content using historic semantic memories, then extracts novel knowledge from the discrepancy between predicted and actual dialogue to update its vector store.**

The `nemori-ai/nemori` repository implements an innovative **predict-calibrate learning** approach to continuous knowledge acquisition. Unlike simple extraction-based memory systems, Nemori uses a two-phase LLM-driven process that treats prediction errors as learning opportunities, enabling the system to refine its understanding of user preferences and facts through natural conversation.

## The Two-Phase Predict-Calibrate Learning Loop

### Phase 1: Prediction (Generating Anticipated Content)

When a new conversation episode arrives, Nemori first retrieves the most relevant historic statements from its ChromaDB vector store using `max_statements_for_prediction` and `statement_similarity_threshold` parameters. In [`src/generation/prediction_correction_engine.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prediction_correction_engine.py), the `_predict_episode` method feeds these retrieved statements along with the episode title to an LLM prompt template.

The LLM generates **predicted episode content**—a simulation of what the conversation would look like based solely on existing knowledge. This prediction step serves as a baseline, representing the system's current understanding of the user's context and preferences.

### Phase 2: Calibration (Extracting Novel Knowledge)

After recording the actual episode messages, Nemori executes the calibration phase through `_extract_knowledge_from_comparison` in [`prediction_correction_engine.py`](https://github.com/nemori-ai/nemori/blob/main/prediction_correction_engine.py). This method compares the **original messages** against the **predicted content** using a second LLM prompt designed to identify high-value knowledge statements present in reality but missing from the prediction.

The extracted novel statements are converted into `SemanticMemory` objects via `_convert_to_semantic_memories` and persisted back to the vector index. This **prediction-correction** mechanism transforms prediction errors into structured learning, ensuring the system only stores genuinely new information rather than redundant data.

## Core Implementation in prediction_correction_engine.py

The `PredictionCorrectionEngine` class orchestrates the learning loop through its primary entry point `learn_from_episode_simplified`. This method coordinates four key operations:

- **`_retrieve_relevant_statements`** – Performs ChromaDB vector search to fetch historic context based on embedding similarity.
- **`_predict_episode`** – Generates anticipated dialogue using the LLM client with configurable `prediction_temperature`.
- **`_extract_knowledge_from_comparison`** – Identifies discrepancies between predicted and actual content to find new facts.
- **`_convert_to_semantic_memories`** – Transforms extracted strings into persistent `SemanticMemory` models for vector storage.

The implementation maintains statelessness beyond the vector store, allowing each incoming episode to immediately improve the system's user model without requiring batch processing or external state management.

## Code Example: Running the Predict-Calibrate Loop

The following example demonstrates how to initialize the engine and process a new conversation episode:

```python
from nemori.src.generation.prediction_correction_engine import PredictionCorrectionEngine
from nemori.src.utils.llm_client import LLMClient
from nemori.src.utils.embedding_client import EmbeddingClient
from nemori.src.config import MemoryConfig

# Initialize dependencies

llm = LLMClient(...)
embed = EmbeddingClient(...)
config = MemoryConfig(...)
engine = PredictionCorrectionEngine(llm, embed, config, vector_search=my_chroma_instance)

# Process new episode

new_memories = engine.learn_from_episode_simplified(
    user_id="user_123",
    new_episode=new_ep,  # Episode object

    existing_statements=existing_kb,  # List of SemanticMemory

)

```

To inspect the intermediate prediction step:

```python
predicted = engine._predict_episode(
    episode_title=new_ep.title,
    relevant_statements=["User likes Python", "Works at Acme Corp"]
)
print(predicted)

```

And to see the extracted knowledge after calibration:

```python
extracted = engine._extract_knowledge_from_comparison(
    original_messages=new_ep.original_messages,
    predicted_episode=predicted,
)
print(extracted)  # List of high-value knowledge statements

```

## Configuration Parameters for Fine-Tuning

The predict-calibrate mechanism exposes several configuration options in `MemoryConfig` that control retrieval and generation behavior:

- **`max_statements_for_prediction`** – Limits the number of historic statements retrieved from ChromaDB to prevent context window overflow.
- **`statement_similarity_threshold`** – Sets the minimum embedding similarity score for including historic statements in the prediction context.
- **`prediction_temperature`** – Controls the creativity of the LLM when generating anticipated episode content; lower values produce more deterministic predictions.

These parameters allow developers to balance between prediction accuracy and computational cost, or to adjust how aggressively the system filters historical context before generating predictions.

## Summary

- Nemori's **predict-calibrate learning** mechanism uses a two-phase LLM-driven process to acquire knowledge from conversations.
- The **prediction phase** generates anticipated dialogue based on retrieved historic statements from ChromaDB.
- The **calibration phase** extracts novel knowledge by comparing actual messages against predictions, storing only genuinely new information as `SemanticMemory` objects.
- The implementation resides primarily in [`src/generation/prediction_correction_engine.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prediction_correction_engine.py), orchestrated by `PredictionCorrectionEngine.learn_from_episode_simplified`.
- Configuration parameters like `max_statements_for_prediction` and `prediction_temperature` allow fine-tuning of the retrieval and generation process.

## Frequently Asked Questions

### What is the difference between predict-calibrate learning and standard RAG?

Standard Retrieval-Augmented Generation (RAG) retrieves context to answer queries but does not systematically learn from the conversation. Nemori's **predict-calibrate learning** actively compares LLM-generated predictions against actual outcomes to identify and store new knowledge, creating a continuous learning loop that improves future predictions based on previous errors.

### How does Nemori prevent duplicate knowledge from being stored?

During the **calibration phase**, the `_extract_knowledge_from_comparison` method specifically prompts the LLM to identify statements present in the actual conversation but missing from the prediction. By design, this extracts only **discrepant** or novel information, ensuring that redundant facts already captured in the vector store are not duplicated.

### Can the predict-calibrate mechanism work with different LLM providers?

Yes. The `PredictionCorrectionEngine` accepts an `LLMClient` abstraction defined in [`src/utils/llm_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/llm_client.py). As long as the client implements the required interface for text and JSON generation, the predict-calibrate learning mechanism remains agnostic to the underlying LLM provider, whether OpenAI, Anthropic, or local models.

### What role does ChromaDB play in the predict-calibrate learning process?

ChromaDB serves as the **vector search backend** for the prediction phase. The engine queries ChromaDB via `_retrieve_relevant_statements` to fetch historic `SemanticMemory` objects whose embeddings are similar to the new episode title. These retrieved statements provide the context necessary for the LLM to generate accurate predictions, grounding the learning process in the user's existing knowledge base.