How Nemori's Predict-Calibrate Learning Mechanism Works

Nemori's predict-calibrate learning mechanism implements a continuous prediction-correction loop that generates anticipated conversation content using historic semantic memories, then extracts novel knowledge from the discrepancy between predicted and actual dialogue to update its vector store.

The nemori-ai/nemori repository implements an innovative predict-calibrate learning approach to continuous knowledge acquisition. Unlike simple extraction-based memory systems, Nemori uses a two-phase LLM-driven process that treats prediction errors as learning opportunities, enabling the system to refine its understanding of user preferences and facts through natural conversation.

The Two-Phase Predict-Calibrate Learning Loop

Phase 1: Prediction (Generating Anticipated Content)

When a new conversation episode arrives, Nemori first retrieves the most relevant historic statements from its ChromaDB vector store using max_statements_for_prediction and statement_similarity_threshold parameters. In src/generation/prediction_correction_engine.py, the _predict_episode method feeds these retrieved statements along with the episode title to an LLM prompt template.

The LLM generates predicted episode content—a simulation of what the conversation would look like based solely on existing knowledge. This prediction step serves as a baseline, representing the system's current understanding of the user's context and preferences.

Phase 2: Calibration (Extracting Novel Knowledge)

After recording the actual episode messages, Nemori executes the calibration phase through _extract_knowledge_from_comparison in prediction_correction_engine.py. This method compares the original messages against the predicted content using a second LLM prompt designed to identify high-value knowledge statements present in reality but missing from the prediction.

The extracted novel statements are converted into SemanticMemory objects via _convert_to_semantic_memories and persisted back to the vector index. This prediction-correction mechanism transforms prediction errors into structured learning, ensuring the system only stores genuinely new information rather than redundant data.

Core Implementation in prediction_correction_engine.py

The PredictionCorrectionEngine class orchestrates the learning loop through its primary entry point learn_from_episode_simplified. This method coordinates four key operations:

  • _retrieve_relevant_statements – Performs ChromaDB vector search to fetch historic context based on embedding similarity.
  • _predict_episode – Generates anticipated dialogue using the LLM client with configurable prediction_temperature.
  • _extract_knowledge_from_comparison – Identifies discrepancies between predicted and actual content to find new facts.
  • _convert_to_semantic_memories – Transforms extracted strings into persistent SemanticMemory models for vector storage.

The implementation maintains statelessness beyond the vector store, allowing each incoming episode to immediately improve the system's user model without requiring batch processing or external state management.

Code Example: Running the Predict-Calibrate Loop

The following example demonstrates how to initialize the engine and process a new conversation episode:

from nemori.src.generation.prediction_correction_engine import PredictionCorrectionEngine
from nemori.src.utils.llm_client import LLMClient
from nemori.src.utils.embedding_client import EmbeddingClient
from nemori.src.config import MemoryConfig

# Initialize dependencies

llm = LLMClient(...)
embed = EmbeddingClient(...)
config = MemoryConfig(...)
engine = PredictionCorrectionEngine(llm, embed, config, vector_search=my_chroma_instance)

# Process new episode

new_memories = engine.learn_from_episode_simplified(
    user_id="user_123",
    new_episode=new_ep,  # Episode object

    existing_statements=existing_kb,  # List of SemanticMemory

)

To inspect the intermediate prediction step:

predicted = engine._predict_episode(
    episode_title=new_ep.title,
    relevant_statements=["User likes Python", "Works at Acme Corp"]
)
print(predicted)

And to see the extracted knowledge after calibration:

extracted = engine._extract_knowledge_from_comparison(
    original_messages=new_ep.original_messages,
    predicted_episode=predicted,
)
print(extracted)  # List of high-value knowledge statements

Configuration Parameters for Fine-Tuning

The predict-calibrate mechanism exposes several configuration options in MemoryConfig that control retrieval and generation behavior:

  • max_statements_for_prediction – Limits the number of historic statements retrieved from ChromaDB to prevent context window overflow.
  • statement_similarity_threshold – Sets the minimum embedding similarity score for including historic statements in the prediction context.
  • prediction_temperature – Controls the creativity of the LLM when generating anticipated episode content; lower values produce more deterministic predictions.

These parameters allow developers to balance between prediction accuracy and computational cost, or to adjust how aggressively the system filters historical context before generating predictions.

Summary

  • Nemori's predict-calibrate learning mechanism uses a two-phase LLM-driven process to acquire knowledge from conversations.
  • The prediction phase generates anticipated dialogue based on retrieved historic statements from ChromaDB.
  • The calibration phase extracts novel knowledge by comparing actual messages against predictions, storing only genuinely new information as SemanticMemory objects.
  • The implementation resides primarily in src/generation/prediction_correction_engine.py, orchestrated by PredictionCorrectionEngine.learn_from_episode_simplified.
  • Configuration parameters like max_statements_for_prediction and prediction_temperature allow fine-tuning of the retrieval and generation process.

Frequently Asked Questions

What is the difference between predict-calibrate learning and standard RAG?

Standard Retrieval-Augmented Generation (RAG) retrieves context to answer queries but does not systematically learn from the conversation. Nemori's predict-calibrate learning actively compares LLM-generated predictions against actual outcomes to identify and store new knowledge, creating a continuous learning loop that improves future predictions based on previous errors.

How does Nemori prevent duplicate knowledge from being stored?

During the calibration phase, the _extract_knowledge_from_comparison method specifically prompts the LLM to identify statements present in the actual conversation but missing from the prediction. By design, this extracts only discrepant or novel information, ensuring that redundant facts already captured in the vector store are not duplicated.

Can the predict-calibrate mechanism work with different LLM providers?

Yes. The PredictionCorrectionEngine accepts an LLMClient abstraction defined in src/utils/llm_client.py. As long as the client implements the required interface for text and JSON generation, the predict-calibrate learning mechanism remains agnostic to the underlying LLM provider, whether OpenAI, Anthropic, or local models.

What role does ChromaDB play in the predict-calibrate learning process?

ChromaDB serves as the vector search backend for the prediction phase. The engine queries ChromaDB via _retrieve_relevant_statements to fetch historic SemanticMemory objects whose embeddings are similar to the new episode title. These retrieved statements provide the context necessary for the LLM to generate accurate predictions, grounding the learning process in the user's existing knowledge base.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →