Nemori Evaluation Benchmarks: LoCoMo and LongMemEval Explained

Nemori uses two primary evaluation benchmarks—LoCoMo (Long-Context Memory) and LongMemEval—to measure retrieval accuracy, long-context reasoning, and LLM alignment using automated LLM judges and lexical metrics.

The nemori-ai/nemori repository provides self-contained evaluation suites that stress-test the memory system against public datasets. These benchmarks verify how well Nemori retrieves and aligns long-context conversations across diverse question categories and 100k-token contexts.

LoCoMo Benchmark

The LoCoMo (Long-Context Memory) benchmark evaluates Nemori’s ability to retrieve and align long-context conversations across four question categories: Multi-Hop, Temporal-Reasoning, Open-Domain, and Single-Hop.

Pipeline Components

The LoCoMo evaluation follows a four-stage pipeline implemented in evaluation/locomo/:

  1. Data Ingestion – evaluation/locomo/add.py loads the LoCoMo JSON dataset and streams conversations into Nemori through the NemoriMemory facade:
from nemori import NemoriMemory, MemoryConfig

cfg = MemoryConfig(...)
memory = NemoriMemory(config=cfg)
memory.add_messages(user_id="locomo_user", messages=conversation)
  1. Parallel Search – evaluation/locomo/search.py executes concurrent vector, BM25, and original-text searches against stored episodes, writing raw results to locomo/results.json.

  2. Scoring – evaluation/locomo/evals.py computes BLEU-1 and F1 lexical scores, then invokes the LLM judge via metrics/llm_judge.evaluate_llm_judge to generate an LLM alignment score.

  3. Aggregation – evaluation/locomo/generate_scores.py aggregates per-category metrics and generates the final score table displayed in the README.

Metrics and Scoring

LoCoMo employs three complementary metrics to judge response quality:

  • BLEU-1: Lexical overlap between generated and reference answers
  • F1: Token-level precision and recall for semantic similarity
  • LLM-alignment score: Binary judgment from an LLM judge (configured to use OpenAI’s gpt-4.1-mini or gpt-4o-mini) determining if the response matches the gold answer

LongMemEval Benchmark

LongMemEval tests Nemori on a 100k-token context benchmark that stresses temporal reasoning, knowledge updates, and personalized preference recall.

Async LLM Judge Implementation

The evaluator in evaluation/longmemeval/evals.py implements an asynchronous evaluation pipeline:

  1. Loads a results file (e.g., longmem_results.json) produced by any Nemori run
  2. Selects prompt templates based on question type (temporal-reasoning, knowledge-update, single-session-preference, or default)
  3. Sends prompts to the configured LLM (self.model) via the async OpenAI client
  4. Parses JSON responses into a Grade model recording binary correctness
  5. Calculates statistics including overall accuracy, per-type accuracy, average response time, and error-case analysis

Evaluation Metrics

LongMemEval reports four key performance indicators:

  • Overall accuracy: Percentage of correctly answered questions
  • Accuracy per question type: Breakdown across temporal reasoning, knowledge updates, and preference recall
  • Average response time: Latency measurements for the evaluation loop
  • Error-case analysis: Structured list of failed retrievals for debugging

Running the Benchmarks

Execute the LoCoMo pipeline from the repository root to reproduce the published metrics:

cd evaluation
python locomo/add.py              # Ingest LoCoMo dataset into Nemori

python locomo/search.py           # Perform parallel searches

python locomo/evals.py            # Compute BLEU/F1 and LLM scores

python locomo/generate_scores.py  # Produce the final comparison table

Run the LongMemEval benchmark against existing results:

python evaluation/longmemeval/evals.py longmem_results.json \
  --model gpt-4.1-mini \
  --output report.json

Both scripts are self-contained and require only the OpenAI API for LLM judging—no external databases or services beyond the Nemori memory instance itself.

Summary

  • LoCoMo evaluates long-context retrieval across four question categories using evaluation/locomo/add.py, search.py, evals.py, and generate_scores.py
  • LongMemEval assesses 100k-token context handling via evaluation/longmemeval/evals.py with async LLM judges
  • Both benchmarks use LLM-based judges (gpt-4.1-mini or gpt-4o-mini) for binary correctness determination
  • BLEU-1 and F1 provide lexical and semantic scoring alongside LLM alignment metrics
  • All evaluation scripts operate against the standard NemoriMemory API without requiring external service dependencies

Frequently Asked Questions

What LLM models does Nemori use for benchmark judging?

Nemori configures OpenAI’s gpt-4.1-mini or gpt-4o-mini as the default judge for both LoCoMo and LongMemEval. The evaluate_llm_judge function in the metrics module sends prompts to these models via the async OpenAI client and parses the JSON responses into binary correctness flags.

How does the LoCoMo pipeline handle data ingestion?

The evaluation/locomo/add.py script loads the LoCoMo JSON dataset and streams each conversation into Nemori using the NemoriMemory facade. It initializes a MemoryConfig instance matching the quick-start configuration, then calls add_messages() with a consistent user_id to populate the memory store for subsequent retrieval testing.

What question categories does LongMemEval test?

LongMemEval evaluates four specific question types: temporal-reasoning (chronological inference across sessions), knowledge-update (handling revised facts), single-session-preference (recalling user preferences within one conversation), and a default category for general retrieval. The evaluator selects prompt templates dynamically based on these categories in evaluation/longmemeval/evals.py.

Can I run these benchmarks without external services?

Yes. Both benchmark suites are self-contained and only require the Nemori memory instance and OpenAI API access for the LLM judge. The scripts read local data files (LoCoMo JSON or LongMemEval results) and invoke the standard Nemori API; no additional databases, vector stores, or third-party evaluation platforms are necessary beyond the configured OpenAI client.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →