Nemori Evaluation Benchmarks: LoCoMo and LongMemEval Explained
Nemori uses two primary evaluation benchmarks—LoCoMo (Long-Context Memory) and LongMemEval—to measure retrieval accuracy, long-context reasoning, and LLM alignment using automated LLM judges and lexical metrics.
The nemori-ai/nemori repository provides self-contained evaluation suites that stress-test the memory system against public datasets. These benchmarks verify how well Nemori retrieves and aligns long-context conversations across diverse question categories and 100k-token contexts.
LoCoMo Benchmark
The LoCoMo (Long-Context Memory) benchmark evaluates Nemori’s ability to retrieve and align long-context conversations across four question categories: Multi-Hop, Temporal-Reasoning, Open-Domain, and Single-Hop.
Pipeline Components
The LoCoMo evaluation follows a four-stage pipeline implemented in evaluation/locomo/:
- Data Ingestion –
evaluation/locomo/add.pyloads the LoCoMo JSON dataset and streams conversations into Nemori through theNemoriMemoryfacade:
from nemori import NemoriMemory, MemoryConfig
cfg = MemoryConfig(...)
memory = NemoriMemory(config=cfg)
memory.add_messages(user_id="locomo_user", messages=conversation)
-
Parallel Search –
evaluation/locomo/search.pyexecutes concurrent vector, BM25, and original-text searches against stored episodes, writing raw results tolocomo/results.json. -
Scoring –
evaluation/locomo/evals.pycomputes BLEU-1 and F1 lexical scores, then invokes the LLM judge viametrics/llm_judge.evaluate_llm_judgeto generate an LLM alignment score. -
Aggregation –
evaluation/locomo/generate_scores.pyaggregates per-category metrics and generates the final score table displayed in the README.
Metrics and Scoring
LoCoMo employs three complementary metrics to judge response quality:
- BLEU-1: Lexical overlap between generated and reference answers
- F1: Token-level precision and recall for semantic similarity
- LLM-alignment score: Binary judgment from an LLM judge (configured to use OpenAI’s
gpt-4.1-miniorgpt-4o-mini) determining if the response matches the gold answer
LongMemEval Benchmark
LongMemEval tests Nemori on a 100k-token context benchmark that stresses temporal reasoning, knowledge updates, and personalized preference recall.
Async LLM Judge Implementation
The evaluator in evaluation/longmemeval/evals.py implements an asynchronous evaluation pipeline:
- Loads a results file (e.g.,
longmem_results.json) produced by any Nemori run - Selects prompt templates based on question type (
temporal-reasoning,knowledge-update,single-session-preference, or default) - Sends prompts to the configured LLM (
self.model) via the async OpenAI client - Parses JSON responses into a
Grademodel recording binary correctness - Calculates statistics including overall accuracy, per-type accuracy, average response time, and error-case analysis
Evaluation Metrics
LongMemEval reports four key performance indicators:
- Overall accuracy: Percentage of correctly answered questions
- Accuracy per question type: Breakdown across temporal reasoning, knowledge updates, and preference recall
- Average response time: Latency measurements for the evaluation loop
- Error-case analysis: Structured list of failed retrievals for debugging
Running the Benchmarks
Execute the LoCoMo pipeline from the repository root to reproduce the published metrics:
cd evaluation
python locomo/add.py # Ingest LoCoMo dataset into Nemori
python locomo/search.py # Perform parallel searches
python locomo/evals.py # Compute BLEU/F1 and LLM scores
python locomo/generate_scores.py # Produce the final comparison table
Run the LongMemEval benchmark against existing results:
python evaluation/longmemeval/evals.py longmem_results.json \
--model gpt-4.1-mini \
--output report.json
Both scripts are self-contained and require only the OpenAI API for LLM judging—no external databases or services beyond the Nemori memory instance itself.
Summary
- LoCoMo evaluates long-context retrieval across four question categories using
evaluation/locomo/add.py,search.py,evals.py, andgenerate_scores.py - LongMemEval assesses 100k-token context handling via
evaluation/longmemeval/evals.pywith async LLM judges - Both benchmarks use LLM-based judges (
gpt-4.1-miniorgpt-4o-mini) for binary correctness determination - BLEU-1 and F1 provide lexical and semantic scoring alongside LLM alignment metrics
- All evaluation scripts operate against the standard
NemoriMemoryAPI without requiring external service dependencies
Frequently Asked Questions
What LLM models does Nemori use for benchmark judging?
Nemori configures OpenAI’s gpt-4.1-mini or gpt-4o-mini as the default judge for both LoCoMo and LongMemEval. The evaluate_llm_judge function in the metrics module sends prompts to these models via the async OpenAI client and parses the JSON responses into binary correctness flags.
How does the LoCoMo pipeline handle data ingestion?
The evaluation/locomo/add.py script loads the LoCoMo JSON dataset and streams each conversation into Nemori using the NemoriMemory facade. It initializes a MemoryConfig instance matching the quick-start configuration, then calls add_messages() with a consistent user_id to populate the memory store for subsequent retrieval testing.
What question categories does LongMemEval test?
LongMemEval evaluates four specific question types: temporal-reasoning (chronological inference across sessions), knowledge-update (handling revised facts), single-session-preference (recalling user preferences within one conversation), and a default category for general retrieval. The evaluator selects prompt templates dynamically based on these categories in evaluation/longmemeval/evals.py.
Can I run these benchmarks without external services?
Yes. Both benchmark suites are self-contained and only require the Nemori memory instance and OpenAI API access for the LLM judge. The scripts read local data files (LoCoMo JSON or LongMemEval results) and invoke the standard Nemori API; no additional databases, vector stores, or third-party evaluation platforms are necessary beyond the configured OpenAI client.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →