# Nemori Evaluation Benchmarks: LoCoMo and LongMemEval Explained

> Discover Nemori's evaluation benchmarks, LoCoMo and LongMemEval. Learn how they measure retrieval accuracy and LLM alignment for long-context reasoning.

- Repository: [Nemori AI/nemori](https://github.com/nemori-ai/nemori)
- Tags: benchmarks
- Published: 2026-03-08

---

**Nemori uses two primary evaluation benchmarks—LoCoMo (Long-Context Memory) and LongMemEval—to measure retrieval accuracy, long-context reasoning, and LLM alignment using automated LLM judges and lexical metrics.**

The `nemori-ai/nemori` repository provides self-contained evaluation suites that stress-test the memory system against public datasets. These benchmarks verify how well Nemori retrieves and aligns long-context conversations across diverse question categories and 100k-token contexts.

## LoCoMo Benchmark

The **LoCoMo** (Long-Context Memory) benchmark evaluates Nemori’s ability to retrieve and align long-context conversations across four question categories: **Multi-Hop**, **Temporal-Reasoning**, **Open-Domain**, and **Single-Hop**.

### Pipeline Components

The LoCoMo evaluation follows a four-stage pipeline implemented in `evaluation/locomo/`:

1. **Data Ingestion** – [`evaluation/locomo/add.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/locomo/add.py) loads the LoCoMo JSON dataset and streams conversations into Nemori through the `NemoriMemory` facade:

```python
from nemori import NemoriMemory, MemoryConfig

cfg = MemoryConfig(...)
memory = NemoriMemory(config=cfg)
memory.add_messages(user_id="locomo_user", messages=conversation)

```

2. **Parallel Search** – [`evaluation/locomo/search.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/locomo/search.py) executes concurrent vector, BM25, and original-text searches against stored episodes, writing raw results to [`locomo/results.json`](https://github.com/nemori-ai/nemori/blob/main/locomo/results.json).

3. **Scoring** – [`evaluation/locomo/evals.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/locomo/evals.py) computes **BLEU-1** and **F1** lexical scores, then invokes the LLM judge via `metrics/llm_judge.evaluate_llm_judge` to generate an *LLM alignment* score.

4. **Aggregation** – [`evaluation/locomo/generate_scores.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/locomo/generate_scores.py) aggregates per-category metrics and generates the final score table displayed in the README.

### Metrics and Scoring

LoCoMo employs three complementary metrics to judge response quality:

- **BLEU-1**: Lexical overlap between generated and reference answers
- **F1**: Token-level precision and recall for semantic similarity
- **LLM-alignment score**: Binary judgment from an LLM judge (configured to use OpenAI’s `gpt-4.1-mini` or `gpt-4o-mini`) determining if the response matches the gold answer

## LongMemEval Benchmark

**LongMemEval** tests Nemori on a **100k-token context** benchmark that stresses temporal reasoning, knowledge updates, and personalized preference recall.

### Async LLM Judge Implementation

The evaluator in [`evaluation/longmemeval/evals.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/longmemeval/evals.py) implements an asynchronous evaluation pipeline:

1. Loads a results file (e.g., [`longmem_results.json`](https://github.com/nemori-ai/nemori/blob/main/longmem_results.json)) produced by any Nemori run
2. Selects prompt templates based on question type (`temporal-reasoning`, `knowledge-update`, `single-session-preference`, or default)
3. Sends prompts to the configured LLM (`self.model`) via the async OpenAI client
4. Parses JSON responses into a `Grade` model recording **binary correctness**
5. Calculates statistics including overall accuracy, per-type accuracy, average response time, and error-case analysis

### Evaluation Metrics

LongMemEval reports four key performance indicators:

- **Overall accuracy**: Percentage of correctly answered questions
- **Accuracy per question type**: Breakdown across temporal reasoning, knowledge updates, and preference recall
- **Average response time**: Latency measurements for the evaluation loop
- **Error-case analysis**: Structured list of failed retrievals for debugging

## Running the Benchmarks

Execute the LoCoMo pipeline from the repository root to reproduce the published metrics:

```bash
cd evaluation
python locomo/add.py              # Ingest LoCoMo dataset into Nemori

python locomo/search.py           # Perform parallel searches

python locomo/evals.py            # Compute BLEU/F1 and LLM scores

python locomo/generate_scores.py  # Produce the final comparison table

```

Run the LongMemEval benchmark against existing results:

```bash
python evaluation/longmemeval/evals.py longmem_results.json \
  --model gpt-4.1-mini \
  --output report.json

```

Both scripts are **self-contained** and require only the OpenAI API for LLM judging—no external databases or services beyond the Nemori memory instance itself.

## Summary

- **LoCoMo** evaluates long-context retrieval across four question categories using [`evaluation/locomo/add.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/locomo/add.py), [`search.py`](https://github.com/nemori-ai/nemori/blob/main/search.py), [`evals.py`](https://github.com/nemori-ai/nemori/blob/main/evals.py), and [`generate_scores.py`](https://github.com/nemori-ai/nemori/blob/main/generate_scores.py)
- **LongMemEval** assesses 100k-token context handling via [`evaluation/longmemeval/evals.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/longmemeval/evals.py) with async LLM judges
- Both benchmarks use **LLM-based judges** (`gpt-4.1-mini` or `gpt-4o-mini`) for binary correctness determination
- **BLEU-1** and **F1** provide lexical and semantic scoring alongside LLM alignment metrics
- All evaluation scripts operate against the standard `NemoriMemory` API without requiring external service dependencies

## Frequently Asked Questions

### What LLM models does Nemori use for benchmark judging?

Nemori configures OpenAI’s `gpt-4.1-mini` or `gpt-4o-mini` as the default judge for both LoCoMo and LongMemEval. The `evaluate_llm_judge` function in the metrics module sends prompts to these models via the async OpenAI client and parses the JSON responses into binary correctness flags.

### How does the LoCoMo pipeline handle data ingestion?

The [`evaluation/locomo/add.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/locomo/add.py) script loads the LoCoMo JSON dataset and streams each conversation into Nemori using the `NemoriMemory` facade. It initializes a `MemoryConfig` instance matching the quick-start configuration, then calls `add_messages()` with a consistent `user_id` to populate the memory store for subsequent retrieval testing.

### What question categories does LongMemEval test?

LongMemEval evaluates four specific question types: **temporal-reasoning** (chronological inference across sessions), **knowledge-update** (handling revised facts), **single-session-preference** (recalling user preferences within one conversation), and a default category for general retrieval. The evaluator selects prompt templates dynamically based on these categories in [`evaluation/longmemeval/evals.py`](https://github.com/nemori-ai/nemori/blob/main/evaluation/longmemeval/evals.py).

### Can I run these benchmarks without external services?

Yes. Both benchmark suites are **self-contained** and only require the Nemori memory instance and OpenAI API access for the LLM judge. The scripts read local data files (LoCoMo JSON or LongMemEval results) and invoke the standard Nemori API; no additional databases, vector stores, or third-party evaluation platforms are necessary beyond the configured OpenAI client.