Supermemory Performance Benchmarks: LongMemEval, LoCoMo, and ConvoMem Results

Supermemory ranks #1 on LongMemEval (81.6%), LoCoMo, and ConvoMem through its Memory Engine and hybrid search architecture, with reproducible evaluation via the open-source MemoryBench CLI.

The supermemoryai/supermemory repository provides an open-source memory layer for AI applications that currently leads the three most-cited AI-memory benchmarks. Understanding Supermemory performance benchmarks requires examining both the quantitative results on standardized datasets and the architectural components that enable these scores. The project ships with MemoryBench, a reproducible framework that validates these claims against the same judges and random seeds used for official results.

Supermemory Performance Benchmark Results

Supermemory consistently achieves top positions across the dominant evaluation suites for long-context and personalized memory systems:

Benchmark Measurement Focus Result
LongMemEval Long-term memory across sessions with knowledge updates and temporal reasoning 81.6% — #1
LoCoMo Fact recall across extended conversations including single-hop, multi-hop, temporal, and adversarial queries #1
ConvoMem Personalization and user preference learning over conversational history #1

These scores are generated using the MemoryBench framework, an open-source evaluation suite located in the repository that enables side-by-side comparison with other memory providers like Mem0 and Zep.

Architecture Behind the Benchmark Scores

Supermemory’s benchmark dominance stems from a tightly integrated stack designed to optimize the specific dimensions each test measures. The architecture is documented in apps/docs/memorybench/architecture.mdx and consists of four primary layers:

Memory Engine

The Memory Engine extracts factual statements from conversations, tracks incremental updates, resolves contradictions, and applies automatic forgetting for stale facts. This layer directly powers the temporal reasoning scores on LongMemEval and ensures preference consistency required by ConvoMem.

User Profiles

User Profiles combine static facts with the most recent dynamic context, refreshing on every request in approximately 50 ms. This provides a single source of truth for personalization, which is critical for achieving high scores on ConvoMem’s personalization tasks. The profile retrieval is implemented in packages/tools/src/ai-sdk.ts.

Hybrid Search merges traditional RAG document retrieval with memory-based recall in a single query pipeline. This dual-mode approach enables robust handling of multi-hop and adversarial questions in the LoCoMo benchmark by combining vector similarity with graph-based relationship traversal.

Connectors and File Processing

Real-time Connectors sync data from Google Drive, Gmail, Notion, and GitHub to keep the knowledge base current, while File Processing handles PDF OCR, video transcription, and code AST-aware chunking. These components ensure high-quality context ingestion that improves long-term recall metrics across all benchmarks.

MemoryBench: Reproducible Evaluation Framework

MemoryBench provides deterministic, reproducible evaluation through a standardized CLI that eliminates variance between runs. The framework uses adapters for LongMemEval, LoCoMo, and ConvoMem datasets with a GPT-4o judge to ensure consistent scoring.

To reproduce the published scores locally, use the CLI command documented in apps/docs/memorybench/cli.mdx:

bun run src/index.ts run -p supermemory -b longmemeval -j gpt-4o -r my-run

This command outputs a MemScore summary containing precision, recall, and temporal consistency metrics. The validation schemas that ensure data integrity during benchmarking are defined in packages/validation/schemas.ts.

Running Performance Benchmarks

Benchmarking via CLI

The primary method for running Supermemory performance benchmarks uses the MemoryBench CLI with Bun:


# Install dependencies

bun i

# Run LongMemEval against Supermemory with GPT-4o judging

bun run src/index.ts run -p supermemory -b longmemeval -j gpt-4o -r my-run

The CLI prints a formatted table showing MemScore metrics and allows comparison across different memory providers using identical datasets and judges.

Profiling via TypeScript SDK

For programmatic evaluation, the Supermemory client provides profile retrieval with integrated hybrid search. This example from packages/tools/src/ai-sdk.ts demonstrates fetching a user profile and relevant memories in a single round-trip:

import Supermemory from "supermemory";

const client = new Supermemory({ apiKey: process.env.SUPERMEMORY_API_KEY });

async function evaluateProfile() {
  const { profile, searchResults } = await client.profile({
    containerTag: "user_123",
    query: "What does the user like to code with?"
  });

  console.log("Profile:", profile);
  console.log("Relevant Memories:", searchResults);
}

evaluateProfile();

The profile() method executes in approximately 50ms and returns both the auto-maintained user profile and semantic search results, representing the core of the User Profiles layer that powers high ConvoMem scores.

Extending MemoryBench with Custom Datasets

MemoryBench supports custom benchmarks through the BenchmarkAdapter interface. To add a new dataset, implement the interface in a new file and register it in src/benchmarks/index.ts as described in apps/docs/memorybench/extend-benchmark.mdx:

// src/benchmarks/mybenchmark/index.ts
import { BenchmarkAdapter } from "@memorybench/types";

export class MyBenchmark implements BenchmarkAdapter {
  name = "mybenchmark";
  
  async loadQuestions() {
    return await import("./questions.json");
  }
  
  async parseResponse(raw: string) {
    const data = JSON.parse(raw);
    return { answer: data.text, confidence: data.conf };
  }
}

Register the implementation:

// src/benchmarks/index.ts
import { MyBenchmark } from "./mybenchmark";

export const benchmarks = {
  ...existingBenchmarks,
  mybenchmark: MyBenchmark,
};

Execute the custom benchmark using the same CLI syntax:

bun run src/index.ts run -p supermemory -b mybenchmark -j gpt-4o -r my-run

Summary

  • Supermemory achieves 81.6% on LongMemEval and #1 rankings on LoCoMo and ConvoMem through its specialized memory architecture.
  • The MemoryBench framework provides reproducible evaluation via CLI using standardized GPT-4o judges and deterministic random seeds.
  • Core architectural components include the Memory Engine for temporal reasoning, User Profiles for personalization, and Hybrid Search for multi-hop queries.
  • Benchmark adapters are extensible via TypeScript, allowing custom datasets to be evaluated against Supermemory using the same methodology as official benchmarks.

Frequently Asked Questions

What is MemoryBench in Supermemory?

MemoryBench is an open-source evaluation framework shipped with Supermemory that provides reproducible benchmarking against LongMemEval, LoCoMo, and ConvoMem datasets. It uses a GPT-4o judge and standardized adapters to generate MemScore summaries that can be compared side-by-side with other memory providers like Mem0 and Zep.

How does Supermemory achieve #1 on LongMemEval?

Supermemory scores 81.6% on LongMemEval through its Memory Engine, which explicitly tracks temporal updates and resolves contradictions across conversation sessions. The engine automatically forgets stale facts while preserving critical knowledge, directly addressing the benchmark’s focus on long-term memory with knowledge updates.

Can I run benchmarks against my own memory implementation?

Yes, the MemoryBench framework is extensible. You can implement the BenchmarkAdapter interface in TypeScript, register your provider in src/benchmarks/index.ts, and run comparative evaluations using the CLI command bun run src/index.ts run -p your-provider -b longmemeval -j gpt-4o.

Where are the benchmark validation schemas defined?

Validation schemas for memory entries, relations, and API payloads—which ensure consistent benchmark input—are defined in packages/validation/schemas.ts. These Zod schemas enforce data contracts between the benchmark runner and the Supermemory API, guaranteeing that all evaluated providers process identically structured requests.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →