# How Supermemory's AI Model Works Internally: A Deep Dive into the Memory Middleware

> Discover how Supermemory's AI model works internally. Learn about its memory middleware that enhances models like GPT-4 by retrieving and injecting relevant context for improved performance.

- Repository: [supermemory/supermemory](https://github.com/supermemoryai/supermemory)
- Tags: deep-dive
- Published: 2026-03-25

---

**Supermemory does not replace underlying language models like GPT-4 or Claude; instead, it provides a thin middleware layer that retrieves relevant memories, injects them into system prompts, and optionally stores conversations back to persistent storage.**

Supermemory is an open-source memory layer for AI applications that enables persistent, context-aware conversations without modifying base language model weights. According to the supermemoryai/supermemory source code, the system operates as a **middleware wrapper** around existing Vercel AI SDK models, intercepting generation calls to enrich prompts with contextual memories while maintaining full compatibility with LanguageModelV2 and LanguageModelV3 implementations.

## The Four-Stage Memory Pipeline

The Supermemory AI model architecture processes every language model interaction through four distinct stages, implemented primarily in [`packages/tools/src/vercel/middleware.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/middleware.ts) and [`packages/tools/src/vercel/index.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/index.ts).

### Stage 1: Context Creation

The pipeline begins when `withSupermemory` (implemented in `wrapVercelLanguageModel` at [`packages/tools/src/vercel/index.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/index.ts)) receives the base language model and configuration parameters. This function validates the `SUPERMEMORY_API_KEY` environment variable and constructs a `SupermemoryMiddlewareContext` via `createSupermemoryContext`, which also normalizes the API base URL through `normalizeBaseUrl`.

The context object contains:

- A ready-to-use Supermemory API client (`new Supermemory({ apiKey, baseURL })`)
- A configurable `Logger` instance for verbose debugging
- The specified container tag (typically a user ID) for namespacing memories
- A per-turn `MemoryCache<string>` that stores memory strings to avoid duplicate API calls

### Stage 2: Parameter Transformation

Every call to `model.doGenerate` or `model.doStream` passes through `transformParamsWithMemory` before reaching the underlying model. This function extracts the user's last message and determines whether memory retrieval is necessary based on the configured `mode` option.

If the cache does not contain a valid entry for the current turn (keyed by `makeTurnKey`), the system calls `buildMemoriesText` to fetch memories from the Supermemory API via `supermemoryProfileSearch`. Depending on the mode setting:

- **`"profile"`**: Retrieves static and dynamic profile information without a search query
- **`"query"`**: Performs semantic search using the extracted user message text
- **`"full"`**: Merges both profile data and query-based search results

The fetched memories are then injected into the system prompt via `injectMemoriesIntoParams` (located in [`packages/tools/src/vercel/memory-prompt.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/memory-prompt.ts)), which either appends to an existing system message or creates a new one. This function operates immutably, returning transformed parameters without mutating the original object.

### Stage 3: Model Invocation

After parameter enrichment, the wrapper forwards the transformed options to the original model's `doGenerate` or `doStream` methods unchanged. This architectural decision preserves full feature parity with the underlying Vercel AI SDK versions 5 and 6, ensuring that streaming, tool calling, and other advanced features work transparently.

### Stage 4: Optional Persistence

When the `addMemory` configuration is set to `"always"`, the system captures the assistant's response after generation completes. For streaming responses, the `flush` handler accumulates `generatedText` before storage; for standard generation, `extractAssistantResponseText` parses the response content.

The `saveMemoryAfterResponse` function then persists the interaction to the Supermemory API, choosing between `addConversation` for thread continuity or `client.add` for single document storage. This process records metadata including the container tag, custom conversation IDs, and content length for observability.

## Core Implementation Files

Understanding the **Supermemory AI model** requires familiarity with these specific source files:

- **[`packages/tools/src/vercel/index.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/index.ts)**: Exports the public `withSupermemory` wrapper and `wrapVercelLanguageModel` factory
- **[`packages/tools/src/vercel/middleware.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/middleware.ts)**: Contains the core middleware logic including `transformParamsWithMemory`, `createSupermemoryContext`, `makeTurnKey`, and `saveMemoryAfterResponse`
- **[`packages/tools/src/vercel/memory-prompt.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/memory-prompt.ts)**: Implements `injectMemoriesIntoParams` for system prompt manipulation
- **[`packages/tools/src/shared/cache.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/shared/cache.ts)**: Defines the `MemoryCache` class and deterministic key generation logic
- **[`packages/tools/src/shared/memory-client.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/shared/memory-client.ts)**: Handles API communication via `buildMemoriesText` and `supermemoryProfileSearch`
- **[`packages/tools/src/shared/logger.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/shared/logger.ts)**: Provides the verbose logging implementation used throughout the middleware

## Implementation Examples

### Basic Usage with Memory Injection

```typescript
import { withSupermemory } from "@supermemory/tools/ai-sdk"
import { openai } from "@ai-sdk/openai"
import { generateText } from "ai"

const modelWithMemory = withSupermemory(openai("gpt-4"), "user-123", {
  conversationId: "conv-001",
  mode: "full",
  addMemory: "always",
  verbose: true,
})

const result = await generateText({
  model: modelWithMemory,
  messages: [{ role: "user", content: "What projects am I currently working on?" }],
})

console.log(result.text)

```

### Streaming with Contextual Memory

```typescript
import { withSupermemory } from "@supermemory/tools/ai-sdk"
import { openai } from "@ai-sdk/openai"
import { streamText } from "ai"

const model = withSupermemory(openai("gpt-4o-mini"), "team-42", {
  mode: "query",
})

const { stream } = await streamText({
  model,
  messages: [{ role: "user", content: "Summarize the last meeting notes." }],
})

for await (const chunk of stream) {
  process.stdout.write(chunk.textDelta)
}

```

### Custom Memory Prompt Templates

```typescript
import { withSupermemory } from "@supermemory/tools/ai-sdk"
import { openai } from "@ai-sdk/openai"

const model = withSupermemory(openai("gpt-4"), "org-789", {
  promptTemplate: (data) => `
    ## Your Personal Context

    ${data.userMemories}

    ## Relevant Project Insights

    ${data.generalSearchMemories}
  `.trim(),
})

```

The `promptTemplate` function receives structured memory data and returns a formatted string, defined in [`packages/tools/src/shared/prompt-builder.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/shared/prompt-builder.ts).

## Summary

- **Supermemory operates as middleware**: It wraps existing language models without replacing them, intercepting calls to inject contextual memories into system prompts.
- **Four-stage pipeline**: Context creation → Parameter transformation (with caching) → Model invocation → Optional persistence.
- **Caching mechanism**: Per-turn `MemoryCache` uses deterministic keys from `makeTurnKey` to prevent redundant API calls during repeated or streaming operations.
- **Three retrieval modes**: `"profile"` for user data, `"query"` for semantic search, and `"full"` for combined context.
- **Immutable transformations**: The `injectMemoriesIntoParams` function in [`packages/tools/src/vercel/memory-prompt.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/memory-prompt.ts) enriches prompts without mutating original parameters.
- **Full SDK compatibility**: Works transparently with Vercel AI SDK 5 and 6, supporting both `doGenerate` and `doStream` methods.

## Frequently Asked Questions

### Does Supermemory replace the underlying AI model?

No, Supermemory does not replace or retrain underlying language models. According to the source code in [`packages/tools/src/vercel/middleware.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/middleware.ts), it provides a thin wrapper that intercepts calls to models like GPT-4 or Claude, enriches the system prompt with retrieved memories via `injectMemoriesIntoParams`, and then forwards the request to the original model's unchanged `doGenerate` or `doStream` methods.

### How does Supermemory prevent duplicate memory API calls?

The system implements a per-turn `MemoryCache` defined in [`packages/tools/src/shared/cache.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/shared/cache.ts). Before fetching memories, `transformParamsWithMemory` generates a deterministic cache key using `makeTurnKey`, which incorporates the container tag, conversation ID, mode, and user message text. If the same turn is processed again (common in streaming scenarios), the cached memory string is reused instead of making redundant API requests.

### What is the difference between the three memory modes?

The `mode` parameter controls memory retrieval behavior in `buildMemoriesText`: `"profile"` fetches static and dynamic user profile data without performing a search query; `"query"` executes a semantic search using the current user message as the search vector; and `"full"` combines both profile information and query-based search results for maximum context. These modes allow developers to optimize latency and relevance based on specific use cases.

### How does conversation persistence work?

When `addMemory` is set to `"always"`, the `saveMemoryAfterResponse` function in [`packages/tools/src/vercel/middleware.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/tools/src/vercel/middleware.ts) captures the assistant's output after generation completes. For streaming responses, the text accumulates in the `flush` handler; for standard generation, it extracts content via `extractAssistantResponseText`. The system then stores the interaction via the Supermemory API as either a conversation thread (using `addConversation`) or a standalone document (using `client.add`), tagging it with metadata for future retrieval.