How Supermemory's AI Model Works Internally: A Deep Dive into the Memory Middleware
Supermemory does not replace underlying language models like GPT-4 or Claude; instead, it provides a thin middleware layer that retrieves relevant memories, injects them into system prompts, and optionally stores conversations back to persistent storage.
Supermemory is an open-source memory layer for AI applications that enables persistent, context-aware conversations without modifying base language model weights. According to the supermemoryai/supermemory source code, the system operates as a middleware wrapper around existing Vercel AI SDK models, intercepting generation calls to enrich prompts with contextual memories while maintaining full compatibility with LanguageModelV2 and LanguageModelV3 implementations.
The Four-Stage Memory Pipeline
The Supermemory AI model architecture processes every language model interaction through four distinct stages, implemented primarily in packages/tools/src/vercel/middleware.ts and packages/tools/src/vercel/index.ts.
Stage 1: Context Creation
The pipeline begins when withSupermemory (implemented in wrapVercelLanguageModel at packages/tools/src/vercel/index.ts) receives the base language model and configuration parameters. This function validates the SUPERMEMORY_API_KEY environment variable and constructs a SupermemoryMiddlewareContext via createSupermemoryContext, which also normalizes the API base URL through normalizeBaseUrl.
The context object contains:
- A ready-to-use Supermemory API client (
new Supermemory({ apiKey, baseURL })) - A configurable
Loggerinstance for verbose debugging - The specified container tag (typically a user ID) for namespacing memories
- A per-turn
MemoryCache<string>that stores memory strings to avoid duplicate API calls
Stage 2: Parameter Transformation
Every call to model.doGenerate or model.doStream passes through transformParamsWithMemory before reaching the underlying model. This function extracts the user's last message and determines whether memory retrieval is necessary based on the configured mode option.
If the cache does not contain a valid entry for the current turn (keyed by makeTurnKey), the system calls buildMemoriesText to fetch memories from the Supermemory API via supermemoryProfileSearch. Depending on the mode setting:
"profile": Retrieves static and dynamic profile information without a search query"query": Performs semantic search using the extracted user message text"full": Merges both profile data and query-based search results
The fetched memories are then injected into the system prompt via injectMemoriesIntoParams (located in packages/tools/src/vercel/memory-prompt.ts), which either appends to an existing system message or creates a new one. This function operates immutably, returning transformed parameters without mutating the original object.
Stage 3: Model Invocation
After parameter enrichment, the wrapper forwards the transformed options to the original model's doGenerate or doStream methods unchanged. This architectural decision preserves full feature parity with the underlying Vercel AI SDK versions 5 and 6, ensuring that streaming, tool calling, and other advanced features work transparently.
Stage 4: Optional Persistence
When the addMemory configuration is set to "always", the system captures the assistant's response after generation completes. For streaming responses, the flush handler accumulates generatedText before storage; for standard generation, extractAssistantResponseText parses the response content.
The saveMemoryAfterResponse function then persists the interaction to the Supermemory API, choosing between addConversation for thread continuity or client.add for single document storage. This process records metadata including the container tag, custom conversation IDs, and content length for observability.
Core Implementation Files
Understanding the Supermemory AI model requires familiarity with these specific source files:
packages/tools/src/vercel/index.ts: Exports the publicwithSupermemorywrapper andwrapVercelLanguageModelfactorypackages/tools/src/vercel/middleware.ts: Contains the core middleware logic includingtransformParamsWithMemory,createSupermemoryContext,makeTurnKey, andsaveMemoryAfterResponsepackages/tools/src/vercel/memory-prompt.ts: ImplementsinjectMemoriesIntoParamsfor system prompt manipulationpackages/tools/src/shared/cache.ts: Defines theMemoryCacheclass and deterministic key generation logicpackages/tools/src/shared/memory-client.ts: Handles API communication viabuildMemoriesTextandsupermemoryProfileSearchpackages/tools/src/shared/logger.ts: Provides the verbose logging implementation used throughout the middleware
Implementation Examples
Basic Usage with Memory Injection
import { withSupermemory } from "@supermemory/tools/ai-sdk"
import { openai } from "@ai-sdk/openai"
import { generateText } from "ai"
const modelWithMemory = withSupermemory(openai("gpt-4"), "user-123", {
conversationId: "conv-001",
mode: "full",
addMemory: "always",
verbose: true,
})
const result = await generateText({
model: modelWithMemory,
messages: [{ role: "user", content: "What projects am I currently working on?" }],
})
console.log(result.text)
Streaming with Contextual Memory
import { withSupermemory } from "@supermemory/tools/ai-sdk"
import { openai } from "@ai-sdk/openai"
import { streamText } from "ai"
const model = withSupermemory(openai("gpt-4o-mini"), "team-42", {
mode: "query",
})
const { stream } = await streamText({
model,
messages: [{ role: "user", content: "Summarize the last meeting notes." }],
})
for await (const chunk of stream) {
process.stdout.write(chunk.textDelta)
}
Custom Memory Prompt Templates
import { withSupermemory } from "@supermemory/tools/ai-sdk"
import { openai } from "@ai-sdk/openai"
const model = withSupermemory(openai("gpt-4"), "org-789", {
promptTemplate: (data) => `
## Your Personal Context
${data.userMemories}
## Relevant Project Insights
${data.generalSearchMemories}
`.trim(),
})
The promptTemplate function receives structured memory data and returns a formatted string, defined in packages/tools/src/shared/prompt-builder.ts.
Summary
- Supermemory operates as middleware: It wraps existing language models without replacing them, intercepting calls to inject contextual memories into system prompts.
- Four-stage pipeline: Context creation → Parameter transformation (with caching) → Model invocation → Optional persistence.
- Caching mechanism: Per-turn
MemoryCacheuses deterministic keys frommakeTurnKeyto prevent redundant API calls during repeated or streaming operations. - Three retrieval modes:
"profile"for user data,"query"for semantic search, and"full"for combined context. - Immutable transformations: The
injectMemoriesIntoParamsfunction inpackages/tools/src/vercel/memory-prompt.tsenriches prompts without mutating original parameters. - Full SDK compatibility: Works transparently with Vercel AI SDK 5 and 6, supporting both
doGenerateanddoStreammethods.
Frequently Asked Questions
Does Supermemory replace the underlying AI model?
No, Supermemory does not replace or retrain underlying language models. According to the source code in packages/tools/src/vercel/middleware.ts, it provides a thin wrapper that intercepts calls to models like GPT-4 or Claude, enriches the system prompt with retrieved memories via injectMemoriesIntoParams, and then forwards the request to the original model's unchanged doGenerate or doStream methods.
How does Supermemory prevent duplicate memory API calls?
The system implements a per-turn MemoryCache defined in packages/tools/src/shared/cache.ts. Before fetching memories, transformParamsWithMemory generates a deterministic cache key using makeTurnKey, which incorporates the container tag, conversation ID, mode, and user message text. If the same turn is processed again (common in streaming scenarios), the cached memory string is reused instead of making redundant API requests.
What is the difference between the three memory modes?
The mode parameter controls memory retrieval behavior in buildMemoriesText: "profile" fetches static and dynamic user profile data without performing a search query; "query" executes a semantic search using the current user message as the search vector; and "full" combines both profile information and query-based search results for maximum context. These modes allow developers to optimize latency and relevance based on specific use cases.
How does conversation persistence work?
When addMemory is set to "always", the saveMemoryAfterResponse function in packages/tools/src/vercel/middleware.ts captures the assistant's output after generation completes. For streaming responses, the text accumulates in the flush handler; for standard generation, it extracts content via extractAssistantResponseText. The system then stores the interaction via the Supermemory API as either a conversation thread (using addConversation) or a standalone document (using client.add), tagging it with metadata for future retrieval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →