How Supermemory Handles Data Storage: Cloudflare D1, pgvector, and Knowledge Graphs
Supermemory stores all user data—raw documents, semantic chunks, and vector embeddings—in a Cloudflare D1/Hyperdrive relational database augmented with the pgvector extension, processing content through an edge-based ingestion pipeline that constructs a living knowledge graph.
Supermemory is an open-source memory layer for AI applications designed to persist both structured metadata and high-dimensional vector embeddings in a unified backend. The platform's storage architecture leverages Cloudflare's edge-native infrastructure to minimize latency while maintaining ACID compliance for relationship-heavy data.
Cloud-Native Storage Stack
Supermemory consolidates user data into a single Cloudflare-powered database that combines traditional relational capabilities with vector search extensions. The platform declares its required bindings in [CLAUDE.md](https://github.com/supermemoryai/supermemory/blob/main/CLAUDE.md), specifying four core storage layers:
- Primary Database: Cloudflare D1 with Hyperdrive (SQL-compatible) stores structured tables for documents, memories, user profiles, projects, and knowledge-graph relationships.
- Vector Store: The pgvector extension (PostgreSQL-compatible) running on D1/Hyperdrive manages high-dimensional embeddings for semantic similarity search.
- Key-Value Layer: Cloudflare KV persists small configuration objects, feature flags, and ephemeral session data.
- Compute Layer: Cloudflare Workers plus AI bindings execute the ingestion pipeline at the edge, handling extraction, chunking, and embedding generation.
According to the self-hosting documentation in apps/docs/deployment/self-hosting.mdx, any PostgreSQL-compatible database must support the pgvector extension to function as the storage backend.
Data Flow: From Upload to Knowledge Graph
The ingestion pipeline transforms raw content into queryable memories through six distinct stages:
- Raw Ingestion: Content uploads via
POST /v3/documentsare stored as raw blobs in the primary database. - Extraction: The system parses text and metadata from PDFs, images, videos, and URLs.
- Semantic Chunking: Text splits into paragraph-level semantic chunks rather than fixed-size windows.
- Vector Embedding: Each chunk is sent to the AI provider, generating 1536-dimensional vectors that are persisted in the pgvector column.
- Relationship Indexing: The system creates three relationship types—Updates, Extends, and Derives—linking memories into a living knowledge graph as described in [
skills/supermemory/references/architecture.md](https://github.com/supermemoryai/supermemory/blob/main/skills/supermemory/references/architecture.md#memory-storage-system). - Semantic Search: Query requests hit
/v3/search, executing cosine-similarity queries against the vector column using logic from [packages/memory-graph/src/lib/similarity.ts](https://github.com/supermemoryai/supermemory/blob/main/packages/memory-graph/src/lib/similarity.ts).
Schema Design and Type Safety
Supermemory uses Drizzle-ORM with Zod validation schemas defined in [packages/validation/schemas.ts](https://github.com/supermemoryai/supermemory/blob/main/packages/validation/schemas.ts). The schema enforces strict typing for vector dimensions and relationship semantics:
// packages/validation/schemas.ts (excerpt)
export const schemas = {
memory: z.object({
id: z.string().uuid(),
content: z.string(),
vector: z.array(z.number()).length(1536), // pgvector column
isStatic: z.boolean().default(false),
containerTag: z.string(),
metadata: z.record(z.unknown()).optional(),
}),
relationship: z.object({
fromId: z.string().uuid(),
toId: z.string().uuid(),
type: z.enum(["updates", "extends", "derives"]),
}),
};
The memory table stores the embedding vector as a fixed-length array of 1536 floats, while the relationship table implements a typed graph structure connecting memories through semantic lineage.
Programmatic Storage with the TypeScript SDK
Developers interact with the storage layer through the official SDK, where the createMemory method abstracts the entire pipeline. Defined in [apps/mcp/src/client.ts](https://github.com/supermemoryai/supermemory/blob/main/apps/mcp/src/client.ts) and consumed by the server in [apps/mcp/src/server.ts](https://github.com/supermemoryai/supermemory/blob/main/apps/mcp/src/server.ts), this method automatically executes extraction, chunking, embedding, and graph indexing:
import { createClient } from "@supermemoryai/tools";
const client = createClient({
apiKey: process.env.SUPERMEMORY_API_KEY,
baseUrl: "https://api.supermemory.ai",
});
// Automatically runs: extract → chunk → embed → index → store
await client.createMemory({
content: "User prefers TypeScript with strict mode",
containerTag: "project-123",
metadata: { source: "user-settings" },
});
Summary
- Supermemory uses Cloudflare D1/Hyperdrive as its primary database, combining relational ACID transactions with vector search via the pgvector extension.
- The ingestion pipeline automatically processes raw content into 1536-dimensional embeddings and links them via Updates, Extends, and Derives relationships to form a queryable knowledge graph.
- Cloudflare KV handles configuration and session state, while Cloudflare Workers execute the extraction and embedding pipeline at the edge.
- Type-safe schemas in
packages/validation/schemas.tsenforce data integrity using Drizzle-ORM and Zod validation.
Frequently Asked Questions
What database does Supermemory use for production deployments?
Supermemory uses Cloudflare D1 paired with Hyperdrive for production workloads. According to the source code in CLAUDE.md, the platform binds to Hyperdrive for SQL-compatible storage, while the self-hosting guide in self-hosting.mdx confirms that any PostgreSQL-compatible database with the pgvector extension can serve as the backend.
How does Supermemory store vector embeddings for semantic search?
Vector embeddings are stored as 1536-dimensional arrays in a dedicated column using the pgvector extension. The schema definition in packages/validation/schemas.ts validates that all vectors are exactly 1536 floats, and the similarity search logic in packages/memory-graph/src/lib/similarity.ts executes cosine-similarity calculations against this column.
What types of relationships exist in the Supermemory knowledge graph?
The knowledge graph supports three relationship types defined in the Zod schema: "updates" (indicating versioned replacements), "extends" (indicating additive information), and "derives" (indicating inferred or computed knowledge). These relationships link memory nodes together and are stored in the relational database alongside the vector embeddings.
Can I self-host Supermemory with my own database infrastructure?
Yes, Supermemory supports self-hosting on any PostgreSQL-compatible database that includes the pgvector extension. The deployment documentation specifies that you must configure the HYPERDRIVE binding to point to your PostgreSQL instance, and the system will automatically handle schema migrations through Drizzle-ORM.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →