# How Supermemory Handles Data Storage: Cloudflare D1, pgvector, and Knowledge Graphs

> Discover how Supermemory stores user data in Cloudflare D1 with pgvector, building a knowledge graph via an edge-based ingestion pipeline.

- Repository: [supermemory/supermemory](https://github.com/supermemoryai/supermemory)
- Tags: deep-dive
- Published: 2026-03-25

---

**Supermemory stores all user data—raw documents, semantic chunks, and vector embeddings—in a Cloudflare D1/Hyperdrive relational database augmented with the pgvector extension, processing content through an edge-based ingestion pipeline that constructs a living knowledge graph.**

Supermemory is an open-source memory layer for AI applications designed to persist both structured metadata and high-dimensional vector embeddings in a unified backend. The platform's storage architecture leverages Cloudflare's edge-native infrastructure to minimize latency while maintaining ACID compliance for relationship-heavy data.

## Cloud-Native Storage Stack

Supermemory consolidates user data into a **single Cloudflare-powered database** that combines traditional relational capabilities with vector search extensions. The platform declares its required bindings in [[`CLAUDE.md`](https://github.com/supermemoryai/supermemory/blob/main/CLAUDE.md)](https://github.com/supermemoryai/supermemory/blob/main/CLAUDE.md), specifying four core storage layers:

- **Primary Database**: **Cloudflare D1 with Hyperdrive** (SQL-compatible) stores structured tables for documents, memories, user profiles, projects, and knowledge-graph relationships.
- **Vector Store**: The **pgvector** extension (PostgreSQL-compatible) running on D1/Hyperdrive manages high-dimensional embeddings for semantic similarity search.
- **Key-Value Layer**: **Cloudflare KV** persists small configuration objects, feature flags, and ephemeral session data.
- **Compute Layer**: **Cloudflare Workers plus AI bindings** execute the ingestion pipeline at the edge, handling extraction, chunking, and embedding generation.

According to the self-hosting documentation in [`apps/docs/deployment/self-hosting.mdx`](https://github.com/supermemoryai/supermemory/blob/main/apps/docs/deployment/self-hosting.mdx#L56-L58), any PostgreSQL-compatible database must support the **pgvector** extension to function as the storage backend.

## Data Flow: From Upload to Knowledge Graph

The ingestion pipeline transforms raw content into queryable memories through six distinct stages:

1. **Raw Ingestion**: Content uploads via `POST /v3/documents` are stored as raw blobs in the primary database.
2. **Extraction**: The system parses text and metadata from PDFs, images, videos, and URLs.
3. **Semantic Chunking**: Text splits into paragraph-level semantic chunks rather than fixed-size windows.
4. **Vector Embedding**: Each chunk is sent to the AI provider, generating **1536-dimensional vectors** that are persisted in the pgvector column.
5. **Relationship Indexing**: The system creates three relationship types—**Updates**, **Extends**, and **Derives**—linking memories into a living knowledge graph as described in [[`skills/supermemory/references/architecture.md`](https://github.com/supermemoryai/supermemory/blob/main/skills/supermemory/references/architecture.md)](https://github.com/supermemoryai/supermemory/blob/main/skills/supermemory/references/architecture.md#memory-storage-system).
6. **Semantic Search**: Query requests hit `/v3/search`, executing cosine-similarity queries against the vector column using logic from [[`packages/memory-graph/src/lib/similarity.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/memory-graph/src/lib/similarity.ts)](https://github.com/supermemoryai/supermemory/blob/main/packages/memory-graph/src/lib/similarity.ts).

## Schema Design and Type Safety

Supermemory uses **Drizzle-ORM** with Zod validation schemas defined in [[`packages/validation/schemas.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/validation/schemas.ts)](https://github.com/supermemoryai/supermemory/blob/main/packages/validation/schemas.ts). The schema enforces strict typing for vector dimensions and relationship semantics:

```typescript
// packages/validation/schemas.ts (excerpt)
export const schemas = {
  memory: z.object({
    id: z.string().uuid(),
    content: z.string(),
    vector: z.array(z.number()).length(1536), // pgvector column
    isStatic: z.boolean().default(false),
    containerTag: z.string(),
    metadata: z.record(z.unknown()).optional(),
  }),

  relationship: z.object({
    fromId: z.string().uuid(),
    toId: z.string().uuid(),
    type: z.enum(["updates", "extends", "derives"]),
  }),
};

```

The `memory` table stores the embedding vector as a fixed-length array of 1536 floats, while the `relationship` table implements a typed graph structure connecting memories through semantic lineage.

## Programmatic Storage with the TypeScript SDK

Developers interact with the storage layer through the official SDK, where the `createMemory` method abstracts the entire pipeline. Defined in [[`apps/mcp/src/client.ts`](https://github.com/supermemoryai/supermemory/blob/main/apps/mcp/src/client.ts)](https://github.com/supermemoryai/supermemory/blob/main/apps/mcp/src/client.ts) and consumed by the server in [[`apps/mcp/src/server.ts`](https://github.com/supermemoryai/supermemory/blob/main/apps/mcp/src/server.ts)](https://github.com/supermemoryai/supermemory/blob/main/apps/mcp/src/server.ts), this method automatically executes extraction, chunking, embedding, and graph indexing:

```typescript
import { createClient } from "@supermemoryai/tools";

const client = createClient({
  apiKey: process.env.SUPERMEMORY_API_KEY,
  baseUrl: "https://api.supermemory.ai",
});

// Automatically runs: extract → chunk → embed → index → store
await client.createMemory({
  content: "User prefers TypeScript with strict mode",
  containerTag: "project-123",
  metadata: { source: "user-settings" },
});

```

## Summary

- **Supermemory** uses **Cloudflare D1/Hyperdrive** as its primary database, combining relational ACID transactions with vector search via the **pgvector** extension.
- The ingestion pipeline automatically processes raw content into **1536-dimensional embeddings** and links them via **Updates**, **Extends**, and **Derives** relationships to form a queryable knowledge graph.
- **Cloudflare KV** handles configuration and session state, while **Cloudflare Workers** execute the extraction and embedding pipeline at the edge.
- Type-safe schemas in [`packages/validation/schemas.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/validation/schemas.ts) enforce data integrity using Drizzle-ORM and Zod validation.

## Frequently Asked Questions

### What database does Supermemory use for production deployments?

Supermemory uses **Cloudflare D1** paired with **Hyperdrive** for production workloads. According to the source code in [`CLAUDE.md`](https://github.com/supermemoryai/supermemory/blob/main/CLAUDE.md), the platform binds to Hyperdrive for SQL-compatible storage, while the self-hosting guide in `self-hosting.mdx` confirms that any PostgreSQL-compatible database with the **pgvector** extension can serve as the backend.

### How does Supermemory store vector embeddings for semantic search?

Vector embeddings are stored as **1536-dimensional arrays** in a dedicated column using the **pgvector** extension. The schema definition in [`packages/validation/schemas.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/validation/schemas.ts) validates that all vectors are exactly 1536 floats, and the similarity search logic in [`packages/memory-graph/src/lib/similarity.ts`](https://github.com/supermemoryai/supermemory/blob/main/packages/memory-graph/src/lib/similarity.ts) executes cosine-similarity calculations against this column.

### What types of relationships exist in the Supermemory knowledge graph?

The knowledge graph supports three relationship types defined in the Zod schema: **"updates"** (indicating versioned replacements), **"extends"** (indicating additive information), and **"derives"** (indicating inferred or computed knowledge). These relationships link memory nodes together and are stored in the relational database alongside the vector embeddings.

### Can I self-host Supermemory with my own database infrastructure?

Yes, Supermemory supports self-hosting on any PostgreSQL-compatible database that includes the **pgvector** extension. The deployment documentation specifies that you must configure the `HYPERDRIVE` binding to point to your PostgreSQL instance, and the system will automatically handle schema migrations through Drizzle-ORM.