What Is the Database Schema for Supermemory? A Complete Technical Reference
Supermemory uses a relational database schema managed by Drizzle ORM and enforced through Zod validation, with core entities including documents, chunks, memory entries, spaces, and connections defined in packages/validation/schemas.ts.
The Supermemory database schema powers an AI-native knowledge base that ingests files, extracts semantic meaning, and enables collaborative organization. Defined primarily in packages/validation/schemas.ts (lines 61‑309), the schema leverages Drizzle ORM for type-safe SQL generation and Zod for runtime data validation. This architecture ensures that the TypeScript types used throughout the application match the actual database constraints.
Core Entity Tables
The schema organizes data into logical entities that handle content ingestion, vectorization, and access control.
Document Table
The Document table stores every uploaded or linked file, from PDFs to web pages. In the DocumentSchema defined at packages/validation/schemas.ts, the table uses an id field (documentId) as its primary key and includes columns for title, type, status, visibility, createdAt, updatedAt, and a flexible metadata field typed as z.record(z.string(), z.union([z.string(), z.number(), z.boolean()])).
MemoryEntry Table
The MemoryEntry table captures processed insights and summaries extracted from document chunks. Each row contains an id, documentId, the extracted memory text, and optional vector embeddings stored in memoryEmbedding as z.array(z.number()).nullable(). The memoryRelations column uses z.record(MemoryRelationEnum).default({}) to track semantic relationships like "updates" or "extends" between memories, while memoryEmbeddingModel records which model generated the vectors.
Space Table
Spaces act as logical containers that group documents and collaborators. The SpaceSchema defines columns for id, name, ownerId, createdAt, and updatedAt, providing the foundation for multi-tenant organization of knowledge bases.
Chunk Table
The Chunk table represents subdivisions of documents prepared for semantic indexing. Using a composite id field alongside documentId, chunks store content, a nullable embedding vector as z.array(z.number()).nullable(), and a type field constrained by ChunkTypeEnum. This table bridges raw documents and their searchable vector representations.
Supporting Tables and Junctions
Beyond core entities, the schema includes infrastructure for integrations, processing pipelines, and many-to-many relationships.
Connection Table
External integrations with Notion, Google Drive, and OneDrive reside in the Connection table. Defined in ConnectionSchema, it stores provider (constrained by ConnectionProviderEnum), status, and provider-specific metadata as a record type, enabling OAuth-based data synchronization.
OrganizationSettings Table
Global configuration for teams lives in OrganizationSettings, keyed by orgId. The table stores plan information, feature flags in the features column, and timestamps, allowing per-organization customization of the Supermemory instance.
ProcessingStep
The ProcessingStep schema tracks ingestion pipeline progress without a dedicated primary key, functioning as an embedded structure. It records stepName, status, startedAt, and finishedAt to monitor asynchronous document processing workflows.
Junction Tables
Many-to-many relationships use explicit junction tables:
- DocumentsToSpaces: Links
documentIdtospaceIdfor multi-space document organization - SpacesToMembers: Associates
spaceIdwithuserIdand assigns arole(constrained bySpaceRoleEnum) for collaborative access control
Enums and Type Constraints
The schema enforces data integrity through TypeScript-native enums defined in packages/validation/schemas.ts:
- VisibilityEnum (
"public" | "private" | "unlisted") – line 8 - DocumentTypeEnum (
"pdf" | "url" | "image" | "video" | "text") – line 11 - DocumentStatusEnum (
"processing" | "ready" | "error") – line 26 - ChunkTypeEnum (
"text" | "image") – line 103 - ConnectionProviderEnum (
"notion" | "googleDrive" | "oneDrive") – line 126 - MemoryRelationEnum (
"updates" | "extends" | "derives") – line 239 - SpaceRoleEnum (
"owner" | "admin" | "editor" | "viewer") – line 296 - RequestTypeEnum (
"search" | "upsert" | "delete") – line 169
These enums translate to database constraints through Drizzle ORM's type system, ensuring invalid values cannot persist at the database level.
Implementation in the Codebase
The schema definitions in packages/validation/schemas.ts (lines 61‑309) serve as the central source of truth. Database operations materialize these definitions through:
packages/lib/queries.ts: Contains Drizzle-ORM query helpers that map to the schema tables, such as thedeleteDocumentmutation at line 111packages/lib/api.ts: Implements HTTP endpoints like/v3/documentsand/v3/searchthat expose the database schema to clientsapps/web: React hooks inhooks/use-…-mutations.tsfiles consume these APIs, ensuring end-to-end type safety from database to UI
Query Examples
Interacting with the schema uses Drizzle ORM's type-safe query builder:
// Select a document by ID
import { db } from '@/lib/db';
import { documents } from '@/db/schema';
const doc = await db
.select()
.from(documents)
.where(eq(documents.id, 'doc_123'))
.limit(1);
Inserting memory entries demonstrates how Zod validation aligns with database operations:
// Insert a new memory entry
import { memoryEntries } from '@/db/schema';
import { sql } from 'drizzle-orm';
await db.insert(memoryEntries).values({
id: sql`gen_random_uuid()`,
documentId: 'doc_123',
memory: 'User wants to remember the pricing model',
memoryEmbedding: null,
memoryRelations: {}, // Uses MemoryRelationEnum keys
createdAt: new Date(),
});
The Zod schemas ensure that values passed to these queries match the database constraints defined in packages/validation/schemas.ts.
Summary
- Supermemory's database schema is a relational model built on Drizzle ORM with Zod validation, defined in
packages/validation/schemas.ts. - Core entities include Document, Chunk, MemoryEntry, Space, Connection, and OrganizationSettings, plus junction tables for many-to-many relationships.
- Type safety is enforced through eight distinct enums (VisibilityEnum, DocumentTypeEnum, etc.) that constrain column values at both the application and database levels.
- Query patterns in
packages/lib/queries.tsandpackages/lib/api.tsprovide type-safe access to these tables using Drizzle ORM's SQL-like syntax.
Frequently Asked Questions
What ORM does Supermemory use for database operations?
Supermemory uses Drizzle ORM to manage database interactions. The ORM reads table definitions from Zod schemas in packages/validation/schemas.ts and generates type-safe SQL queries, as seen in packages/lib/queries.ts where operations like deleteDocument construct actual SQL statements from the schema definitions.
How does Supermemory handle vector embeddings in the database?
The schema stores vector embeddings as nullable arrays of numbers in the Chunk and MemoryEntry tables. Specifically, the embedding column in Chunk and memoryEmbedding in MemoryEntry use the type z.array(z.number()).nullable(), allowing the application to store high-dimensional vectors for semantic search while handling cases where embeddings have not yet been generated.
Where are the many-to-many relationships defined in the Supermemory schema?
Junction tables for many-to-many relationships are explicitly defined as DocumentsToSpaces and SpacesToMembers in packages/validation/schemas.ts. DocumentsToSpaces links documents to spaces via documentId and spaceId, while SpacesToMembers manages membership through spaceId, userId, and a role field constrained by SpaceRoleEnum.
What database constraints ensure data integrity in Supermemory?
Data integrity is maintained through Zod validation schemas combined with Drizzle ORM's type system. Eight enums—including DocumentStatusEnum, VisibilityEnum, and SpaceRoleEnum—constrain possible values at the application level, while the relational structure enforces referential integrity between documents, chunks, spaces, and users.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →