Understanding User Prefixes in y-gui's Data Storage System
User prefixes in y-gui's data storage system provide multi-tenant isolation by namespace-partitioning all user data within shared R2 buckets and D1 databases using a unique identifier derived from the user's email.
y-gui is a multi-tenant SaaS platform where numerous users store integrations, chat histories, bot configurations, and other persistent data. To maintain strict data isolation without provisioning separate storage resources for each tenant, the system implements a user prefix strategy that prepends a unique identifier to every database record and object storage key.
How User Prefixes Enable Multi-Tenant Isolation
The prefix architecture serves as a logical namespace that prevents data collision across tenants while enabling efficient bulk operations.
Namespace Isolation in Object Storage
All objects stored in Cloudflare R2 follow the key pattern {user_prefix}/{resource_path}. For example, an integration configuration for a user with email alice@example.com is stored at:
e99a18c428cb38d5f260853678922e03_alice_at_example_dot_com/integration_config.jsonl
This guarantees that objects belonging to different users never clash, even when stored within a single shared bucket.
Tenant Separation in D1 Databases
Every table in the D1 SQL database—including integration, chat, bot, and mcp_server—includes a mandatory user_prefix column. As implemented in backend/src/repository/d1/integration-d1-repository.ts (lines 14-22), queries always filter by this column:
// backend/src/repository/d1/integration-d1-repository.ts
async getIntegrationsByPrefix(prefix: string): Promise<Integration[]> {
const result = await this.db
.prepare('SELECT * FROM integration WHERE user_prefix = ?')
.bind(prefix)
.all();
return result.results as Integration[];
}
This schema design enforces tenant separation at the storage layer, making joins and lookups straightforward while preventing cross-tenant data leakage.
Generating Unique User Prefixes
The system generates deterministic prefixes from user email addresses using the calculateUserPrefix function in backend/src/utils/user.ts (lines 14-28):
// backend/src/utils/user.ts
export async function calculateUserPrefix(email?: string): Promise<string> {
if (!email) return '';
const md5Hash = await calculateMd5(email);
const replacedEmail = email.replace(/@/g, '_at_').replace(/\./g, '_dot_');
// → md5(email)_user_at_example_dot_com
return `${md5Hash}_${replacedEmail}`;
}
The algorithm combines an MD5 hash of the email (providing a fixed-length, collision-resistant component) with a sanitized version of the email address (replacing @ with _at_ and . with _dot_). This produces a filesystem-safe string that uniquely identifies the tenant while remaining human-readable for debugging purposes.
Enumerating User Prefixes for Bulk Operations
To perform operations across all tenants—such as refreshing OAuth tokens—the system must enumerate all existing prefixes. The listUserPrefixes function in backend/src/utils/storage.ts (lines 13-38) implements this by scanning the R2 bucket:
// backend/src/utils/storage.ts
export async function listUserPrefixes(r2Bucket: R2Bucket): Promise<string[]> {
const prefixes = new Set<string>();
const objects = await r2Bucket.list();
for (const object of objects.objects) {
const parts = object.key.split('/');
if (parts.length > 1) {
const prefix = parts.slice(0, -1).join('/');
if (prefix) prefixes.add(prefix);
}
}
return Array.from(prefixes);
}
This function extracts the directory portion of every object key (everything before the final /), aggregates unique values into a Set, and returns them as an array. As utilized in backend/src/index.ts, this enables the system to iterate over every tenant during startup to refresh tokens without scanning unrelated data.
Practical Implementation Examples
Storing Integration Configurations
When a logged-in user saves an integration, the system calculates their prefix and constructs the R2 key:
import { calculateUserPrefix } from '../utils/user';
import { r2Bucket } from '../services/r2';
async function saveIntegrationConfig(email: string, config: object) {
const prefix = await calculateUserPrefix(email);
const key = `${prefix}/integration_config.jsonl`;
await r2Bucket.put(key, JSON.stringify(config));
}
The resulting object key includes the MD5 hash and sanitized email, ensuring unique placement within the shared bucket.
Querying User-Specific Data
Database queries always filter by the user_prefix column to enforce tenant isolation:
import { calculateUserPrefix } from '../utils/user';
import { chatRepository } from '../repository/d1/chat-d1-repository';
async function getUserChats(email: string) {
const prefix = await calculateUserPrefix(email);
return chatRepository.getAllChatsForUserPrefix(prefix);
}
This translates to a SQL WHERE user_prefix = ? clause, guaranteeing only that user's chat records are returned.
Refreshing Tokens Across All Tenants
For administrative tasks affecting every user, the system enumerates prefixes and processes each tenant:
import { listUserPrefixes } from '../utils/storage';
import { tokenRefresh } from '../utils/token-refresh';
async function refreshAllTokens(r2: R2Bucket) {
const prefixes = await listUserPrefixes(r2);
for (const prefix of prefixes) {
await tokenRefresh.refreshForUserPrefix(prefix, r2);
}
}
This pattern enables bulk operations without requiring a separate user registry or external database.
Summary
- User prefixes provide namespace isolation for multi-tenant data storage in y-gui's shared R2 buckets and D1 databases.
- Prefix generation combines MD5 hashing with character sanitization to create unique, filesystem-safe identifiers from user emails.
- Storage patterns require every R2 key and SQL record to include the
user_prefix, enforcing tenant separation at the data layer. - Bulk operations leverage
listUserPrefixesto enumerate tenants for system-wide tasks like token refresh without scanning unrelated data.
Frequently Asked Questions
How does y-gui handle data storage for users who are not logged in?
When no user is authenticated, calculateUserPrefix returns an empty string, causing the system to use the default (empty) prefix. This allows unauthenticated use while maintaining the same storage architecture, with data stored at the root level of the bucket or with NULL/empty prefix values in the database.
Why does y-gui use MD5 hashing for the user prefix instead of a UUID?
The implementation uses MD5 hashing to generate a fixed-length, deterministic identifier from the user's email. This approach ensures that the same email always produces the same prefix (enabling idempotent lookups) while providing a collision-resistant component that remains consistent across sessions. The hash is combined with a sanitized version of the email to maintain human readability for debugging purposes.
Can administrators query data across all users in the y-gui system?
Yes, administrators can perform cross-tenant operations by utilizing the listUserPrefixes function, which scans the R2 bucket to discover all unique prefixes. As implemented in backend/src/index.ts, this enables bulk operations such as refreshing OAuth tokens for every user without requiring a separate user registry or external database query.
What database tables include the user_prefix column?
According to the source code analysis, every table in the D1 database schema includes a user_prefix column to enforce tenant isolation. This includes tables for integration, chat, bot, and mcp_server, among others. The consistent schema design ensures that all SQL queries can filter by prefix to return only the requesting user's data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →