How Workspace Document Isolation Works in AnythingLLM: A Technical Deep Dive
Workspace document isolation in AnythingLLM is enforced through three tightly-coupled mechanisms: database-level filtering by workspaceId, vector store namespacing using workspace slugs, and user-workspace relationship checks that prevent unauthorized access.
AnythingLLM implements strict workspace document isolation to ensure that documents, embeddings, and conversations remain segregated between different workspaces and users. This architectural pattern prevents data leakage across organizational boundaries while allowing flexible multi-tenant deployments. The isolation strategy spans the Prisma ORM layer, vector database namespaces, and API-level permission gates.
Database-Level Isolation with Prisma
Workspace-Specific Table Schemas
The foundation of document isolation resides in the database schema design. All workspace-related tables include a mandatory workspaceId column that creates a foreign key relationship to the parent workspace:
workspaces: The parent table containing workspace metadata and slugsworkspace_documents: Stores document metadata withworkspaceIdfilteringworkspace_parsed_files: Tracks parsed file chunks scoped to specific workspacesworkspace_users: Manages the many-to-many relationship between users and workspaces
Every query against these tables explicitly filters by workspaceId, ensuring that data retrieval operations cannot accidentally cross workspace boundaries.
Query Filtering in workspace.js
The server/models/workspace.js file implements the core isolation logic through methods like Workspace.get and Workspace.getWithUser. These functions automatically append workspaceId constraints to every Prisma query:
// From workspace.js#L73-L88
const workspace = await prisma.workspaces.findFirst({
where: {
slug: workspaceSlug,
// Automatically scoped to user's accessible workspaces
workspace_users: {
some: { user_id: user?.id }
}
},
include: {
documents: true, // Returns only workspace_documents with matching workspaceId
workspace_users: true
}
});
The Document.forWorkspace(workspaceId) method in server/models/documents.js similarly enforces isolation by returning only rows where workspaceId matches the provided parameter.
Vector Store Namespace Isolation
Slug-Based Namespacing
While database isolation prevents metadata leakage, AnythingLLM extends this protection to the vector store layer through namespace segregation. Each workspace receives a unique namespace based on its slug—a URL-safe version of the workspace name stored in the workspaces.slug column.
When documents are embedded, the VectorDb.addDocumentToNamespace function receives workspace.slug as the namespace parameter. This creates a physical separation in the vector database (Pinecone, Chroma, LanceDB, etc.) that prevents similarity searches from accessing vectors outside the current workspace context.
Document Embedding Workflow
The Document.addDocuments method in server/models/documents.js (lines 83-110) orchestrates the end-to-end isolation:
// From documents.js#L83-L110
static async addDocuments(workspace, filePaths, userId) {
// 1. Verify user has access to this workspace
const hasAccess = await Workspace.getWithUser(userId, { id: workspace.id });
if (!hasAccess) throw new Error("Unauthorized");
// 2. Process and embed documents into workspace-specific namespace
for (const filePath of filePaths) {
await VectorDb.addDocumentToNamespace(
workspace.slug, // Isolated namespace
filePath,
{ metadata: { workspaceId: workspace.id } }
);
}
// 3. Create database records scoped to workspaceId
await prisma.workspace_documents.createMany({
data: filePaths.map(path => ({
workspaceId: workspace.id, // Foreign key isolation
docpath: path,
userId: userId
}))
});
}
This workflow ensures that vectors and metadata remain synchronized within the workspace boundary throughout the document lifecycle.
Access Control and User Permissions
The workspace_users Relationship Table
The workspace_users table implements a many-to-many relationship that links users to their authorized workspaces. This junction table contains user_id and workspace_id foreign keys, along with role information that determines permission levels.
When the API receives a request, it extracts the user's identity from the authentication token and validates it against this relationship table before executing any workspace-scoped operations.
Role-Based Permission Bypass
The permission system recognizes two privileged roles that bypass standard workspace filtering:
admin: Full system access across all workspacesmanager: Elevated privileges allowing cross-workspace operations
For these roles, the Workspace.getWithUser method skips the workspace_users join and returns the requested workspace directly. Regular users trigger the join condition:
// From workspace.js#L90-L106 (whereWithUser logic)
const whereClause = {
slug: workspaceSlug,
...(user.role === ROLES.admin || user.role === ROLES.manager
? {}
: { workspace_users: { some: { user_id: user.id } } })
};
This multi-layered approach ensures that even if a user somehow obtains a workspace ID or slug, they cannot access documents without explicit membership in the workspace_users table.
Practical Implementation Examples
Creating an Isolated Workspace
When you create a new workspace, AnythingLLM automatically establishes the isolation boundaries:
const { Workspace } = require("./models/workspace");
// Creates a workspace with its own slug/namespace
const { workspace, message } = await Workspace.new(
"Engineering Documentation", // name → slugified for isolation
userId, // creator becomes a member via WorkspaceUser
{ chatProvider: "openai", chatModel: "gpt-4" }
);
This creates a row in workspaces with a unique slug (e.g., engineering-documentation), and the creator is automatically linked via workspace_users.
Adding Documents to a Workspace
Document ingestion automatically routes vectors to the workspace-specific namespace:
const { Document } = require("./models/documents");
const { Workspace } = require("./models/workspace");
// Assume we already have the workspace object from Workspace.get()
await Document.addDocuments(workspace, ["./docs/api-spec.pdf"], userId);
Behind the scenes (documents.js lines 83-110):
VectorDb.addDocumentToNamespace(workspace.slug, …)writes vectors under the namespaceengineering-documentation- The
workspace_documentsrow storesworkspaceIdlinking the file to the workspace
Querying Documents with User Verification
Retrieving documents requires both workspace membership and explicit filtering:
const { Workspace } = require("./models/workspace");
// `req.user` is the authenticated user object
const ws = await Workspace.getWithUser(req.user, { slug: "engineering-documentation" });
console.log(ws.documents); // only documents belonging to that workspace
Workspace.getWithUser (lines 69-78) adds a workspace_users filter for non-admin users, ensuring the returned documents array contains only records whose workspaceId matches the requested workspace.
Vector Search Within Workspace Boundaries
Similarity searches are automatically constrained to the workspace namespace:
const { getVectorDbClass } = require("./utils/helpers");
const VectorDb = getVectorDbClass();
const results = await VectorDb.searchWithinNamespace(
workspace.slug, // namespace = workspace slug
"How do I reset the API key?",
5 // top-k results
);
Only vectors embedded under the same slug (i.e., the same workspace) are examined, preventing cross-workspace leakage during retrieval-augmented generation (RAG) operations.
Summary
Workspace document isolation in AnythingLLM operates through a defense-in-depth strategy that spans the data layer, vector storage, and application logic:
- Database scoping: All tables (
workspace_documents,workspace_parsed_files,workspace_users) include aworkspaceIdcolumn that filters every query through Prisma ORM constraints. - Vector namespace segregation: Each workspace receives a unique slug-based namespace in the vector database, ensuring that similarity searches and embeddings remain physically separated.
- Relationship-based access control: The
workspace_usersjunction table enforces membership validation throughWorkspace.getWithUser, preventing unauthorized access even if workspace identifiers are known.
These mechanisms ensure complete logical isolation where users cannot read, modify, or vector-search documents belonging to workspaces they do not have explicit permission to access.
Frequently Asked Questions
How does AnythingLLM prevent users from accessing documents in other workspaces?
AnythingLLM prevents cross-workspace access through a combination of database foreign key constraints and application-level permission checks. The Workspace.getWithUser method in server/models/workspace.js automatically joins the workspace_users table for non-admin users, ensuring that queries return only workspaces where the user has an explicit membership record. Additionally, all document queries filter by the workspaceId column, creating a database-level guarantee of isolation.
What is the role of workspace slugs in document isolation?
Workspace slugs serve as the isolation boundary in the vector database layer. When documents are embedded, the VectorDb.addDocumentToNamespace function uses the workspace's slug (a URL-safe version of the workspace name) as the namespace identifier. This ensures that vector similarity searches performed via VectorDb.searchWithinNamespace only examine embeddings within that specific namespace, preventing cross-contamination of vector data between workspaces even when using shared vector database infrastructure.
Can administrators bypass workspace isolation boundaries?
Yes, users with admin or manager roles can bypass the standard workspace membership checks. In server/models/workspace.js, the whereWithUser logic checks the user's role before applying the workspace_users filter. If the user has an admin or manager role, the query skips the membership validation and returns the requested workspace directly. This design allows privileged users to manage or audit across all workspaces while maintaining strict isolation for regular users.
How does document isolation affect vector search performance?
Workspace document isolation improves vector search performance through namespace partitioning. By segregating vectors into workspace-specific namespaces based on the workspace slug, the vector database can limit the search scope to a smaller subset of embeddings rather than scanning the entire corpus. The VectorDb.searchWithinNamespace method in the vector store wrapper leverages this partitioning, passing the workspace slug as the namespace parameter to ensure that similarity queries execute against only the relevant workspace's vector collection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →