How Documents Are Organized into Collections Within the Knowledge Base in 5ire

Documents are organized into collections within the knowledge base through a hierarchical three-tier structure where files reference collections via foreign keys and are subdivided into chunks for vector search.

The 5ire application (nanbingxyz/5ire) implements a relational document management system where documents are organized into collections within the knowledge base. This architecture enables efficient storage, retrieval, and semantic search across imported files through a strict Collection → Files → Chunks hierarchy maintained in SQLite.

The Collection-Document Hierarchy

The knowledge base structures data into three distinct levels to maintain logical organization and support vector search operations.

Collections as Logical Containers

A collection serves as the top-level organizational unit, identified by id, name, and optional metadata such as memo and favorite flags. The TypeScript interface definition in [src/types/knowledge.d.ts](https://github.com/nanbingxyz/5ire/blob/main/src/types/knowledge.d.ts#L1-L10) establishes the contract for these containers, ensuring type safety across the application.

File-to-Collection Relationships

Each document or file belongs to exactly one collection through a foreign-key constraint. In [src/main/database/schema/tables.ts](https://github.com/nanbingxyz/5ire/blob/main/src/main/database/schema/tables.ts#L158-L171), the knowledge_files table (or the newer documents table) includes a collectionId column that references knowledge_collections.id, enforcing referential integrity at the database level.

To enable similarity search, files are split into chunks stored in the document_chunks table. Each chunk maintains a documentId foreign key linking it back to its parent file. This creates a complete chain: Collection → Files → Chunks. The structure allows the system to retrieve specific text segments across all documents inside a collection.

Database Schema Implementation

The SQLite schema in src/main/database/schema/tables.ts defines the physical storage layer for this hierarchy. The schema creates three primary tables that enforce how documents are organized into collections within the knowledge base:

  • knowledge_collections: Stores collection metadata
  • knowledge_files: Stores file data with collectionId as a foreign key
  • document_chunks: Stores vector-searchable segments with documentId as a foreign key

This relational design supports cascading operations—deleting a collection automatically removes associated files and chunks—while enabling efficient queries that join collections with their document counts.

Managing Collections Programmatically

The [src/stores/useKnowledgeStore.ts](https://github.com/nanbingxyz/5ire/blob/main/src/stores/useKnowledgeStore.ts) file implements the API for interacting with this hierarchy. The store provides type-safe methods to create collections, import files, and query relationships.

Create a new collection using the store's createCollection method (lines 64-77):

await useKnowledgeStore.getState().createCollection({
  name: 'Project Docs',
  memo: 'Technical specs',
});
// → INSERT into `knowledge_collections`

Add a document to an existing collection by specifying the collectionId (lines 72-84):

await useKnowledgeStore.getState().createFile({
  collectionId: '<collection-id>',
  name: 'API_guide.pdf',
  size: 102400,
  numOfChunks: 12,
});
// → INSERT into `knowledge_files`

Retrieve all collections with their associated file counts using listCollections (lines 58-62):

const collections = await useKnowledgeStore.getState().listCollections();
// → SELECT with LEFT JOIN on `knowledge_files`

Fetch all files belonging to a specific collection:

const files = await useKnowledgeStore.getState().listFiles('<collection-id>');
// → SELECT from `knowledge_files` where `collectionId` = ?

Summary

  • Collections act as logical containers with metadata, defined in src/types/knowledge.d.ts.
  • Documents are organized into collections via the collectionId foreign key in the knowledge_files table.
  • Chunks enable vector search by linking to parent documents through documentId in document_chunks.
  • The three-tier hierarchy (Collection → Files → Chunks) supports cascading deletes and efficient retrieval.
  • The useKnowledgeStore API provides type-safe methods for managing these relationships according to the schema in src/main/database/schema/tables.ts.

Frequently Asked Questions

What is the relationship between collections and documents in 5ire?

Documents are organized into collections within the knowledge base using a one-to-many relationship. Each document row in the knowledge_files table stores a collectionId foreign key that references exactly one collection, while a single collection can contain multiple documents.

How does the database enforce document organization?

The SQLite schema in src/main/database/schema/tables.ts enforces organization through foreign key constraints. The collectionId column in the knowledge_files table ensures every document belongs to a valid collection, preventing orphaned records and enabling cascading deletion of documents when their parent collection is removed.

What happens to documents when a collection is deleted?

When a collection is deleted, the relational database schema automatically removes all associated documents from knowledge_files and their corresponding chunks from document_chunks through cascading delete constraints. This maintains data integrity and ensures no orphaned files remain in the knowledge base.

How are document chunks linked to collections?

Document chunks are linked indirectly to collections through their parent documents. Each chunk in the document_chunks table contains a documentId foreign key referencing knowledge_files.id, which in turn references knowledge_collections.id via collectionId. This chain enables vector similarity searches scoped to specific collections.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →