How holaOS Handles Data Processing and Management: A Deep Dive into the SQLite-First Architecture

holaOS stores all user-generated data in versioned SQLite databases managed by the runtime-state-store package, with workspace-isolated files, typed migrations, and vector-powered semantic search.

The open-source holaOS project from Holaboss AI takes a deliberate database-per-workspace approach to data processing and management. Rather than relying on external services, the entire platform persists workspace files, interaction logs, semantic embeddings, and integration graphs in local SQLite databases with strict schema versioning. This article explores the three-layer storage architecture, the migration engine that enables safe evolution, and how the system surfaces data through typed APIs and channel gateways.


The Three-Layer Storage Architecture

holaOS organizes data across three logical layers, each optimized for different access patterns and lifecycle requirements.

Root State: Global Host Configuration

The host-state.db serves as the system-wide registry. Created on startup, it holds:

  • Globally installed plugins and capabilities
  • A folded monolith of legacy workspace rows consolidated from earlier versions
  • The HOST_STATE_MONOLITH_FOLDED_MARKER_KEY constant (line 42 of runtime/state-store/src/store.ts) that prevents re-processing deleted workspaces

This root database is intentionally minimal—most data lives in workspace-scoped files.

Workspace Runtime: Isolated Per-Workspace Databases

Each workspace maintains its own SQLite file at .holaboss/runtime.db. The path is constructed by workspaceRuntimeDbPathForWorkspacePath (line 1463) using constants:

WORKSPACE_RUNTIME_DIRNAME = ".holaboss"
WORKSPACE_RUNTIME_DB_FILENAME = "runtime.db"

These databases contain:

Data Type Stored In
Session history and turn results sessions, turn_results tables
Semantic memory entries with embeddings semantic_memory_search_docs, semantic_memory_nodes
Integration knowledge graphs integration_trees, integration_leaves, integration_node_embeddings
Workspace artifacts and metadata Various artifact tables

Workspace identity is protected by a lock file mechanism (workspace_id.lock, line 27) with WORKSPACE_IDENTITY_LOCK_RETRY_ATTEMPTS (line 32) retries. Failures surface as RuntimeStateStoreWorkspaceError types (lines 88-95) for precise debugging.

Control-Plane: Optional Cross-Workspace Analytics

When HOLABOSS_CONTROL_PLANE_DB_PATH is set, holaOS builds a read-only control-plane.db aggregating integration trees and other cross-workspace data. Migration 030-integration-tables-control-plane-only.ts isolates these tables to prevent root database bloat while enabling global discovery queries.


Database Engine and Migration System

SQLite with Extensions

All databases open via better-sqlite3 (line 1, store.ts) with optional sqlite-vec (line 8) for vector operations. This combination provides:

  • Synchronous, performant SQLite access
  • Native vector similarity search without external services
  • Portable, file-based storage suitable for local-first deployments

Versioned Schema Migrations

The MigrationRunner (line 10) executes sequential migrations from runtime/state-store/src/migrations/. Each migration exports an up function that:

  1. Modifies schema (CREATE TABLE, ALTER COLUMN, etc.)
  2. Backfills existing data when necessary
  3. Logs MigrationLogEvents for audit trails

The runner continues on best-effort failure (line 16484) so partially migrated databases still start—a pragmatic choice for development environments.

// Example migration structure (033-session-org-id.ts pattern)
export async function up(db: Database) {
  await db.exec(`
    ALTER TABLE sessions ADD COLUMN org_id TEXT;
    UPDATE sessions SET org_id = 'default' WHERE org_id IS NULL;
  `);
}

Runtime Data Access via Typed APIs

The RuntimeStateStore interface (@holaboss/runtime-state-store) abstracts all database operations. Key methods consumed by the API server include:

  • listSemanticMemorySearchDocs() – retrieves vector-indexed documents
  • searchSemanticMemorySearchDocs(query) – executes hybrid vector+lexical search
  • getSemanticMemoryNode(nodeId) – fetches specific semantic entities
  • listIntegrationTrees(workspaceId) – reads integration knowledge graphs

The API server module runtime/api-server/src/workspace-memory.ts demonstrates typical usage:

import type {
  IntegrationTreeRecord,
  InteractionEntityRecord,
  RuntimeStateStore,
  SemanticMemoryCategory,
} from "@holaboss/runtime-state-store";

export async function getWorkspaceMemory(
  store: RuntimeStateStore,
  workspaceId: string,
) {
  const trees = await store.listIntegrationTrees(workspaceId);
  const interactions = await store.listInteractionEntities(workspaceId);
  // Combine, rank, and return hybrid retrieval result
}

This abstraction ensures callers never write raw SQL while enabling the store to optimize queries internally.


Semantic Memory and Hybrid Retrieval

holaOS implements per-workspace semantic search using sqlite-vec embeddings. The retrieval pipeline in workspace-memory.ts operates in four stages:

  1. Tokenization (tokenize, line 86) normalizes the query
  2. Vector first pass fetches candidates via queryMemoryModelEmbedding
  3. Lexical fallback (DIRECT_LEXICAL_SCAN_*) supplements when vector coverage is insufficient
  4. Result combination via buildMemoryHybridRetrievalResult (line 17 import)

Tunable constants control this behavior:

VECTOR_FIRST_PASS_LIMIT_FLOOR = 20
VECTOR_FIRST_PASS_LIMIT_CEILING = 100
LEXICAL_SUPPORT_LIMIT_FLOOR = 10
LEXICAL_SUPPORT_LIMIT_CEILING = 50
SEMANTIC_MEMORY_PLANNER_STATS_MARKER = 29  // Bump to force ANALYZE

The WorkspaceMemoryExecutionProfile (lines 11-14) lets callers disable embeddings or LLM reranking for latency-sensitive paths.


Integration Knowledge Graph

External service integrations are modeled as trees and leaves in dedicated tables:

  • integration_trees – root nodes per integration instance
  • integration_leaves – actionable endpoints or resources
  • integration_node_embeddings – vector representations for semantic discovery

These tables are excluded from root consolidation via INTEGRATION_GRAPH_ROOT_SKIP_TABLES (line 81) because they can be rebuilt from workspace configuration. Migration 030 optionally migrates them to the control-plane database for cross-workspace analytics without duplicating into every workspace file.


Channel Gateway: Formatting Data for External Services

The channel-gateway package (runtime/channel-gateway) transforms internal semantic nodes into platform-specific messages. Each connector implements a format module:

Telegram HTML formatter (runtime/channel-gateway/src/format/telegram-html.ts):

export function formatTelegramHtml(node: WorkspaceSemanticNode): string {
  return `<b>${escapeHtml(node.title)}</b>\n${escapeHtml(node.body)}`;
}

WeChat connector (runtime/channel-gateway/src/connectors/wechat.ts) handles authentication, rate limiting, and message egress to WeChat servers.

This format-aware gateway ensures semantic memory retrieved via the store can reach users on any supported platform with appropriate rendering.


Complete Data Flow Example

import { RuntimeStateStore } from "@holaboss/runtime-state-store";
import { getWorkspaceMemory } from "runtime/api-server/src/workspace-memory.js";
import { formatTelegramHtml } from "runtime/channel-gateway/src/format/telegram-html.js";

// 1. Initialize store (creates DBs, runs pending migrations)
const store = new RuntimeStateStore();

// 2. Retrieve hybrid memory results for a workspace
const memory = await getWorkspaceMemory(store, "proj-2024-q3");

// 3. Format top result for Telegram delivery
const message = formatTelegramHtml(memory.hits[0].semanticNode);

// 4. Egress via platform connector (connector.send omitted)

All symbols trace to source files in runtime/state-store/src/store.ts, runtime/api-server/src/workspace-memory.ts, and runtime/channel-gateway/src/format/telegram-html.ts.


Summary

  • holaOS uses SQLite as the primary data store, with separate files for host state, each workspace, and optional control-plane analytics
  • Versioned migrations in runtime/state-store/src/migrations/ guarantee backward-compatible schema evolution
  • Workspace isolation through file-per-workspace databases protects data boundaries and enables portability
  • Semantic search combines sqlite-vec embeddings with lexical fallback for hybrid retrieval
  • Typed APIs via RuntimeStateStore abstract all data access, consumed by API server and channel gateway modules
  • Channel formatters transform internal data for Telegram, Slack, Discord, WeChat, and WeCom platforms

Frequently Asked Questions

What database does holaOS use for data storage?

holaOS uses SQLite for all persistent storage, managed through the better-sqlite3 library with optional sqlite-vec extension for vector search. The system creates three types of database files: a global host-state.db, individual runtime.db files per workspace, and an optional control-plane.db for cross-workspace analytics.

How does holaOS handle schema changes and database migrations?

The MigrationRunner class sequentially applies TypeScript migration files from runtime/state-store/src/migrations/. Each migration contains an up function that modifies schema and backfills data. The runner logs all operations and continues on partial failure, allowing the system to start even with incomplete migrations.

Where is semantic search data stored in holaOS?

Semantic embeddings and searchable documents live in per-workspace SQLite databases under .holaboss/runtime.db, specifically in tables like semantic_memory_search_docs and semantic_memory_nodes. The sqlite-vec extension enables native vector similarity search without external vector databases.

How does holaOS prevent data corruption during concurrent access?

Each workspace uses a file-based lock (workspace_id.lock) with configurable retry attempts (WORKSPACE_IDENTITY_LOCK_RETRY_ATTEMPTS). The RuntimeStateStore wraps errors in RuntimeStateStoreWorkspaceError types to provide specific failure diagnostics when lock acquisition fails.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →