How the AI Coding Assistant Works in holaOS: A Deep Dive into the pi-coding-agent Architecture

The AI coding assistant in holaOS is built on the pi-coding-agent framework and processes natural-language requests through a four-layer pipeline: UI channel gateway, runtime API server, harness-host orchestrator, and model harness (Codex or Claude-Code).

This article explains the internals of holaOS's AI coding assistant, which transforms user prompts into executable code using a stateless, tool-augmented architecture. We'll trace the exact flow from user input to final output, referencing the actual source files in the holaboss-ai/holaOS repository.

Architecture Overview: Four Layers of the AI Coding Assistant

The holaOS AI coding assistant operates through four cooperating layers. Each layer has distinct responsibilities and communicates via well-defined interfaces.

Layer Location Core Responsibility
UI → Channel Gateway runtime/channel-gateway/src/manager.ts Formats and routes messages between desktop UI and runtime server
Runtime API Server runtime/api-server/src/app.ts Accepts HTTP requests and initializes Pi sessions
Harness-Host (Pi) runtime/harness-host/src/pi.ts Orchestrates workspace loading, tool discovery, and model spawning
Model Harness runtime/harnesses/src/codex.ts Runs the actual LLM CLI and handles bidirectional streaming

The design emphasizes stateless execution per run while persisting turn-level data for contextual continuity across sessions.

Step 1: Request Initiation via the Channel Gateway

When a user submits a coding request in the holaOS desktop app, the journey begins at the channel gateway.

The gateway in runtime/channel-gateway/src/manager.ts receives a JSON-RPC run_started event from the UI. It formats this message for the specific chat client in use—whether Slack, Discord, Telegram, or others—then forwards the payload to the runtime API server.

// Conceptual flow through the channel gateway
// Source: runtime/channel-gateway/src/manager.ts
const gatewayEvent = {
  jsonrpc: "2.0",
  method: "run_started",
  params: {
    workspace_id: "my-workspace",
    session_id: "sess-123",
    instruction: "Write a React component that fetches data"
  }
};

This abstraction allows holaOS to support multiple client interfaces without modifying downstream components.

Step 2: Runtime API Server Creates the Pi Session

The runtime API server at runtime/api-server/src/app.ts exposes the primary HTTP interface for the AI coding assistant. It accepts POST /api/v1/run requests with a HarnessHostPiRequest body.

Upon receiving a request, the server:

  1. Validates the payload
  2. Creates a Pi session via runtime/harness-host/src/pi.ts
  3. Delegates execution to the harness-host
import { request } from "packages/runtime-client/src/request.ts";

await request({
  method: "POST",
  url: "/api/v1/run",
  body: {
    workspace_id: "my-workspace",
    session_id: "sess-123",
    instruction: "Create a TypeScript function that validates an email address",
    model_id: "gpt-5.5",
    provider_id: "openai_codex",
    tools: {},           // Host exposes default coding tools
    attachments: [],     // Optional files for model context
  },
});

The API server acts as a thin orchestration layer—heavy lifting happens in the harness-host.

Step 3: Harness-Host Orchestrates the Coding Run

The harness-host in runtime/harness-host/src/pi.ts is the core brain of the AI coding assistant. It performs four critical operations before spawning any model:

Workspace Skill Loading

loadHarnessWorkspaceSkills() discovers and loads user-defined skills from the workspace configuration. These skills extend the assistant's capabilities with domain-specific knowledge.

MCP Tool Discovery

discoverHarnessMcpTools() identifies available Model Context Protocol (MCP) tools. MCP is an open standard for connecting LLMs to external systems, and holaOS uses it to expose file systems, databases, APIs, and other resources.

Runtime Configuration Building

buildHarnessHostRequest() assembles the complete runtime config including:

  • model_id: Which LLM to use (e.g., gpt-5.5)
  • tools: Available function definitions
  • attachments: Files the model can reference
  • MCP servers: External tool endpoints

Model Harness Selection and Spawning

Based on the provider_id, the host selects a harness implementation. By default, it uses Codex via runtime/harnesses/src/codex.ts, though Claude-Code is also supported.

// In runtime/harness-host/src/pi.ts — spawning the model harness
const harness = await bindHarnessHostPlugin("codex");
const childProcess = spawn("codex", [
  "app-server",
  "--listen", "stdio://"
], {
  stdio: ["pipe", "pipe", "pipe"]
});

The host communicates with this child process over bidirectional stdio using JSON-RPC 2.0.

Step 4: Model Harness Streams and Tool Execution

Once spawned, the model harness runs the actual LLM CLI. For Codex, this is codex app-server --listen stdio:// as defined in runtime/harnesses/src/codex.ts.

Event Streaming

The harness streams JSON-RPC events back to the host:

  • assistant text deltas → forwarded to UI for typing animation
  • tool_use requests → host resolves and executes
  • run_completed → finalization signal

Tool Schema Sanitization

Before tools reach the model, sanitizeToolSchemas() processes them in runtime/harness-host/src/pi.ts. This function:

  • Removes unsupported JSON-Schema keywords
  • Normalizes union types
  • Logs schema modifications for debugging

Tool Call Resolution

When the assistant emits a tool_use event:

// In runtime/harness-host/src/pi.ts
const toolDef = resolveToolByName(toolName);
const result = await runTool(toolDef, toolArgs, {
  // Tool implementations from @earendil-works/pi-coding-agent
  ...createCodingTools()
});
await sendToolResult({
  tool_call_id: toolCallId,
  result,
});

The host proxies MCP servers, runs the actual tool implementation, and streams results back. This cycle repeats until the assistant signals completion.

State Persistence with the State-Store

Though each run is stateless, holaOS persists conversational data via runtime/state-store/src/store.ts. This enables:

  • Cross-session context retrieval
  • Audit trails for assistant actions
  • Debugging and replay capabilities
import { getMemoryEntries } from "runtime/state-store/src/store.ts";

const entries = await getMemoryEntries({
  workspaceId: "my-workspace",
  sessionId: "sess-123",
  sourceTypes: ["assistant_turn"],
});

console.log(entries.map(e => e.assistantText).join("\n"));

The store records assistant_turn entries containing concatenated assistant replies, tool results, and attachment references.

Complete Request Flow Diagram


┌─────────────┐     run_started (JSON-RPC)      ┌─────────────────┐
│ holaOS UI   │ ──────────────────────────────→ │ Channel Gateway │
│ (Desktop)   │                                   │  (manager.ts)   │
└─────────────┘                                   └────────┬────────┘
                                                           │
                              POST /api/v1/run            ↓
┌─────────────┘                                   ┌─────────────────┐
│                                               │  API Server     │
│                                               │   (app.ts)      │
│                                               └────────┬────────┘
│                                                        │
│                              Pi session + harness     ↓
│                                               ┌─────────────────┐
│                                               │  Harness-Host   │
│                                               │    (pi.ts)      │
│                                               │  • load skills  │
│                                               │  • discover MCP │
│                                               │  • spawn model  │
│                                               └────────┬────────┘
│                                                        │
│                              stdio JSON-RPC           ↓
│                                               ┌─────────────────┐
└──────────────────────────────────────────────→│  Model Harness  │
                                                │  (codex.ts)     │
                                                │  ↔ LLM CLI      │
                                                └─────────────────┘

Supported Model Providers

The AI coding assistant in holaOS supports multiple backends through its harness abstraction:

Provider Harness File CLI Command
OpenAI Codex runtime/harnesses/src/codex.ts codex app-server --listen stdio://
Claude-Code Equivalent harness claude-code with stdio interface

The harness interface is designed for extension—new providers implement the same JSON-RPC streaming contract.

Summary

  • The holaOS AI coding assistant uses a four-layer pipelined architecture: channel gateway, API server, harness-host, and model harness.
  • Stateless execution per run ensures reliability, while the state-store enables contextual continuity.
  • MCP tool discovery and schema sanitization happen in runtime/harness-host/src/pi.ts before any model interaction.
  • The Codex harness at runtime/harnesses/src/codex.ts spawns the LLM as a child process with bidirectional stdio streaming.
  • All turn-level data persists to runtime/state-store/src/store.ts for retrieval and debugging.

Frequently Asked Questions

What is pi-coding-agent in holaOS?

pi-coding-agent is the underlying framework that powers holaOS's AI coding assistant. It provides the orchestration layer for workspace skills, tool discovery, and model harness management. The framework is implemented primarily in runtime/harness-host/src/pi.ts and integrates with OpenAI's Codex and Anthropic's Claude-Code.

How does holaOS handle tool calls from the AI assistant?

Tool calls flow through a sanitization and resolution pipeline. When the model emits a tool_use event, the harness-host in pi.ts resolves the tool by name, sanitizes its schema to remove unsupported JSON-Schema keywords, executes the implementation from createCodingTools(), and streams the result back via JSON-RPC. This happens transparently to the user while the UI shows a typing indicator.

Can holaOS work with different LLM providers?

Yes, through its pluggable harness system. The default implementation uses OpenAI Codex via runtime/harnesses/src/codex.ts, but the architecture supports alternative providers like Claude-Code through equivalent harness implementations. Each harness must implement the same JSON-RPC 2.0 streaming interface over stdio.

Where is conversation history stored in holaOS?

Turn-level data persists to the state-store at runtime/state-store/src/store.ts. This includes assistant text, tool results, and attachment references keyed by workspace and session ID. While individual runs are stateless, the store enables later sessions to query prior context using getMemoryEntries() with filters for sourceTypes: ["assistant_turn"].

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →