How HolaOS Handles Real‑Time AI Inference: Architecture, MCP Protocol, and Streaming Pipeline
HolaOS handles real‑time AI inference through a layered architecture that couples the desktop UI, a local runtime bundle, and a Model Context Protocol (MCP) server to stream tokens from multiple providers with persistent usage tracking.
The HolaOS open‑source project implements a provider‑agnostic inference pipeline that keeps the interface responsive while abstracting differences between OpenAI, Anthropic, Kimi, and other model backends. This article breaks down the exact mechanism—from model selection to token streaming—using the actual source code from the holaboss-ai/holaOS repository.
Model Catalog and Selection
Before any inference begins, HolaOS consults a static catalog to resolve model metadata. The catalog in apps/desktop/shared/model-catalog.ts defines every invokable model including model‑id, supported modalities, reasoning flags, and default "thinking" values.
When the UI initiates a request, it calls catalogMetadataForProviderModel to retrieve the appropriate configuration:
import { catalogMetadataForProviderModel } from '@/shared/model-catalog';
const meta = catalogMetadataForProviderModel('holaboss_model_proxy', 'gpt-5.4');
// meta => { label: 'GPT-5.4', reasoning: true, … }
This decouples the UI from provider‑specific details, allowing the same interface to work across heterogeneous backends.
MCP Bridge: The Runtime Abstraction Layer
HolaOS communicates with model providers through a Model Context Protocol (MCP) client embedded in the runtime bundle. The MCP protocol abstracts transport and authentication differences, exposing a uniform JSON‑RPC interface to the desktop process.
The desktop sends inference requests to the runtime, which forwards them to the selected provider's endpoint. Per‑request usage statistics are tracked according to the type definitions in apps/docs/worker-configuration.d.ts:
// Usage statistics for the inference request
interface InferenceUsage {
inputTokens: number;
outputTokens: number;
cachedInput: number;
// … additional provider‑specific fields
}
This runtime layer enables provider‑agnostic real‑time AI inference without modifying the desktop codebase when new providers are added.
Streaming Response Handling
The MCP client establishes a streaming HTTP connection—using SSE or WebSocket when the provider supports it—to minimize latency. Tokens are emitted from the runtime as they arrive and piped back to the renderer through Electron's IPC channel.
The desktop UI consumes this stream via the useInference hook in apps/desktop/src/hooks/useInference.ts:
import { useInference } from '@/hooks/useInference';
function ChatBox() {
const { response, isLoading, start } = useInference();
const send = async (prompt: string) => {
await start({
modelId: 'gpt-5.4',
providerId: 'holaboss_model_proxy',
prompt,
// streaming defaults to true
});
};
return (
<>
<textarea onBlur={e => send(e.target.value)} />
<pre>{isLoading ? '▍' : response}</pre>
</>
);
}
Internally, the MCP client processes the stream and dispatches tokens to the renderer:
export async function runInference(req: InferenceRequest) {
const { modelId, providerId, prompt } = req;
const endpoint = resolveProviderEndpoint(providerId, modelId);
const resp = await fetch(endpoint, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ prompt, stream: true })
});
for await (const chunk of resp.body!.getReader()) {
const token = decodeChunk(chunk);
sendToRenderer('inference-token', token);
}
await stateStore.recordInferenceUsage(req, usageStats);
}
This architecture delivers sub‑second token visibility to users while keeping the main thread unblocked.
State‑Store Bookkeeping for Analytics
Every inference call persists metadata to a SQLite‑backed state store at runtime/state-store/src/store.ts. The store records:
- requestedModel: The model ID specified by the user
- effectiveModel: The actual model served (may differ due to routing or fallbacks)
- Usage counters: input‑tokens, output‑tokens, cached‑input, and derived cost metrics
import { Store } from './store';
await Store.recordInferenceUsage({
requestedModel: 'gpt-5.4',
effectiveModel: 'gpt-5.4',
usage: { inputTokens: 42, outputTokens: 120, cachedInput: 0 },
});
This persistent layer enables the UI to display real‑time cost and performance metrics without external telemetry dependencies.
End‑to‑End Data Flow
The complete real‑time AI inference pipeline in HolaOS follows this sequence:
- UI → User triggers
useInference().start()with model and prompt - Catalog →
catalogMetadataForProviderModelresolves provider configuration - MCP Runtime → Formats JSON‑RPC request and opens streaming connection to provider
- Provider → Returns tokens via SSE/WebSocket as they are generated
- Runtime → Decodes chunks and forwards via Electron IPC to renderer
- UI → React state updates incrementally, rendering tokens as they arrive
- State Store → Usage stats written to SQLite for analytics
This design prioritizes latency, modularity, and observability without sacrificing type safety or cross‑platform compatibility.
Summary
- Model catalog (
model-catalog.ts) provides unified metadata resolution across providers - MCP runtime abstracts transport differences and manages streaming connections
- Electron IPC pipes tokens from runtime to renderer for immediate UI feedback
- SQLite state store persists usage metrics locally for cost tracking and analytics
- Provider‑agnostic architecture allows swapping backends without desktop code changes
Frequently Asked Questions
What protocol does HolaOS use to communicate with AI providers?
HolaOS uses the Model Context Protocol (MCP), a JSON‑RPC style abstraction implemented in the runtime bundle. This protocol standardizes request formatting, authentication, and streaming semantics across OpenAI, Anthropic, Kimi, and other providers.
How does HolaOS achieve low‑latency token streaming?
The MCP client opens streaming HTTP connections using Server‑Sent Events or WebSocket when available. Tokens are decoded and forwarded through Electron's IPC channel immediately upon arrival, eliminating batch‑waiting delays and keeping the UI responsive.
Where does HolaOS store inference usage data?
Usage statistics are persisted to a local SQLite database via the state‑store module at runtime/state-store/src/store.ts. Records include requested and effective model IDs, token counts, and caching metrics—enabling offline cost analysis without external services.
Can HolaOS work with multiple AI providers simultaneously?
Yes. The model catalog maps generic model IDs to provider‑specific endpoints, and the MCP runtime resolves these dynamically. The same useInference hook can target any configured provider by changing the providerId parameter, with no UI code modifications required.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →