What AI Models Does holaOS Use? A Complete Technical Guide to the Model Routing System
holaOS uses a flexible model-routing layer that primarily targets OpenAI's GPT-5 series while supporting Anthropic Claude, Google Gemini, Ollama, and custom providers through a unified API resolution system.
The open-source holaOS repository (holaboss-ai/holaOS) implements a provider-agnostic AI orchestration platform capable of routing requests to multiple large language model families. Understanding what AI models holaOS uses requires examining the model-routing architecture in runtime/harnesses/src/model-routing.ts, which normalizes identifiers, selects appropriate API endpoints, and enforces per-model budgets.
Primary OpenAI GPT-5 Models
holaOS treats the GPT-5 series as its primary model family, with built-in detection logic and default configurations optimized for these models.
GPT-5 Pattern Detection
The system identifies OpenAI models using a regex pattern matcher defined in isOpenAiGpt5Model:
function isOpenAiGpt5Model(modelId: string): boolean {
return /^gpt-5(?:[.-]|$)/.test(modelId);
}
This function resides in runtime/harnesses/src/model-routing.ts and triggers OpenAI-compatible routing logic for any model ID beginning with gpt-5.
Default Model Configuration
The default agent model for new workspaces is hardcoded as gpt-5.4 in runtime/api-server/src/workspace-runtime-plan.ts:
const DEFAULT_AGENT_MODEL = "gpt-5.4";
Supported GPT-5 Variants
The catalog of officially supported OpenAI models is enumerated in runtime/harnesses/src/codex.ts (lines 29-34):
| Model ID | Display Name | Provider |
|---|---|---|
gpt-5.5 |
GPT-5.5 | OpenAI |
gpt-5.5-mini |
GPT-5.5 Mini | OpenAI |
gpt-5.4 |
GPT-5.4 | OpenAI |
gpt-5.4-mini |
GPT-5.4 Mini | OpenAI |
gpt-5.4-pro |
GPT-5.4 Pro | OpenAI |
gpt-5.4-nano |
GPT-5.4 Nano | OpenAI |
gpt-5.3-codex |
GPT-5.3 Codex | OpenAI |
gpt-5 |
GPT-5 | OpenAI |
Third-Party AI Provider Support
Beyond OpenAI, holaOS integrates with multiple AI providers through provider-specific detection logic in resolveHarnessModelApi.
Anthropic Claude Integration
Models like claude-2 and claude-3-sonnet are detected when model_proxy_provider === "anthropic_native" and route to the anthropic-messages API.
Google Gemini Support
Gemini models (gemini-1.0-pro, gemini-1.5-pro) are routed through the google-generative-ai API when model_proxy_provider === "google_compatible" and provider_id === "gemini_direct".
Local and Open-Source Options
The routing layer supports local inference through Ollama (identified by ollama_* prefixes or models starting with llama, qwen3:, or gpt-oss:) and Qwen models (prefixed with qwen/). These disable specific features like the developer role and store capabilities to maintain compatibility.
Channel Gateway Connectors
Enterprise messaging platforms including Minimax, Dingtalk, WeChat, Discord, Slack, Feishu, and QQ are handled through dedicated connectors in runtime/channel-gateway/src/connectors/*.
The Model Routing Implementation
The core routing decision occurs in resolveHarnessModelApi (lines 300-312 of model-routing.ts):
export function resolveHarnessModelApi(request: HarnessModelRoutingRequest): HarnessModelApi {
const normalizedProvider = request.model_client.model_proxy_provider.trim().toLowerCase();
if (normalizedProvider === "anthropic_native") return "anthropic-messages";
if (shouldUseNativeGoogleProvider(request)) return "google-generative-ai";
if (shouldUseOpenAiResponsesProvider(request)) return "openai-responses";
return "openai-completions";
}
This function normalizes the provider string and selects the appropriate API harness based on the request configuration.
Budget and Token Limit Configuration
holaOS enforces per-model constraints through the knownModelBudgetOverride function. For example, GPT-5.4 models receive specific token allocations:
case "gpt-5.4":
case "gpt-5.4-pro":
return { contextWindow: 1_050_000, maxTokens: 128_000 };
These definitions appear in runtime/harnesses/src/model-routing.ts (lines 78-85) and ensure that API requests respect model-specific context windows and output limits.
Practical Implementation Examples
Routing a GPT-5.4 Request
To build a routing request for the default GPT-5.4 model:
import { resolveHarnessModelProfile } from "./runtime/harnesses/src/model-routing";
const request = {
provider_id: "openai_direct",
model_id: "gpt-5.4",
thinking_value: "medium",
model_client: {
model_proxy_provider: "openai_compatible",
api_key: process.env.OPENAI_API_KEY!,
base_url: "https://api.openai.com/v1",
},
};
const profile = resolveHarnessModelProfile(request, {
modelCatalog: {},
fallbackBudget: undefined,
});
console.log(profile);
The returned profile includes the API type (openai-responses), base URL, reasoning capabilities, and budget constraints (contextWindow: 1_050_000, maxTokens: 128_000).
Configuring Anthropic Claude
For Anthropic models, change the provider configuration:
const anthropicRequest = {
provider_id: "anthropic_direct",
model_id: "claude-3-sonnet-20240229",
thinking_value: null,
model_client: {
model_proxy_provider: "anthropic_native",
api_key: process.env.ANTHROPIC_API_KEY!,
base_url: "https://api.anthropic.com",
},
};
const profile = resolveHarnessModelProfile(anthropicRequest, { modelCatalog: {} });
console.log(profile.api); // "anthropic-messages"
Custom Model Catalog Override
Runtime configuration loading allows custom budget definitions via the HOLABOSS_RUNTIME_CONFIG_PATH environment variable pointing to a JSON file:
{
"models": {
"openai/gpt-5.4": {
"context_window": 1200000,
"max_tokens": 128000,
"cost": { "input": 0.0005, "output": 0.0015, "cacheRead": 0, "cacheWrite": 0 }
}
}
}
The router reads these values through runtimeConfigModelCatalog() (lines 66-84 of model-routing.ts), enabling custom pricing and token limits without code changes.
Summary
- holaOS primarily uses OpenAI's GPT-5 series (gpt-5.4 as default, plus gpt-5.5, mini, pro, and nano variants) as defined in
runtime/harnesses/src/codex.ts. - Multi-provider architecture supports Anthropic Claude, Google Gemini, Ollama, and Qwen through the
resolveHarnessModelApirouter inruntime/harnesses/src/model-routing.ts. - Pattern-based detection uses regex (
/^gpt-5(?:[.-]|$)/) to identify OpenAI models and provider strings for third-party services. - Configurable budgets enforce context windows (up to 1,050,000 tokens for GPT-5.4) and costs through
knownModelBudgetOverride. - Runtime customization allows new models and providers via environment-specific configuration files without modifying core routing logic.
Frequently Asked Questions
What is the default AI model in holaOS?
The default agent model is gpt-5.4, hardcoded in runtime/api-server/src/workspace-runtime-plan.ts. This default applies to new workspaces unless explicitly overridden by user configuration or runtime environment variables.
Can holaOS run local AI models?
Yes. holaOS supports Ollama for local inference, detecting models with ollama_* prefixes or specific local identifiers like llama, qwen3:, and gpt-oss:. The routing layer automatically disables unsupported features (such as the developer role and store persistence) when targeting local endpoints.
How does holaOS handle API authentication for different providers?
Authentication is handled per-request through the model_client object, which includes api_key and base_url fields. The resolveHarnessModelApi function normalizes the model_proxy_provider string to determine which API harness (OpenAI, Anthropic, or Google) to invoke, passing the corresponding credentials to the selected endpoint.
Where are model budgets and token limits configured?
Per-model constraints are defined in the knownModelBudgetOverride function within runtime/harnesses/src/model-routing.ts (lines 78-85). Additionally, operators can specify custom budgets by setting HOLABOSS_RUNTIME_CONFIG_PATH to a JSON file containing model-specific context_window, max_tokens, and cost parameters, which the system loads via runtimeConfigModelCatalog().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →