What AI Models Does holaOS Use? A Complete Technical Guide to the Model Routing System

holaOS uses a flexible model-routing layer that primarily targets OpenAI's GPT-5 series while supporting Anthropic Claude, Google Gemini, Ollama, and custom providers through a unified API resolution system.

The open-source holaOS repository (holaboss-ai/holaOS) implements a provider-agnostic AI orchestration platform capable of routing requests to multiple large language model families. Understanding what AI models holaOS uses requires examining the model-routing architecture in runtime/harnesses/src/model-routing.ts, which normalizes identifiers, selects appropriate API endpoints, and enforces per-model budgets.

Primary OpenAI GPT-5 Models

holaOS treats the GPT-5 series as its primary model family, with built-in detection logic and default configurations optimized for these models.

GPT-5 Pattern Detection

The system identifies OpenAI models using a regex pattern matcher defined in isOpenAiGpt5Model:

function isOpenAiGpt5Model(modelId: string): boolean {
  return /^gpt-5(?:[.-]|$)/.test(modelId);
}

This function resides in runtime/harnesses/src/model-routing.ts and triggers OpenAI-compatible routing logic for any model ID beginning with gpt-5.

Default Model Configuration

The default agent model for new workspaces is hardcoded as gpt-5.4 in runtime/api-server/src/workspace-runtime-plan.ts:

const DEFAULT_AGENT_MODEL = "gpt-5.4";

Supported GPT-5 Variants

The catalog of officially supported OpenAI models is enumerated in runtime/harnesses/src/codex.ts (lines 29-34):

Model ID Display Name Provider
gpt-5.5 GPT-5.5 OpenAI
gpt-5.5-mini GPT-5.5 Mini OpenAI
gpt-5.4 GPT-5.4 OpenAI
gpt-5.4-mini GPT-5.4 Mini OpenAI
gpt-5.4-pro GPT-5.4 Pro OpenAI
gpt-5.4-nano GPT-5.4 Nano OpenAI
gpt-5.3-codex GPT-5.3 Codex OpenAI
gpt-5 GPT-5 OpenAI

Third-Party AI Provider Support

Beyond OpenAI, holaOS integrates with multiple AI providers through provider-specific detection logic in resolveHarnessModelApi.

Anthropic Claude Integration

Models like claude-2 and claude-3-sonnet are detected when model_proxy_provider === "anthropic_native" and route to the anthropic-messages API.

Google Gemini Support

Gemini models (gemini-1.0-pro, gemini-1.5-pro) are routed through the google-generative-ai API when model_proxy_provider === "google_compatible" and provider_id === "gemini_direct".

Local and Open-Source Options

The routing layer supports local inference through Ollama (identified by ollama_* prefixes or models starting with llama, qwen3:, or gpt-oss:) and Qwen models (prefixed with qwen/). These disable specific features like the developer role and store capabilities to maintain compatibility.

Channel Gateway Connectors

Enterprise messaging platforms including Minimax, Dingtalk, WeChat, Discord, Slack, Feishu, and QQ are handled through dedicated connectors in runtime/channel-gateway/src/connectors/*.

The Model Routing Implementation

The core routing decision occurs in resolveHarnessModelApi (lines 300-312 of model-routing.ts):

export function resolveHarnessModelApi(request: HarnessModelRoutingRequest): HarnessModelApi {
  const normalizedProvider = request.model_client.model_proxy_provider.trim().toLowerCase();
  if (normalizedProvider === "anthropic_native") return "anthropic-messages";
  if (shouldUseNativeGoogleProvider(request)) return "google-generative-ai";
  if (shouldUseOpenAiResponsesProvider(request)) return "openai-responses";
  return "openai-completions";
}

This function normalizes the provider string and selects the appropriate API harness based on the request configuration.

Budget and Token Limit Configuration

holaOS enforces per-model constraints through the knownModelBudgetOverride function. For example, GPT-5.4 models receive specific token allocations:

case "gpt-5.4":
case "gpt-5.4-pro":
  return { contextWindow: 1_050_000, maxTokens: 128_000 };

These definitions appear in runtime/harnesses/src/model-routing.ts (lines 78-85) and ensure that API requests respect model-specific context windows and output limits.

Practical Implementation Examples

Routing a GPT-5.4 Request

To build a routing request for the default GPT-5.4 model:

import { resolveHarnessModelProfile } from "./runtime/harnesses/src/model-routing";

const request = {
  provider_id: "openai_direct",
  model_id: "gpt-5.4",
  thinking_value: "medium",
  model_client: {
    model_proxy_provider: "openai_compatible",
    api_key: process.env.OPENAI_API_KEY!,
    base_url: "https://api.openai.com/v1",
  },
};

const profile = resolveHarnessModelProfile(request, {
  modelCatalog: {},
  fallbackBudget: undefined,
});

console.log(profile);

The returned profile includes the API type (openai-responses), base URL, reasoning capabilities, and budget constraints (contextWindow: 1_050_000, maxTokens: 128_000).

Configuring Anthropic Claude

For Anthropic models, change the provider configuration:

const anthropicRequest = {
  provider_id: "anthropic_direct",
  model_id: "claude-3-sonnet-20240229",
  thinking_value: null,
  model_client: {
    model_proxy_provider: "anthropic_native",
    api_key: process.env.ANTHROPIC_API_KEY!,
    base_url: "https://api.anthropic.com",
  },
};

const profile = resolveHarnessModelProfile(anthropicRequest, { modelCatalog: {} });
console.log(profile.api); // "anthropic-messages"

Custom Model Catalog Override

Runtime configuration loading allows custom budget definitions via the HOLABOSS_RUNTIME_CONFIG_PATH environment variable pointing to a JSON file:

{
  "models": {
    "openai/gpt-5.4": {
      "context_window": 1200000,
      "max_tokens": 128000,
      "cost": { "input": 0.0005, "output": 0.0015, "cacheRead": 0, "cacheWrite": 0 }
    }
  }
}

The router reads these values through runtimeConfigModelCatalog() (lines 66-84 of model-routing.ts), enabling custom pricing and token limits without code changes.

Summary

  • holaOS primarily uses OpenAI's GPT-5 series (gpt-5.4 as default, plus gpt-5.5, mini, pro, and nano variants) as defined in runtime/harnesses/src/codex.ts.
  • Multi-provider architecture supports Anthropic Claude, Google Gemini, Ollama, and Qwen through the resolveHarnessModelApi router in runtime/harnesses/src/model-routing.ts.
  • Pattern-based detection uses regex (/^gpt-5(?:[.-]|$)/) to identify OpenAI models and provider strings for third-party services.
  • Configurable budgets enforce context windows (up to 1,050,000 tokens for GPT-5.4) and costs through knownModelBudgetOverride.
  • Runtime customization allows new models and providers via environment-specific configuration files without modifying core routing logic.

Frequently Asked Questions

What is the default AI model in holaOS?

The default agent model is gpt-5.4, hardcoded in runtime/api-server/src/workspace-runtime-plan.ts. This default applies to new workspaces unless explicitly overridden by user configuration or runtime environment variables.

Can holaOS run local AI models?

Yes. holaOS supports Ollama for local inference, detecting models with ollama_* prefixes or specific local identifiers like llama, qwen3:, and gpt-oss:. The routing layer automatically disables unsupported features (such as the developer role and store persistence) when targeting local endpoints.

How does holaOS handle API authentication for different providers?

Authentication is handled per-request through the model_client object, which includes api_key and base_url fields. The resolveHarnessModelApi function normalizes the model_proxy_provider string to determine which API harness (OpenAI, Anthropic, or Google) to invoke, passing the corresponding credentials to the selected endpoint.

Where are model budgets and token limits configured?

Per-model constraints are defined in the knownModelBudgetOverride function within runtime/harnesses/src/model-routing.ts (lines 78-85). Additionally, operators can specify custom budgets by setting HOLABOSS_RUNTIME_CONFIG_PATH to a JSON file containing model-specific context_window, max_tokens, and cost parameters, which the system loads via runtimeConfigModelCatalog().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →