# What AI Models Does moeru-ai/airi Use?

> Discover the AI models powering moeru-ai/airi. Explore its multimodal capabilities leveraging LLMs, speech synthesis, and vision from OpenAI, Anthropic, ElevenLabs, and local engines.

- Repository: [Moeru AI/airi](https://github.com/moeru-ai/airi)
- Tags: internals
- Published: 2026-03-08

---

**moeru-ai/airi is a multimodal AI platform that orchestrates large language models (LLMs), speech synthesis, transcription, and vision models from providers including OpenAI, Anthropic, ElevenLabs, and local engines like Kokoro, unified behind a single provider abstraction layer.**

moeru-ai/airi functions as a **multimodal AI companion** capable of talking, hearing, seeing, and acting. Rather than hardcoding a single model, the platform implements a dynamic provider architecture that abstracts diverse underlying model families. This design allows the UI to select appropriate AI models based on user configuration, enabling both cloud-based and local inference.

## Core Model Categories in moeru-ai/airi

The platform categorizes AI models by capability, integrating specialized models for each modality behind unified interfaces.

### Large Language Models (LLMs) for Chat

AIRI leverages conversational LLMs for natural-language interaction. The primary models include **OpenAI's `gpt-4o` and `gpt-4o-mini`**, alongside **Anthropic's `claude-3`**, **Meta's `llama-3`**, and **Google's `gemma-2`**. These models power the core chat interface through the `buildOpenAICompatibleProvider` function, which standardizes API interactions across different backends.

In [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts), the OpenAI chat provider is registered via `buildOpenAICompatibleProvider({ id: 'openai-audio-speech', … })`, while compatible providers like Azure AI Foundry, Ollama, and Perplexity reuse the same builder with different identifiers.

### Speech Synthesis and Text-to-Speech (TTS)

AIRI implements multiple TTS strategies for voice output. The platform supports **OpenAI's `tts-1` and `tts-1-hd`**, **ElevenLabs voice models** (e.g., `elevenlabs-en-us-female`), and **local inference via Kokoro**.

The **Kokoro local TTS** provider (`'kokoro-local'`) loads the `kokoro-82m` model using WebGPU acceleration, enabling on-device speech synthesis without network connectivity. This provider is defined in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts) (lines 1504-1512) and accessed through Web Workers.

### Speech Transcription (STT)

For audio input, AIRI integrates **OpenAI's `whisper-1` and `gpt-4o-mini-transcribe`** models, plus third-party services including Microsoft Speech, Alibaba Cloud, and Deepgram. These models convert user speech into text for processing by the chat LLMs.

### Vision and Multimodal Models

Vision capabilities rely on multimodal LLMs such as **OpenAI's `gpt-4o` (vision-enabled)** for image captioning and OCR. These models process visual input alongside text, enabling the platform to "see" and describe images.

### Embedding Models

The platform utilizes **OpenAI's `text-embedding-3`** series for vector search and semantic retrieval. These embeddings power memory and context retrieval features within the application.

## Provider Architecture and Model Discovery

AIRI's flexibility stems from its provider registry pattern. All AI services are registered in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts), which defines metadata, capabilities, and factory functions for each provider.

The **model-discovery logic** resides in [`packages/stage-ui/src/stores/providers/openai-compatible-builder.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers/openai-compatible-builder.ts). This module imports `listModels` from `@xsai/model` to dynamically fetch available model catalogs from compatible endpoints:

```typescript
import { listModels } from '@xsai/model'

const models = await listModels({ 
  baseURL: 'https://api.openai.com/v1/', 
  apiKey: process.env.OPENAI_API_KEY! 
})
console.log(models.map(m => m.id))

```

This discovery mechanism enables AIRI to populate model selection UIs automatically, supporting both official OpenAI endpoints and compatible services like Ollama or Azure AI Foundry.

## Supported AI Providers and Model Families

The platform abstracts the following provider categories:

- **OpenAI (Official)**: Native integration for `gpt-4o`, `gpt-4o-mini`, TTS, and Whisper models via the `createOpenAI` factory from `@xsai-ext/providers/create`.

- **OpenAI-Compatible**: Support for **Azure AI Foundry**, **Anthropic**, **Ollama**, **Perplexity**, and generic OpenAI-compatible endpoints. These use the shared `buildOpenAICompatibleProvider` with customized base URLs and authentication.

- **Local Inference**: **Kokoro** for offline TTS and **Ollama** for local LLMs (supporting models like `gpt-oss`).

- **Specialized Voice Services**: **ElevenLabs** for high-quality voice synthesis, **Microsoft Speech**, **Alibaba Cloud**, **Volcengine**, and **Deepgram** for alternative STT/TTS backends.

- **Search-Augmented**: **Perplexity AI** integration using models like `pplx-7b-chat` for search-grounded conversations.

- **Regional Providers**: **Minimax** models for Chinese-language use cases.

## Code Examples: Integrating AI Models

### Creating an OpenAI Chat Provider

The following pattern instantiates an OpenAI provider for chat capabilities:

```typescript
import { createOpenAI } from '@xsai-ext/providers/create'

const openai = createOpenAI({
  apiKey: process.env.OPENAI_API_KEY!,
  baseUrl: 'https://api.openai.com/v1/'
})

const answer = await openai.chat({
  model: 'gpt-4o',
  messages: [{ role: 'user', content: 'Explain the theory of relativity' }]
})

```

This corresponds to the provider definition at lines 389-398 in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts).

### Using Local Kokoro TTS

For offline speech synthesis, AIRI loads Kokoro models via Web Workers:

```typescript
import { getKokoroWorker, getDefaultKokoroModel } from '@/workers/kokoro'

const worker = await getKokoroWorker()
await worker.loadModel(getDefaultKokoroModel())
const audio = await worker.synthesize('Hello, I am your virtual companion.')

```

The provider registration for this local TTS engine appears at lines 1504-1512 in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts).

### Switching to ElevenLabs Voice

For premium voice synthesis, the platform integrates ElevenLabs:

```typescript
import { createUnElevenLabs } from 'unspeech'

const eleven = createUnElevenLabs('ELEVENLABS_API_KEY')
const voice = await eleven.voice({ voiceId: 'elevenlabs-en-us-female' })
const audio = await voice.speech('Welcome to AIRI!')

```

This provider is configured in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts) (lines 883-891).

## Key Implementation Files

The model integration layer is implemented across these specific source files:

- **[`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts)**: Central registry containing all provider metadata, including OpenAI, Kokoro, ElevenLabs, and compatible endpoints.

- **[`packages/stage-ui/src/stores/providers/openai-compatible-builder.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers/openai-compatible-builder.ts)**: Builder utility that wraps OpenAI-compatible APIs and implements the `listModels` discovery logic.

- **[`packages/stage-ui/src/libs/providers/providers/anthropic/index.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/libs/providers/providers/anthropic/index.ts)**: Claude-specific provider implementation for non-OpenAI LLM support.

- **[`packages/stage-ui/src/libs/providers/providers/azure-ai-foundry/index.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/libs/providers/providers/azure-ai-foundry/index.ts)**: Azure AI Foundry integration for enterprise OpenAI-compatible deployments.

- **[`packages/stage-ui/src/libs/providers/providers/ollama/index.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/libs/providers/providers/ollama/index.ts)**: Local LLM provider supporting `gpt-oss` detection and other Ollama-managed models.

- **[`packages/stage-ui/src/workers/kokoro/constants.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/workers/kokoro/constants.ts)**: Definition of Kokoro model configurations, WebGPU requirements, and default model selections.

## Summary

- moeru-ai/airi operates as a **multimodal platform** combining LLMs, speech, vision, and embedding models rather than relying on a single AI model.
- The **provider abstraction layer** in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts) unifies OpenAI, Anthropic, ElevenLabs, Kokoro, and other backends behind consistent interfaces.
- **Local inference** is supported through Kokoro (TTS) and Ollama (LLMs), enabling offline operation.
- **Dynamic model discovery** via `listModels` from `@xsai/model` allows the UI to automatically detect available models from compatible endpoints.
- The platform supports **specialized modalities** including speech synthesis (OpenAI TTS, ElevenLabs, Kokoro), transcription (Whisper, Deepgram), and vision (GPT-4o).

## Frequently Asked Questions

### Does moeru-ai/airi support local AI models without internet connectivity?

Yes. The platform supports **fully local inference** through the Kokoro provider for text-to-speech (using the `kokoro-82m` model with WebGPU acceleration) and Ollama for local LLMs. These providers are registered in [`packages/stage-ui/src/stores/providers.ts`](https://github.com/moeru-ai/airi/blob/main/packages/stage-ui/src/stores/providers.ts) and load models directly in the browser or local runtime without requiring cloud API keys.

### What speech synthesis models does moeru-ai/airi use?

AIRI integrates **five distinct TTS strategies**: OpenAI's `tts-1` and `tts-1-hd` models, ElevenLabs voice models (such as `elevenlabs-en-us-female`), local **Kokoro** models (`kokoro-82m`), Microsoft Speech, and Alibaba Cloud/Volcengine services. The platform selects between these based on network availability and user preference settings.

### How does moeru-ai/airi discover available models from a provider?

The platform uses the **model-discovery logic** implemented in [`openai-compatible-builder.ts`](https://github.com/moeru-ai/airi/blob/main/openai-compatible-builder.ts), which invokes the `listModels` helper from `@xsai/model`. This function queries the provider's `/models` endpoint (standard in OpenAI-compatible APIs) to retrieve the catalog of available model IDs, enabling dynamic population of the model selection interface without hardcoded lists.

### Can I use Claude or other non-OpenAI models with moeru-ai/airi?

Yes. Through the **OpenAI-compatible provider architecture**, AIRI supports Anthropic's Claude models, Azure AI Foundry deployments, Perplexity AI, and local Ollama instances. The `buildOpenAICompatibleProvider` function standardizes the interface for these backends, allowing models like `claude-3` or `llama-3` to power chat functionality through the same UI components as OpenAI models.