What AI Models Does moeru-ai/airi Use?
moeru-ai/airi is a multimodal AI platform that orchestrates large language models (LLMs), speech synthesis, transcription, and vision models from providers including OpenAI, Anthropic, ElevenLabs, and local engines like Kokoro, unified behind a single provider abstraction layer.
moeru-ai/airi functions as a multimodal AI companion capable of talking, hearing, seeing, and acting. Rather than hardcoding a single model, the platform implements a dynamic provider architecture that abstracts diverse underlying model families. This design allows the UI to select appropriate AI models based on user configuration, enabling both cloud-based and local inference.
Core Model Categories in moeru-ai/airi
The platform categorizes AI models by capability, integrating specialized models for each modality behind unified interfaces.
Large Language Models (LLMs) for Chat
AIRI leverages conversational LLMs for natural-language interaction. The primary models include OpenAI's gpt-4o and gpt-4o-mini, alongside Anthropic's claude-3, Meta's llama-3, and Google's gemma-2. These models power the core chat interface through the buildOpenAICompatibleProvider function, which standardizes API interactions across different backends.
In packages/stage-ui/src/stores/providers.ts, the OpenAI chat provider is registered via buildOpenAICompatibleProvider({ id: 'openai-audio-speech', … }), while compatible providers like Azure AI Foundry, Ollama, and Perplexity reuse the same builder with different identifiers.
Speech Synthesis and Text-to-Speech (TTS)
AIRI implements multiple TTS strategies for voice output. The platform supports OpenAI's tts-1 and tts-1-hd, ElevenLabs voice models (e.g., elevenlabs-en-us-female), and local inference via Kokoro.
The Kokoro local TTS provider ('kokoro-local') loads the kokoro-82m model using WebGPU acceleration, enabling on-device speech synthesis without network connectivity. This provider is defined in packages/stage-ui/src/stores/providers.ts (lines 1504-1512) and accessed through Web Workers.
Speech Transcription (STT)
For audio input, AIRI integrates OpenAI's whisper-1 and gpt-4o-mini-transcribe models, plus third-party services including Microsoft Speech, Alibaba Cloud, and Deepgram. These models convert user speech into text for processing by the chat LLMs.
Vision and Multimodal Models
Vision capabilities rely on multimodal LLMs such as OpenAI's gpt-4o (vision-enabled) for image captioning and OCR. These models process visual input alongside text, enabling the platform to "see" and describe images.
Embedding Models
The platform utilizes OpenAI's text-embedding-3 series for vector search and semantic retrieval. These embeddings power memory and context retrieval features within the application.
Provider Architecture and Model Discovery
AIRI's flexibility stems from its provider registry pattern. All AI services are registered in packages/stage-ui/src/stores/providers.ts, which defines metadata, capabilities, and factory functions for each provider.
The model-discovery logic resides in packages/stage-ui/src/stores/providers/openai-compatible-builder.ts. This module imports listModels from @xsai/model to dynamically fetch available model catalogs from compatible endpoints:
import { listModels } from '@xsai/model'
const models = await listModels({
baseURL: 'https://api.openai.com/v1/',
apiKey: process.env.OPENAI_API_KEY!
})
console.log(models.map(m => m.id))
This discovery mechanism enables AIRI to populate model selection UIs automatically, supporting both official OpenAI endpoints and compatible services like Ollama or Azure AI Foundry.
Supported AI Providers and Model Families
The platform abstracts the following provider categories:
-
OpenAI (Official): Native integration for
gpt-4o,gpt-4o-mini, TTS, and Whisper models via thecreateOpenAIfactory from@xsai-ext/providers/create. -
OpenAI-Compatible: Support for Azure AI Foundry, Anthropic, Ollama, Perplexity, and generic OpenAI-compatible endpoints. These use the shared
buildOpenAICompatibleProviderwith customized base URLs and authentication. -
Local Inference: Kokoro for offline TTS and Ollama for local LLMs (supporting models like
gpt-oss). -
Specialized Voice Services: ElevenLabs for high-quality voice synthesis, Microsoft Speech, Alibaba Cloud, Volcengine, and Deepgram for alternative STT/TTS backends.
-
Search-Augmented: Perplexity AI integration using models like
pplx-7b-chatfor search-grounded conversations. -
Regional Providers: Minimax models for Chinese-language use cases.
Code Examples: Integrating AI Models
Creating an OpenAI Chat Provider
The following pattern instantiates an OpenAI provider for chat capabilities:
import { createOpenAI } from '@xsai-ext/providers/create'
const openai = createOpenAI({
apiKey: process.env.OPENAI_API_KEY!,
baseUrl: 'https://api.openai.com/v1/'
})
const answer = await openai.chat({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Explain the theory of relativity' }]
})
This corresponds to the provider definition at lines 389-398 in packages/stage-ui/src/stores/providers.ts.
Using Local Kokoro TTS
For offline speech synthesis, AIRI loads Kokoro models via Web Workers:
import { getKokoroWorker, getDefaultKokoroModel } from '@/workers/kokoro'
const worker = await getKokoroWorker()
await worker.loadModel(getDefaultKokoroModel())
const audio = await worker.synthesize('Hello, I am your virtual companion.')
The provider registration for this local TTS engine appears at lines 1504-1512 in packages/stage-ui/src/stores/providers.ts.
Switching to ElevenLabs Voice
For premium voice synthesis, the platform integrates ElevenLabs:
import { createUnElevenLabs } from 'unspeech'
const eleven = createUnElevenLabs('ELEVENLABS_API_KEY')
const voice = await eleven.voice({ voiceId: 'elevenlabs-en-us-female' })
const audio = await voice.speech('Welcome to AIRI!')
This provider is configured in packages/stage-ui/src/stores/providers.ts (lines 883-891).
Key Implementation Files
The model integration layer is implemented across these specific source files:
-
packages/stage-ui/src/stores/providers.ts: Central registry containing all provider metadata, including OpenAI, Kokoro, ElevenLabs, and compatible endpoints. -
packages/stage-ui/src/stores/providers/openai-compatible-builder.ts: Builder utility that wraps OpenAI-compatible APIs and implements thelistModelsdiscovery logic. -
packages/stage-ui/src/libs/providers/providers/anthropic/index.ts: Claude-specific provider implementation for non-OpenAI LLM support. -
packages/stage-ui/src/libs/providers/providers/azure-ai-foundry/index.ts: Azure AI Foundry integration for enterprise OpenAI-compatible deployments. -
packages/stage-ui/src/libs/providers/providers/ollama/index.ts: Local LLM provider supportinggpt-ossdetection and other Ollama-managed models. -
packages/stage-ui/src/workers/kokoro/constants.ts: Definition of Kokoro model configurations, WebGPU requirements, and default model selections.
Summary
- moeru-ai/airi operates as a multimodal platform combining LLMs, speech, vision, and embedding models rather than relying on a single AI model.
- The provider abstraction layer in
packages/stage-ui/src/stores/providers.tsunifies OpenAI, Anthropic, ElevenLabs, Kokoro, and other backends behind consistent interfaces. - Local inference is supported through Kokoro (TTS) and Ollama (LLMs), enabling offline operation.
- Dynamic model discovery via
listModelsfrom@xsai/modelallows the UI to automatically detect available models from compatible endpoints. - The platform supports specialized modalities including speech synthesis (OpenAI TTS, ElevenLabs, Kokoro), transcription (Whisper, Deepgram), and vision (GPT-4o).
Frequently Asked Questions
Does moeru-ai/airi support local AI models without internet connectivity?
Yes. The platform supports fully local inference through the Kokoro provider for text-to-speech (using the kokoro-82m model with WebGPU acceleration) and Ollama for local LLMs. These providers are registered in packages/stage-ui/src/stores/providers.ts and load models directly in the browser or local runtime without requiring cloud API keys.
What speech synthesis models does moeru-ai/airi use?
AIRI integrates five distinct TTS strategies: OpenAI's tts-1 and tts-1-hd models, ElevenLabs voice models (such as elevenlabs-en-us-female), local Kokoro models (kokoro-82m), Microsoft Speech, and Alibaba Cloud/Volcengine services. The platform selects between these based on network availability and user preference settings.
How does moeru-ai/airi discover available models from a provider?
The platform uses the model-discovery logic implemented in openai-compatible-builder.ts, which invokes the listModels helper from @xsai/model. This function queries the provider's /models endpoint (standard in OpenAI-compatible APIs) to retrieve the catalog of available model IDs, enabling dynamic population of the model selection interface without hardcoded lists.
Can I use Claude or other non-OpenAI models with moeru-ai/airi?
Yes. Through the OpenAI-compatible provider architecture, AIRI supports Anthropic's Claude models, Azure AI Foundry deployments, Perplexity AI, and local Ollama instances. The buildOpenAICompatibleProvider function standardizes the interface for these backends, allowing models like claude-3 or llama-3 to power chat functionality through the same UI components as OpenAI models.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →