OpenHuman Model Routing: Provider Selection, STT Engine Factory, and BYOK Subscription Blending
OpenHuman uses a trait-based provider abstraction in Rust to dynamically route LLM and STT requests through configurable backends, blending built-in credentials with user-supplied BYOK keys via an encrypted keyring.
OpenHuman separates model selection from service implementation through a robust routing layer written in Rust. This architecture enables dynamic provider selection for both large language model (LLM) inference and speech-to-text (STT) processing while seamlessly integrating Bring-Your-Own-Key (BYOK) subscriptions with default service tiers. Understanding this model routing mechanism is essential for developers configuring self-hosted or hybrid AI deployments in the tinyhumansai/openhuman repository.
Provider Abstraction and Selection Logic
The core routing system relies on a Provider trait that abstracts any remote or local model service.
The Provider Trait and Factory Registration
In src/openhuman/inference/provider.rs, the Provider trait defines the interface for all model services including OpenAI-compatible APIs, Anthropic, Azure, and Ollama. Concrete provider implementations register with a factory, and the Provider::select function examines three critical inputs:
- The request's
modelfield - Any explicit
provideroverrides - The user's subscription tier
This selection logic determines which backend handles each inference request.
Harness Builder and Runtime Invocation
The Harness builder in src/openhuman/harness.rs receives configuration for the desired provider including API key, base URL, and model name. It attaches this configuration to the core context during initialization.
When executing a turn, the CoreRuntime::invoke path calls Provider::invoke, which forwards the request to the selected provider implementation. This indirection ensures the application code remains agnostic to the specific backend handling the request.
Cost-Aware Provider Blending
OpenHuman supports complex routing scenarios where requests distribute across multiple providers based on subscription credits.
Fallback and Weighted Distribution
If a request omits an explicit provider, the router falls back to the default provider defined in the user’s config (config.rs). For subscription tiers that include multiple provider credits—such as blended plans mixing OpenAI and Anthropic quotas—the system uses cost-weighting logic in src/openhuman/inference/blend.rs to split request batches across services optimally.
STT Engine Factory Pattern
The voice subsystem implements the same abstraction pattern through a dedicated factory for speech-to-text providers.
Factory Entry Point and Configuration
The voice::factory::effective_stt_provider function in src/openhuman/voice/factory.rs reads the stt_engine configuration value. Valid options include:
"backend"– Uses the backend service defined by user configuration"elevenlabs"– Routes to ElevenLabs hosted service"openai"– Routes to OpenAI Whisper API"cloud"– Default option after local Whisper engine removal
Provider Construction via Trait Objects
Each configuration option maps to a concrete implementation satisfying the SttProvider trait defined in src/openhuman/voice/providers/*.rs. The factory returns a boxed trait object, allowing the voice pipeline to interact with any STT backend through a uniform interface.
When a user maintains a BYOK subscription for a third-party STT service, the factory merges the custom endpoint and API key from the subscription record into the provider configuration before returning the instance.
BYOK Subscription Integration and Security
OpenHuman’s subscription model, defined in the web3 domain, enables users to attach custom credentials to any supported service while maintaining security through encrypted storage.
Secure Credential Storage
User-provided keys and endpoint URLs reside in the encrypted keyring (src/openhuman/security/keyring.rs). The Subscription struct in src/openhuman/web3/subscription.rs records the service name, credential identifier, and usage limits. This separation ensures sensitive data never exists in plain configuration files.
Dynamic Injection at Runtime
When instantiating any provider—whether for LLM inference, STT processing, or X-402 payment gateways—the factory queries keyring::lookup_subscription to check for matching BYOK entries. If found, the factory injects the custom URL and API key into the provider’s configuration before construction.
Because every provider implements the same trait interface (Provider for LLMs or SttProvider for speech), the core code operates identically regardless of credential source. The routing layer automatically blends BYOK-enabled services with default providers, allowing a single conversation turn to utilize a custom endpoint for one call while falling back to standard services for others.
Practical Implementation Examples
The following examples demonstrate constructing providers with BYOK subscription support:
// Example: building a Harness that prefers a BYOK‑enabled provider
let harness = Harness::builder()
.provider(Provider::openai_compatible(
"https://api.openai.com/v1", // default URL
std::env::var("OPENAI_API_KEY")?,
).model("gpt-4o"))
// The keyring contains a BYOK entry for a cheaper Anthropic endpoint
.subscription_keyring(keyring::load()?)
.build()
.await?;
// The router will automatically select the Anthropic BYOK endpoint
// for calls that match the subscription’s “model = claude‑2.1” rule.
let response = harness.run("Summarise the repository.").await?;
// Example: selecting an STT engine with a custom BYOK subscription
let stt = voice::factory::effective_stt_provider(&config, &keyring)?;
let transcript = stt.transcribe(audio_bytes).await?;
Summary
- Provider abstraction in
src/openhuman/inference/provider.rsenables uniform access to diverse LLM backends through theProvidertrait andProvider::selectlogic. - Cost-aware blending in
src/openhuman/inference/blend.rsdistributes requests across multiple providers based on subscription credits and weighting rules. - STT factory pattern in
src/openhuman/voice/factory.rsinstantiates speech-to-text providers viavoice::factory::effective_stt_provider, supporting both hosted services and BYOK configurations. - BYOK integration leverages the encrypted keyring (
src/openhuman/security/keyring.rs) and subscription records (src/openhuman/web3/subscription.rs) to inject custom credentials at runtime without code changes. - Unified API surface ensures that application logic remains identical whether using built-in credentials or user-supplied keys.
Frequently Asked Questions
How does OpenHuman decide which LLM provider to use for a specific request?
The routing layer examines the request's model field, any explicit provider override parameters, and the user's subscription tier through Provider::select in src/openhuman/inference/provider.rs. If no specific provider is requested, the system falls back to the default configured in config.rs or selects a BYOK-enabled provider if the model matches a subscription rule.
Can I use my own API keys for STT services in OpenHuman?
Yes. OpenHuman supports BYOK (Bring-Your-Own-Key) subscriptions for STT providers through the factory in src/openhuman/voice/factory.rs. When you configure a custom subscription with your API key and endpoint stored in the encrypted keyring, voice::factory::effective_stt_provider automatically merges these credentials into the provider configuration, allowing seamless use of paid services like ElevenLabs or OpenAI Whisper without modifying application code.
What happens when a subscription includes credits for multiple providers?
OpenHuman utilizes the blending logic in src/openhuman/inference/blend.rs to distribute requests across providers based on cost-weighting algorithms. The router can split batches of requests between services—such as allocating traffic across OpenAI and Anthropic—according to the subscription's credit allocation, maximizing cost efficiency while maintaining API compatibility through the unified Provider trait interface.
Where are BYOK credentials stored securely in OpenHuman?
BYOK credentials are encrypted and stored in the keyring implementation at src/openhuman/security/keyring.rs. The Subscription struct in src/openhuman/web3/subscription.rs maintains references to these credentials including service names and usage limits. During provider instantiation, factories query this keyring via keyring::lookup_subscription to inject credentials dynamically, ensuring sensitive keys never reside in plain text configuration files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →