How OpenHuman's STT Provider Factory and Dictation Server Handle Voice and Speech
OpenHuman separates voice processing into a factory-based provider system for speech-to-text/text-to-speech and an always-on dictation server that captures audio, detects voice activity, and routes transcribed commands through a wake-word gate.
The tinyhumansai/openhuman repository implements a modular voice architecture that decouples audio capture from provider selection. This article explains how the STT Provider Factory and Dictation Server work together to process speech, transcribe audio, and route commands to the agent.
Provider Factory Architecture
The voice subsystem centers on a factory pattern implemented in src/openhuman/voice/factory/mod.rs. This module interprets provider strings—compact configuration tokens like "cloud", "piper", or "deepgram:nova-2"—and instantiates the appropriate speech-to-text (STT) or text-to-speech (TTS) backend.
The factory exposes its public API through a series of constructors and helpers:
pub use entry::{
create_stt_provider, create_tts_provider,
default_stt_provider, default_tts_provider,
DEFAULT_PIPER_VOICE, DEFAULT_STT_MODEL,
};
pub use helpers::{effective_stt_provider, effective_tts_provider};
pub use traits::{SttProvider, SttResult, TtsProvider};
All factory operations log with the [voice-factory] prefix, while concrete providers log under [voice-stt] or [voice-tts]. The factory delegates API key resolution to lookup_key_for_slug in src/openhuman/voice/factory/helpers.rs, which retrieves credentials from the global configuration.
Building STT and TTS Providers
The factory generates boxed trait objects through two primary entry points:
create_stt_provider(name, api_key, config)– Returns either aCloudSttProvider(routing to the OpenHuman backend proxy) or anExternalSttProvider(direct third-party API calls). Notably, the factory has no local STT branch; local Whisper inference was removed, with local processing delegated to thevoice_server.stt_engineconfiguration.create_tts_provider– Follows a similar pattern but includes a specialPiperTtsProviderfor local subprocess-based synthesis.
Provider implementations live in src/openhuman/voice/factory/stt_providers.rs and src/openhuman/voice/factory/tts_providers.rs, while the trait definitions reside in src/openhuman/voice/factory/traits.rs. Each provider must implement the transcribe method for STT or the speak method for TTS, accepting configuration objects and audio data.
Dictation Server Pipeline
The always-on listening pipeline is implemented in src/openhuman/voice/always_on.rs. This module maintains a continuous microphone stream and processes speech through several distinct stages:
Audio Capture and Preprocessing
The spawn_capture_thread function creates a cpal input stream that forwards raw interleaved audio chunks to an async processor via an MPSC channel. The system tracks dropped chunks (DROPPED_CHUNKS) and logs them on power-of-two intervals to monitor buffer health.
Voice Activity Detection
The processor task lazily opens a VAD session using tinyvoice::VadSession::open, retrying on failure with SESSION_RETRY_INTERVAL. Audio undergoes downmixing and resampling via tinyvoice::prepare_frames before entering the VAD module, which emits SpeechStart and SpeechEnd events. The system buffers utterances up to MAX_UTTERANCE_SAMPLES (approximately 60 seconds).
Transcription and Wake-Word Gating
When speech ends, the system invokes transcribe_and_deliver, which:
- Encodes the 16 kHz PCM buffer to WAV format using
tinyvoice::encode_wav - Resolves the effective STT provider via
effective_stt_provider(config) - Instantiates the provider through
create_stt_provider - Sends base64-encoded audio to the provider's
transcribemethod
The resulting transcript passes through tinyvoice::extract_command to check against the configured config.voice_server.wake_word. If the wake word is detected, the command string flows to deliver_command; otherwise, the utterance is discarded or triggers an acknowledgement turn ("hello").
Fast-Path Intent Routing
The deliver_command function first attempts lightweight local routing through modules::voice::route for intents like Pause or VolumeUp. If no fast-path intent matches, the command falls back to the standard LLM pipeline via dictation_listener::publish_transcription.
Configuration and Provider String Grammar
The factory parses provider strings according to a specific grammar documented in the source. Users configure providers via the stt_provider and tts_provider config fields.
The resolution logic handles several string forms:
"cloud"or"openhuman"– Resolves to the OpenHuman backend proxy (STT routes to/openai/v1/audio/transcriptions, TTS uses ElevenLabs)"piper"– Local Piper subprocess for TTS only"<slug>:<model>"– Third-party STT APIs such asdeepgram:nova-2oropenai:whisper-1"<slug>"– Bare slug using the provider's default model or voice
API keys are resolved dynamically through lookup_key_for_slug in src/openhuman/voice/factory/helpers.rs, allowing secure credential management without hardcoding.
Implementation Examples
Creating an STT Provider Manually
use openhuman::voice::factory::{create_stt_provider, effective_stt_provider};
use openhuman::config::Config;
// Load the application configuration
let cfg: Config = load_config();
// Resolve the provider name from config (e.g., "openai" or "deepgram:nova-2")
let provider_name = effective_stt_provider(&cfg);
// Instantiate the provider
let stt = create_stt_provider(&provider_name, "", &cfg)
.expect("STT provider must be available");
// Transcribe base64-encoded WAV audio
let wav_base64 = "..."; // base64-encoded 16kHz PCM
let result = stt.transcribe(&cfg, &wav_base64, Some("audio/wav"), None, Some("en")).await?;
println!("Transcribed text: {}", result.value.text);
Starting the Always-On Dictation Server
use openhuman::voice::always_on::start_if_enabled;
use openhuman::config::Config;
let cfg = Config::load_or_init().await.unwrap();
// Launch the microphone pipeline if enabled in config
tokio::spawn(async move {
start_if_enabled(&cfg).await;
});
Handling Fast-Path Intents
use openhuman::voice::ops::deliver_command;
use openhuman::config::Config;
let command = "pause".to_string();
let cfg = Config::load_or_init().await.unwrap();
// Routes to local intent handler or falls back to LLM
deliver_command(&cfg, command).await;
Summary
- The STT Provider Factory in
src/openhuman/voice/factory/mod.rsinstantiates speech providers based on configuration strings like"cloud"or"deepgram:nova-2". - Provider resolution occurs dynamically via
effective_stt_providerandeffective_tts_provider, with API keys fetched securely from the global config. - The Dictation Server in
src/openhuman/voice/always_on.rsmaintains a continuouscpalaudio stream, processes voice activity throughtinyvoice, and transcribes utterances up to 60 seconds. - A wake-word gate filters transcripts before routing; matched commands flow through fast-path intent handlers (
modules::voice::route) or the standard LLM pipeline. - The architecture contains no local STT inference in the factory; local processing requires external configuration via
voice_server.stt_engine.
Frequently Asked Questions
What provider strings does OpenHuman support for STT?
OpenHuman supports "cloud" (OpenHuman backend proxy), "openai:whisper-1", "deepgram:nova-2", and other third-party slugs following the "<slug>:<model>" pattern. Bare slugs like "openai" use provider defaults. The factory resolves these in src/openhuman/voice/factory/mod.rs and maps them to CloudSttProvider or ExternalSttProvider implementations.
How does the Dictation Server handle voice activity detection?
The server uses the tinyvoice crate to run VadSession analysis on audio captured via cpal. The processor task in src/openhuman/voice/always_on.rs prepares frames through tinyvoice::prepare_frames and buffers speech between SpeechStart and SpeechEnd events, enforcing a maximum utterance length of MAX_UTTERANCE_SAMPLES (approximately 60 seconds).
Can OpenHuman run STT locally without cloud providers?
The factory itself has no local STT branch; it only provides CloudSttProvider and ExternalSttProvider. Local Whisper inference was removed from the factory codebase. Local STT processing requires configuring voice_server.stt_engine separately, which operates outside the factory pattern described in src/openhuman/voice/factory/.
How are transcribed commands routed after wake-word detection?
After transcription, tinyvoice::extract_command checks for the configured wake_word. If matched, deliver_command attempts fast-path routing through modules::voice::route for local intents (e.g., volume control). Unmatched commands are published to the dictation bus via dictation_listener::publish_transcription for LLM processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →