# How OpenHuman's STT Provider Factory and Dictation Server Handle Voice and Speech

> Discover how OpenHuman's STT Provider Factory and Dictation Server process voice and speech. Learn about audio capture, voice activity detection, and transcribed command routing.

- Repository: [Tiny Humans/openhuman](https://github.com/tinyhumansai/openhuman)
- Tags: internals
- Published: 2026-08-29

---

**OpenHuman separates voice processing into a factory-based provider system for speech-to-text/text-to-speech and an always-on dictation server that captures audio, detects voice activity, and routes transcribed commands through a wake-word gate.**

The `tinyhumansai/openhuman` repository implements a modular voice architecture that decouples audio capture from provider selection. This article explains how the **STT Provider Factory** and **Dictation Server** work together to process speech, transcribe audio, and route commands to the agent.

## Provider Factory Architecture

The voice subsystem centers on a factory pattern implemented in [`src/openhuman/voice/factory/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/mod.rs). This module interprets **provider strings**—compact configuration tokens like `"cloud"`, `"piper"`, or `"deepgram:nova-2"`—and instantiates the appropriate speech-to-text (STT) or text-to-speech (TTS) backend.

The factory exposes its public API through a series of constructors and helpers:

```rust
pub use entry::{
    create_stt_provider, create_tts_provider,
    default_stt_provider, default_tts_provider,
    DEFAULT_PIPER_VOICE, DEFAULT_STT_MODEL,
};
pub use helpers::{effective_stt_provider, effective_tts_provider};
pub use traits::{SttProvider, SttResult, TtsProvider};

```

All factory operations log with the `[voice-factory]` prefix, while concrete providers log under `[voice-stt]` or `[voice-tts]`. The factory delegates API key resolution to `lookup_key_for_slug` in [`src/openhuman/voice/factory/helpers.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/helpers.rs), which retrieves credentials from the global configuration.

## Building STT and TTS Providers

The factory generates boxed trait objects through two primary entry points:

- **`create_stt_provider(name, api_key, config)`** – Returns either a `CloudSttProvider` (routing to the OpenHuman backend proxy) or an `ExternalSttProvider` (direct third-party API calls). Notably, the factory **has no local STT branch**; local Whisper inference was removed, with local processing delegated to the `voice_server.stt_engine` configuration.
- **`create_tts_provider`** – Follows a similar pattern but includes a special `PiperTtsProvider` for local subprocess-based synthesis.

Provider implementations live in [`src/openhuman/voice/factory/stt_providers.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/stt_providers.rs) and [`src/openhuman/voice/factory/tts_providers.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/tts_providers.rs), while the trait definitions reside in [`src/openhuman/voice/factory/traits.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/traits.rs). Each provider must implement the `transcribe` method for STT or the `speak` method for TTS, accepting configuration objects and audio data.

## Dictation Server Pipeline

The always-on listening pipeline is implemented in [`src/openhuman/voice/always_on.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/always_on.rs). This module maintains a continuous microphone stream and processes speech through several distinct stages:

### Audio Capture and Preprocessing

The `spawn_capture_thread` function creates a `cpal` input stream that forwards raw interleaved audio chunks to an async processor via an MPSC channel. The system tracks dropped chunks (`DROPPED_CHUNKS`) and logs them on power-of-two intervals to monitor buffer health.

### Voice Activity Detection

The processor task lazily opens a VAD session using `tinyvoice::VadSession::open`, retrying on failure with `SESSION_RETRY_INTERVAL`. Audio undergoes downmixing and resampling via `tinyvoice::prepare_frames` before entering the VAD module, which emits `SpeechStart` and `SpeechEnd` events. The system buffers utterances up to `MAX_UTTERANCE_SAMPLES` (approximately 60 seconds).

### Transcription and Wake-Word Gating

When speech ends, the system invokes `transcribe_and_deliver`, which:

1. Encodes the 16 kHz PCM buffer to WAV format using `tinyvoice::encode_wav`
2. Resolves the effective STT provider via `effective_stt_provider(config)`
3. Instantiates the provider through `create_stt_provider`
4. Sends base64-encoded audio to the provider's `transcribe` method

The resulting transcript passes through `tinyvoice::extract_command` to check against the configured `config.voice_server.wake_word`. If the wake word is detected, the command string flows to `deliver_command`; otherwise, the utterance is discarded or triggers an acknowledgement turn (`"hello"`).

### Fast-Path Intent Routing

The `deliver_command` function first attempts lightweight local routing through `modules::voice::route` for intents like `Pause` or `VolumeUp`. If no fast-path intent matches, the command falls back to the standard LLM pipeline via `dictation_listener::publish_transcription`.

## Configuration and Provider String Grammar

The factory parses provider strings according to a specific grammar documented in the source. Users configure providers via the `stt_provider` and `tts_provider` config fields.

The resolution logic handles several string forms:

- `"cloud"` or `"openhuman"` – Resolves to the OpenHuman backend proxy (STT routes to `/openai/v1/audio/transcriptions`, TTS uses ElevenLabs)
- `"piper"` – Local Piper subprocess for TTS only
- `"<slug>:<model>"` – Third-party STT APIs such as `deepgram:nova-2` or `openai:whisper-1`
- `"<slug>"` – Bare slug using the provider's default model or voice

API keys are resolved dynamically through `lookup_key_for_slug` in [`src/openhuman/voice/factory/helpers.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/helpers.rs), allowing secure credential management without hardcoding.

## Implementation Examples

### Creating an STT Provider Manually

```rust
use openhuman::voice::factory::{create_stt_provider, effective_stt_provider};
use openhuman::config::Config;

// Load the application configuration
let cfg: Config = load_config();

// Resolve the provider name from config (e.g., "openai" or "deepgram:nova-2")
let provider_name = effective_stt_provider(&cfg);

// Instantiate the provider
let stt = create_stt_provider(&provider_name, "", &cfg)
    .expect("STT provider must be available");

// Transcribe base64-encoded WAV audio
let wav_base64 = "..."; // base64-encoded 16kHz PCM
let result = stt.transcribe(&cfg, &wav_base64, Some("audio/wav"), None, Some("en")).await?;
println!("Transcribed text: {}", result.value.text);

```

### Starting the Always-On Dictation Server

```rust
use openhuman::voice::always_on::start_if_enabled;
use openhuman::config::Config;

let cfg = Config::load_or_init().await.unwrap();

// Launch the microphone pipeline if enabled in config
tokio::spawn(async move {
    start_if_enabled(&cfg).await;
});

```

### Handling Fast-Path Intents

```rust
use openhuman::voice::ops::deliver_command;
use openhuman::config::Config;

let command = "pause".to_string();
let cfg = Config::load_or_init().await.unwrap();

// Routes to local intent handler or falls back to LLM
deliver_command(&cfg, command).await;

```

## Summary

- The **STT Provider Factory** in [`src/openhuman/voice/factory/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/mod.rs) instantiates speech providers based on configuration strings like `"cloud"` or `"deepgram:nova-2"`.
- **Provider resolution** occurs dynamically via `effective_stt_provider` and `effective_tts_provider`, with API keys fetched securely from the global config.
- The **Dictation Server** in [`src/openhuman/voice/always_on.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/always_on.rs) maintains a continuous `cpal` audio stream, processes voice activity through `tinyvoice`, and transcribes utterances up to 60 seconds.
- A **wake-word gate** filters transcripts before routing; matched commands flow through fast-path intent handlers (`modules::voice::route`) or the standard LLM pipeline.
- The architecture contains **no local STT inference** in the factory; local processing requires external configuration via `voice_server.stt_engine`.

## Frequently Asked Questions

### What provider strings does OpenHuman support for STT?

OpenHuman supports `"cloud"` (OpenHuman backend proxy), `"openai:whisper-1"`, `"deepgram:nova-2"`, and other third-party slugs following the `"<slug>:<model>"` pattern. Bare slugs like `"openai"` use provider defaults. The factory resolves these in [`src/openhuman/voice/factory/mod.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/factory/mod.rs) and maps them to `CloudSttProvider` or `ExternalSttProvider` implementations.

### How does the Dictation Server handle voice activity detection?

The server uses the `tinyvoice` crate to run `VadSession` analysis on audio captured via `cpal`. The processor task in [`src/openhuman/voice/always_on.rs`](https://github.com/tinyhumansai/openhuman/blob/main/src/openhuman/voice/always_on.rs) prepares frames through `tinyvoice::prepare_frames` and buffers speech between `SpeechStart` and `SpeechEnd` events, enforcing a maximum utterance length of `MAX_UTTERANCE_SAMPLES` (approximately 60 seconds).

### Can OpenHuman run STT locally without cloud providers?

The factory itself has no local STT branch; it only provides `CloudSttProvider` and `ExternalSttProvider`. Local Whisper inference was removed from the factory codebase. Local STT processing requires configuring `voice_server.stt_engine` separately, which operates outside the factory pattern described in `src/openhuman/voice/factory/`.

### How are transcribed commands routed after wake-word detection?

After transcription, `tinyvoice::extract_command` checks for the configured `wake_word`. If matched, `deliver_command` attempts fast-path routing through `modules::voice::route` for local intents (e.g., volume control). Unmatched commands are published to the dictation bus via `dictation_listener::publish_transcription` for LLM processing.