How the `--stt none` Mode Works with Audio-Input Capable LLM Models

The --stt none flag disables the speech-to-text module and routes raw audio chunks directly to LLM backends that declare supports_audio_input=True, allowing the model to consume audio natively without intermediate transcription.

The huggingface/speech-to-speech repository provides a modular pipeline for real-time voice conversations. When using audio-input capable LLM models, you can bypass the traditional STT component entirely by setting --stt none, enabling the LLM to process raw audio directly.

Architecture of the --stt none Mode

CLI Argument Parsing

The --stt flag is defined in the pipeline entry point. In src/speech_to_speech/s2s_pipeline.py, the argument parser accepts "none" as a valid value for the STT module (lines 195–200).

Capability Validation

Before constructing the pipeline, the prepare_module_args function validates that the selected LLM backend supports audio input when --stt none is specified. The check occurs at lines 321–324:

if module_kwargs.stt == "none" and not llm_backend.spec.capabilities.supports_audio_input:
    supported = ", ".join(name for name, spec in LLM_BACKENDS.items()
                          if spec.capabilities.supports_audio_input)
    raise ValueError(
        f"--stt none requires an audio-input LLM backend; choose one of: {supported}."
    )

This validation relies on the BackendCapabilities dataclass defined in src/speech_to_speech/backend_registry.py.

Backend Capability Registry

Each LLM backend registers its capabilities in src/speech_to_speech/backend_registry.py. For audio-input capable models like chat-completions, the supports_audio_input flag is set to True (lines 393–406):

LLM_BACKENDS["chat-completions"] = BackendSpec(
    ...,
    capabilities=BackendCapabilities(supports_audio_input=True, supports_llm_proxy=True),
)

Audio Routing Pipeline

When validation passes and stt is set to "none", the pipeline omits the STT handler. Instead, the Voice Activity Detector (VAD) passes audio chunks directly to the AudioInputNotifier located in src/speech_to_speech/LLM/audio_input_notifier.py. This component queues raw audio for the LLM, which processes the waveform as its input prompt and returns text for the TTS stage.

Usage Examples

Command Line


# Route audio directly to the chat-completions LLM without STT transcription

python -m speech_to_speech serve --stt none --llm_backend chat-completions --tts parler-tts

Python API

from speech_to_speech.s2s_pipeline import parse_arguments, prepare_all_args

# Simulate CLI arguments

args = parse_arguments([
    "--stt", "none",
    "--llm_backend", "chat-completions",
    "--tts", "parler-tts"
])

# Validate configuration

prepare_all_args(args)

# Verify audio input support

print(f"Backend: {args.llm_backend.name}")
print(f"Supports audio input: {args.llm_backend.spec.capabilities.supports_audio_input}")

Key Source Files

Summary

  • The --stt none flag disables the speech-to-text module entirely.
  • The pipeline validates that the selected LLM backend has supports_audio_input=True before startup.
  • Audio flows directly from the VAD to the LLM via the AudioInputNotifier.
  • This mode reduces latency by eliminating the transcription step and leverages native audio understanding in models like chat-completions.

Frequently Asked Questions

What happens if I use --stt none with a non-audio LLM backend?

The pipeline raises a ValueError during initialization. The error message lists all available backends that support audio input, as enforced in prepare_module_args within s2s_pipeline.py.

Which LLM backends currently support audio input?

According to the registry in backend_registry.py, the chat-completions backend explicitly sets supports_audio_input=True. Other backends may add support by updating their BackendCapabilities registration.

How does the VAD interact with the LLM when STT is disabled?

The VAD continues to detect voice activity and extract audio chunks. Instead of sending these chunks to an STT handler, it publishes them to the AudioInputNotifier, which the LLM consumes as its input stream.

Can I mix text and audio inputs when using --stt none?

When --stt none is active, the primary input modality is audio. Text handling depends on the specific LLM backend implementation, but the standard pipeline configuration expects audio chunks as the primary prompt source.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →