How to Customize Speaker Identity in Speech-to-Speech: CLI and API Methods

You can customize speaker identity in speech-to-speech by either setting the qwen3_tts_speaker argument during pipeline initialization or by passing a session_voice field in runtime API requests, with the handler validating names against the model's supported speaker list and falling back to the first available voice if necessary.

The huggingface/speech-to-speech repository provides flexible speaker customization for voice synthesis through the Qwen3TTS handler. When you need to customize speaker identity in speech-to-speech applications, the system supports both static configuration and dynamic runtime overrides. This article explores the two primary mechanisms for controlling voice output based on the actual implementation in the source code.

Setting the Default Speaker via CLI Arguments

The primary method for configuring speaker identity relies on the Qwen3TTSHandlerArguments dataclass defined in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py (lines 47-52). This class exposes the qwen3_tts_speaker field, which defaults to "Aiden".

When you instantiate the pipeline with a specific speaker name, the handler stores this value in self.speaker and preserves the original user input in self._initial_speaker (as seen in the __init__ method of Qwen3TTSHandler at line 124). The handler validates the provided name against the list of supported speakers reported by the underlying CustomVoice model. If the name matches a supported speaker, the model synthesizes audio using that voice; otherwise, the system falls back to the first supported speaker.

Runtime Speaker Overrides via API

For dynamic voice switching during active sessions, the Qwen3TTSHandler class implements the _apply_session_voice_override() method (lines 429-442 in src/speech_to_speech/TTS/qwen3_tts_handler.py). This mechanism activates when the client sends a session_voice field in the request payload.

The handler normalizes the input string by lower-casing it and mapping it to the canonical speaker name from the model's get_supported_speakers() list. If the requested name is not supported, the handler logs a warning message (lines 432-435) and retains the current speaker identity rather than crashing the pipeline.

Speaker Resolution and Fallback Mechanism

The internal _resolve_speaker() method (lines 496-508 in qwen3_tts_handler.py) manages the final speaker selection logic. This method ensures that:

  • Explicit arguments passed via --qwen3_tts_speaker take precedence during initialization
  • Runtime overrides via session_voice update the active speaker for subsequent requests
  • When no speaker is explicitly set, the system defaults to the first entry in the supported speakers list
  • Unsupported speaker names trigger graceful degradation with logged warnings

The resolution occurs in sequence: first checking the explicit argument, then applying any session-level override, and finally falling back to the model's first available voice if necessary.

Implementation Examples

Command Line Interface

To specify a different CustomVoice speaker (such as "Vivian") when launching from the command line:

python -m speech_to_speech.run \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
  --qwen3_tts_speaker Vivian \
  --input "Hello, how are you?" \
  --output hello.wav

Python Pipeline Configuration

To configure the speaker programmatically using the argument dataclass:

from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSHandlerArguments
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline

args = Qwen3TTSHandlerArguments(
    qwen3_tts_speaker="Vivian",          # <- custom speaker name

    qwen3_tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
)

pipeline = SpeechToSpeechPipeline(tts_args=args)
audio = pipeline.run_text("Welcome to the demo!")
pipeline.save_wav(audio, "welcome.wav")

Runtime API Request

To override the speaker for a single request via the API:

{
  "model": "speech-to-speech",
  "input": "Can you read this sentence?",
  "session_voice": "Vivian"          // <- overrides the default speaker for this call
}

The server maps "Vivian" to the canonical speaker name reported by the model and synthesizes with that voice for the specific request.

Summary

  • Set static speaker identity via the qwen3_tts_speaker argument in Qwen3TTSHandlerArguments (default: "Aiden")
  • Override voices dynamically at runtime using the session_voice field in API requests
  • The handler validates speaker names against get_supported_speakers() with automatic fallback to the first available voice
  • Unsupported speaker names trigger warnings in _apply_session_voice_override() while preserving the current speaker configuration

Frequently Asked Questions

What is the default speaker in speech-to-speech?

The default speaker is "Aiden", as defined in the Qwen3TTSHandlerArguments dataclass in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py (lines 47-52). This value is used when no explicit speaker is provided during pipeline initialization.

How do I check which speakers are supported by my model?

The underlying CustomVoice model exposes a get_supported_speakers() method that returns the list of valid speaker names. The Qwen3TTSHandler uses this method internally to validate both initialization arguments and runtime overrides against the model's capabilities.

What happens if I specify an unsupported speaker name?

When an unsupported name is requested, the _apply_session_voice_override() method (lines 429-442 in src/speech_to_speech/TTS/qwen3_tts_handler.py) logs a warning message (lines 432-435) and keeps the current speaker identity rather than raising an exception. This ensures graceful degradation while alerting you to the configuration issue.

Can I change the speaker identity mid-session?

Yes, you can change speakers dynamically by including the session_voice field in your API request payload. The handler processes this override immediately, validating the new name against supported speakers and updating the synthesis voice for that specific request without affecting the default configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →