How TextToSpeechService Switches Between Windows Speech Synthesis and ElevenLabs in openclaw-windows-node

The TextToSpeechService dynamically selects between Windows Speech Synthesis and ElevenLabs at runtime by resolving the provider identifier from either command arguments or user settings, then dispatching to the corresponding SpeakWith...Async implementation.

The openclaw-windows-node repository implements a flexible Text-to-Speech (TTS) system that supports multiple synthesis backends. Understanding how the TextToSpeechService class handles provider switching is essential for configuring voice output in OpenClaw's Windows node implementation.

Provider Resolution Strategy

The service determines which backend to use through a two-stage resolution process that prioritizes explicit requests over default configurations.

Merging Command and Configuration Preferences

At the entry point of TextToSpeechService.SpeakAsync, the provider selection logic calls TtsCapability.ResolveProvider to merge inputs:

var provider = TtsCapability.ResolveProvider(args.Provider, _settings.TtsProvider);

This resolution follows a simple priority rule: if the tts.speak command includes a specific Provider value in its arguments (args.Provider), that value takes precedence. When the command omits the provider specification, the system falls back to the user-configured default stored in SettingsManager.TtsProvider.

Supported Provider Constants

The resolution logic relies on string constants defined in TtsCapability.cs to identify valid backends:

  • TtsCapability.WindowsProvider – Local Windows Speech Synthesis engine
  • TtsCapability.ElevenLabsProvider – Cloud-based ElevenLabs API
  • TtsCapability.PiperProvider – Local Piper TTS engine

Runtime Provider Dispatch

Once resolved, the provider string drives a straightforward conditional dispatch chain within TextToSpeechService.cs:

if (string.Equals(provider, TtsCapability.WindowsProvider, ...))
    await SpeakWithWindowsAsync(args, cancellationToken);
else if (string.Equals(provider, TtsCapability.ElevenLabsProvider, ...))
    await SpeakWithElevenLabsAsync(args, cancellationToken);
else if (string.Equals(provider, TtsCapability.PiperProvider, ...))
    await SpeakWithPiperAsync(args, cancellationToken);
else
    throw new InvalidOperationException($"Unsupported TTS provider '{provider}'.");

This structure ensures that each provider follows its own dedicated code path while both ultimately converge on PlayStreamAsync for exclusive audio playback that supports interruption.

Windows Speech Synthesis Path

When the resolved provider matches WindowsProvider, the service invokes SpeakWithWindowsAsync. This method instantiates a Windows.Media.SpeechSynthesis.SpeechSynthesizer, optionally applies a specific voice from settings, synthesizes the text to a WAV stream, and plays it through a MediaPlayer instance.

ElevenLabs Cloud Path

For ElevenLabsProvider, the service calls SpeakWithElevenLabsAsync, which retrieves the API key and voice ID from _settings, then delegates to ElevenLabsTextToSpeechClient.SynthesizeAsync. The returned MP3 byte array converts to an audio stream before playback.

Implementation Examples

You can leverage this switching mechanism programmatically by either accepting defaults or explicitly overriding them per utterance:

// Example 1 – Use the default provider from settings (e.g., Windows)
await nodeService.TextToSpeech?.SpeakAsync(
    new TtsSpeakArgs { Text = "Hello, world!" }, CancellationToken.None);

// Example 2 – Force Eleven Labs for a single utterance
await nodeService.TextToSpeech?.SpeakAsync(
    new TtsSpeakArgs {
        Text = "Hello from Eleven Labs",
        Provider = TtsCapability.ElevenLabsProvider   // overrides the default
    }, CancellationToken.None);

Both calls route through TextToSpeechService.SpeakAsync, which applies the resolution logic and executes the matching backend method.

Summary

  • Priority-based resolution: The TtsCapability.ResolveProvider helper checks command arguments first, then falls back to SettingsManager.TtsProvider for the default backend.
  • String-based dispatch: The service uses simple string comparison against provider constants (WindowsProvider, ElevenLabsProvider) to select implementation paths.
  • Consistent playback: Both Windows Speech and ElevenLabs providers ultimately call PlayStreamAsync, ensuring uniform audio handling and interruption support.
  • Extensible design: The conditional chain in TextToSpeechService.cs allows easy addition of new providers like Piper without modifying the core resolution logic.

Frequently Asked Questions

How does TextToSpeechService prioritize which TTS provider to use?

The service prioritizes the provider specified in the TtsSpeakArgs.Provider property of the individual command. If this value is null or empty, it falls back to the system-wide default stored in SettingsManager.TtsProvider, allowing per-command overrides while maintaining a configurable baseline.

Can I switch between Windows Speech and ElevenLabs within the same application session?

Yes. Because the provider resolution occurs inside SpeakAsync on every call, you can mix providers freely within the same session. Simply pass different Provider values in your TtsSpeakArgs for each utterance, or omit the field to use the current default.

What happens if I specify an unsupported provider name?

If the resolved provider string does not match WindowsProvider, ElevenLabsProvider, or PiperProvider, the service throws an InvalidOperationException with the message "Unsupported TTS provider '{provider}'." This immediate failure prevents silent fallback to unexpected backends.

Where does the ElevenLabs implementation store API credentials?

According to the source code in TextToSpeechService.cs, the SpeakWithElevenLabsAsync method reads the ElevenLabs API key and voice ID from the _settings object (specifically SettingsManager), not from the command arguments, keeping sensitive credentials separate from runtime text parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →