# How TextToSpeechService Switches Between Windows Speech Synthesis and ElevenLabs in openclaw-windows-node

> Discover how openclaw-windows-node's TextToSpeechService switches between Windows Speech Synthesis and ElevenLabs dynamically. Learn provider selection and implementation details.

- Repository: [openclaw/openclaw-windows-node](https://github.com/openclaw/openclaw-windows-node)
- Tags: internals
- Published: 2026-06-05

---

**The TextToSpeechService dynamically selects between Windows Speech Synthesis and ElevenLabs at runtime by resolving the provider identifier from either command arguments or user settings, then dispatching to the corresponding `SpeakWith...Async` implementation.**

The openclaw-windows-node repository implements a flexible Text-to-Speech (TTS) system that supports multiple synthesis backends. Understanding how the `TextToSpeechService` class handles provider switching is essential for configuring voice output in OpenClaw's Windows node implementation.

## Provider Resolution Strategy

The service determines which backend to use through a two-stage resolution process that prioritizes explicit requests over default configurations.

### Merging Command and Configuration Preferences

At the entry point of `TextToSpeechService.SpeakAsync`, the provider selection logic calls `TtsCapability.ResolveProvider` to merge inputs:

```csharp
var provider = TtsCapability.ResolveProvider(args.Provider, _settings.TtsProvider);

```

This resolution follows a simple priority rule: if the `tts.speak` command includes a specific `Provider` value in its arguments (`args.Provider`), that value takes precedence. When the command omits the provider specification, the system falls back to the user-configured default stored in `SettingsManager.TtsProvider`.

### Supported Provider Constants

The resolution logic relies on string constants defined in [`TtsCapability.cs`](https://github.com/openclaw/openclaw-windows-node/blob/main/TtsCapability.cs) to identify valid backends:

- **`TtsCapability.WindowsProvider`** – Local Windows Speech Synthesis engine
- **`TtsCapability.ElevenLabsProvider`** – Cloud-based ElevenLabs API
- **`TtsCapability.PiperProvider`** – Local Piper TTS engine

## Runtime Provider Dispatch

Once resolved, the provider string drives a straightforward conditional dispatch chain within [`TextToSpeechService.cs`](https://github.com/openclaw/openclaw-windows-node/blob/main/TextToSpeechService.cs):

```csharp
if (string.Equals(provider, TtsCapability.WindowsProvider, ...))
    await SpeakWithWindowsAsync(args, cancellationToken);
else if (string.Equals(provider, TtsCapability.ElevenLabsProvider, ...))
    await SpeakWithElevenLabsAsync(args, cancellationToken);
else if (string.Equals(provider, TtsCapability.PiperProvider, ...))
    await SpeakWithPiperAsync(args, cancellationToken);
else
    throw new InvalidOperationException($"Unsupported TTS provider '{provider}'.");

```

This structure ensures that each provider follows its own dedicated code path while both ultimately converge on `PlayStreamAsync` for exclusive audio playback that supports interruption.

### Windows Speech Synthesis Path

When the resolved provider matches `WindowsProvider`, the service invokes `SpeakWithWindowsAsync`. This method instantiates a `Windows.Media.SpeechSynthesis.SpeechSynthesizer`, optionally applies a specific voice from settings, synthesizes the text to a WAV stream, and plays it through a `MediaPlayer` instance.

### ElevenLabs Cloud Path

For `ElevenLabsProvider`, the service calls `SpeakWithElevenLabsAsync`, which retrieves the API key and voice ID from `_settings`, then delegates to `ElevenLabsTextToSpeechClient.SynthesizeAsync`. The returned MP3 byte array converts to an audio stream before playback.

## Implementation Examples

You can leverage this switching mechanism programmatically by either accepting defaults or explicitly overriding them per utterance:

```csharp
// Example 1 – Use the default provider from settings (e.g., Windows)
await nodeService.TextToSpeech?.SpeakAsync(
    new TtsSpeakArgs { Text = "Hello, world!" }, CancellationToken.None);

// Example 2 – Force Eleven Labs for a single utterance
await nodeService.TextToSpeech?.SpeakAsync(
    new TtsSpeakArgs {
        Text = "Hello from Eleven Labs",
        Provider = TtsCapability.ElevenLabsProvider   // overrides the default
    }, CancellationToken.None);

```

Both calls route through `TextToSpeechService.SpeakAsync`, which applies the resolution logic and executes the matching backend method.

## Summary

- **Priority-based resolution**: The `TtsCapability.ResolveProvider` helper checks command arguments first, then falls back to `SettingsManager.TtsProvider` for the default backend.
- **String-based dispatch**: The service uses simple string comparison against provider constants (`WindowsProvider`, `ElevenLabsProvider`) to select implementation paths.
- **Consistent playback**: Both Windows Speech and ElevenLabs providers ultimately call `PlayStreamAsync`, ensuring uniform audio handling and interruption support.
- **Extensible design**: The conditional chain in [`TextToSpeechService.cs`](https://github.com/openclaw/openclaw-windows-node/blob/main/TextToSpeechService.cs) allows easy addition of new providers like Piper without modifying the core resolution logic.

## Frequently Asked Questions

### How does TextToSpeechService prioritize which TTS provider to use?

The service prioritizes the provider specified in the `TtsSpeakArgs.Provider` property of the individual command. If this value is null or empty, it falls back to the system-wide default stored in `SettingsManager.TtsProvider`, allowing per-command overrides while maintaining a configurable baseline.

### Can I switch between Windows Speech and ElevenLabs within the same application session?

Yes. Because the provider resolution occurs inside `SpeakAsync` on every call, you can mix providers freely within the same session. Simply pass different `Provider` values in your `TtsSpeakArgs` for each utterance, or omit the field to use the current default.

### What happens if I specify an unsupported provider name?

If the resolved provider string does not match `WindowsProvider`, `ElevenLabsProvider`, or `PiperProvider`, the service throws an `InvalidOperationException` with the message "Unsupported TTS provider '{provider}'." This immediate failure prevents silent fallback to unexpected backends.

### Where does the ElevenLabs implementation store API credentials?

According to the source code in [`TextToSpeechService.cs`](https://github.com/openclaw/openclaw-windows-node/blob/main/TextToSpeechService.cs), the `SpeakWithElevenLabsAsync` method reads the ElevenLabs API key and voice ID from the `_settings` object (specifically `SettingsManager`), not from the command arguments, keeping sensitive credentials separate from runtime text parameters.