# How to Use the generate_speech Tool with the VoiceStudio MCP Server

> Learn to use the generate_speech tool with the VoiceStudio MCP server to convert text to speech. Get PCM16 bytes or file references efficiently.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-09

---

**The generate_speech tool converts text into synthesized audio through the VoiceStudio MCP server, returning either raw PCM16 bytes or file system references based on the OMNIVOICE_MCP_OUTPUT_MODE environment variable.**

The VoiceStudio repository by debpalash provides a lightweight Multimedia Control Protocol (MCP) server that exposes a text-to-speech helper called `generate_speech`. This function serves as the bridge between the OpenAI-compatible `/v1/audio/speech` HTTP endpoint and underlying TTS engines like ElevenLabs or Azure. Understanding how to configure and invoke this tool is essential for integrating voice synthesis into your applications.

## What generate_speech Does

The `generate_speech` function in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) accepts a text string and voice identifier, then routes the request to the selected synthesis engine. It returns raw **PCM 16-bit audio at 16 kHz**, with the delivery method controlled by the `OMNIVOICE_MCP_OUTPUT_MODE` environment variable.

The tool supports three distinct output behaviors:

- **resources** (default): Returns raw audio bytes directly in the API response with `Content-Type: application/octet-stream`.
- **files**: Saves a temporary WAV file under the directory specified by `OMNIVOICE_MCP_BASE_PATH` and returns a JSON object containing the file path.
- **both**: Returns the raw bytes and writes the file simultaneously, exposing both the audio data and a `note` field with the file location.

When `files` or `both` mode is active, the function validates that the target path remains inside the base directory to prevent directory traversal attacks.

## Configuration Environment Variables

The VoiceStudio MCP server reads these environment variables once per request to determine `generate_speech` behavior:

| Variable | Description | Required/Default |
|----------|-------------|------------------|
| `OMNIVOICE_MCP_OUTPUT_MODE` | Controls output format: `resources`, `files`, or `both` | Default: `resources` |
| `OMNIVOICE_MCP_BASE_PATH` | Absolute root directory for file outputs | Required when mode is `files` or `both` |
| `OMNIVOICE_MCP_TIMEOUT_S` | Maximum seconds to wait for backend synthesis | Default: `120` |

If `OMNIVOICE_MCP_BASE_PATH` is unset when file output is requested, `generate_speech` raises a `ValueError` indicating the base path is required.

## Implementation Details in backend/mcp_server.py

According to the VoiceStudio source code, the `generate_speech` implementation resides in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py). Lines 107-119 handle the parsing of `OMNIVOICE_MCP_OUTPUT_MODE`, while lines 120-141 perform strict validation of the base path to ensure all file writes are constrained to the designated directory. The function also resolves the appropriate TTS engine internally based on the model identifier passed from the API layer.

## How to Call the generate_speech Tool

### Via the OpenAI-Compatible HTTP API

The endpoint defined in [`api/routers/speech.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/speech.py) forwards incoming requests to `generate_speech`, making it accessible through standard HTTP clients.

**Request raw audio bytes (default mode):**

```python
import requests

payload = {
    "model": "elevenlabs/voice-xyz",
    "input": "Hello, world!",
    "voice": "en_us_001"
}

response = requests.post(
    "http://localhost:8000/v1/audio/speech",
    json=payload,
    headers={"Accept": "application/octet-stream"}
)

audio_bytes = response.content  # PCM-16 @ 16kHz

```

**Request file output with path validation:**

```python
import os
import requests

os.environ["OMNIVOICE_MCP_OUTPUT_MODE"] = "files"
os.environ["OMNIVOICE_MCP_BASE_PATH"] = "/tmp/mcp_outputs"

response = requests.post(
    "http://localhost:8000/v1/audio/speech",
    json=payload,
    headers={"Accept": "application/json"}
)

result = response.json()
print("Audio saved at:", result["note"])  # e.g., "/tmp/mcp_outputs/voice-xyz-123.wav"

```

### Direct Python Invocation

For batch processing or internal service integrations, import the tool directly from the backend module:

```python
import os
from backend.mcp_server import generate_speech

os.environ["OMNIVOICE_MCP_OUTPUT_MODE"] = "both"
os.environ["OMNIVOICE_MCP_BASE_PATH"] = "/tmp/mcp_outputs"

audio_bytes, file_note = generate_speech(
    text="Good morning!",
    voice_id="en_us_001",
    engine="elevenlabs"
)

print(f"Returned {len(audio_bytes)} bytes")
print(f"File location: {file_note}")

```

## Security and Path Validation

When operating in `files` or `both` mode, `generate_speech` strictly validates that any generated file path resides within `OMNIVOICE_MCP_BASE_PATH`. This prevents directory traversal attacks that could attempt to write audio files outside the designated output directory. The validation logic in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) checks path containment before any filesystem write operations occur.

## Summary

- The **generate_speech** tool in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) provides text-to-speech synthesis for the VoiceStudio MCP server.
- Behavior is controlled by `OMNIVOICE_MCP_OUTPUT_MODE`, `OMNIVOICE_MCP_BASE_PATH`, and `OMNIVOICE_MCP_TIMEOUT_S`.
- Three output modes exist: `resources` (raw bytes), `files` (saved WAV), and `both` (bytes plus file path).
- The tool enforces path validation to prevent writes outside the configured base directory.
- Access the tool via the OpenAI-compatible `/v1/audio/speech` endpoint in [`api/routers/speech.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/speech.py) or import it directly for programmatic use.

## Frequently Asked Questions

### What audio format does generate_speech return?

The tool returns raw PCM 16-bit audio sampled at 16 kHz. When saving to disk in `files` or `both` mode, the output is wrapped in a WAV container to preserve the sample rate and bit depth metadata.

### How do I switch between output modes?

Set the `OMNIVOICE_MCP_OUTPUT_MODE` environment variable to `resources`, `files`, or `both` before starting the server. The variable is evaluated once per request at runtime, so changes take effect immediately for subsequent calls without requiring a server restart.

### What happens if OMNIVOICE_MCP_BASE_PATH is not set?

If you request `files` or `both` output mode without configuring `OMNIVOICE_MCP_BASE_PATH`, the `generate_speech` function raises a `ValueError` indicating that the base path is required. This check ensures the server never attempts to write audio files to an undefined or potentially insecure location.

### Where is the OpenAI-compatible endpoint implemented?

The HTTP endpoint that wraps `generate_speech` is implemented in [`api/routers/speech.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/speech.py). This router handles POST requests to `/v1/audio/speech`, validates incoming JSON payloads, and forwards parameters to the MCP server backend before returning the synthesized audio in the configured format.