How to Use the generate_speech Tool with the VoiceStudio MCP Server
The generate_speech tool converts text into synthesized audio through the VoiceStudio MCP server, returning either raw PCM16 bytes or file system references based on the OMNIVOICE_MCP_OUTPUT_MODE environment variable.
The VoiceStudio repository by debpalash provides a lightweight Multimedia Control Protocol (MCP) server that exposes a text-to-speech helper called generate_speech. This function serves as the bridge between the OpenAI-compatible /v1/audio/speech HTTP endpoint and underlying TTS engines like ElevenLabs or Azure. Understanding how to configure and invoke this tool is essential for integrating voice synthesis into your applications.
What generate_speech Does
The generate_speech function in backend/mcp_server.py accepts a text string and voice identifier, then routes the request to the selected synthesis engine. It returns raw PCM 16-bit audio at 16 kHz, with the delivery method controlled by the OMNIVOICE_MCP_OUTPUT_MODE environment variable.
The tool supports three distinct output behaviors:
- resources (default): Returns raw audio bytes directly in the API response with
Content-Type: application/octet-stream. - files: Saves a temporary WAV file under the directory specified by
OMNIVOICE_MCP_BASE_PATHand returns a JSON object containing the file path. - both: Returns the raw bytes and writes the file simultaneously, exposing both the audio data and a
notefield with the file location.
When files or both mode is active, the function validates that the target path remains inside the base directory to prevent directory traversal attacks.
Configuration Environment Variables
The VoiceStudio MCP server reads these environment variables once per request to determine generate_speech behavior:
| Variable | Description | Required/Default |
|---|---|---|
OMNIVOICE_MCP_OUTPUT_MODE |
Controls output format: resources, files, or both |
Default: resources |
OMNIVOICE_MCP_BASE_PATH |
Absolute root directory for file outputs | Required when mode is files or both |
OMNIVOICE_MCP_TIMEOUT_S |
Maximum seconds to wait for backend synthesis | Default: 120 |
If OMNIVOICE_MCP_BASE_PATH is unset when file output is requested, generate_speech raises a ValueError indicating the base path is required.
Implementation Details in backend/mcp_server.py
According to the VoiceStudio source code, the generate_speech implementation resides in backend/mcp_server.py. Lines 107-119 handle the parsing of OMNIVOICE_MCP_OUTPUT_MODE, while lines 120-141 perform strict validation of the base path to ensure all file writes are constrained to the designated directory. The function also resolves the appropriate TTS engine internally based on the model identifier passed from the API layer.
How to Call the generate_speech Tool
Via the OpenAI-Compatible HTTP API
The endpoint defined in api/routers/speech.py forwards incoming requests to generate_speech, making it accessible through standard HTTP clients.
Request raw audio bytes (default mode):
import requests
payload = {
"model": "elevenlabs/voice-xyz",
"input": "Hello, world!",
"voice": "en_us_001"
}
response = requests.post(
"http://localhost:8000/v1/audio/speech",
json=payload,
headers={"Accept": "application/octet-stream"}
)
audio_bytes = response.content # PCM-16 @ 16kHz
Request file output with path validation:
import os
import requests
os.environ["OMNIVOICE_MCP_OUTPUT_MODE"] = "files"
os.environ["OMNIVOICE_MCP_BASE_PATH"] = "/tmp/mcp_outputs"
response = requests.post(
"http://localhost:8000/v1/audio/speech",
json=payload,
headers={"Accept": "application/json"}
)
result = response.json()
print("Audio saved at:", result["note"]) # e.g., "/tmp/mcp_outputs/voice-xyz-123.wav"
Direct Python Invocation
For batch processing or internal service integrations, import the tool directly from the backend module:
import os
from backend.mcp_server import generate_speech
os.environ["OMNIVOICE_MCP_OUTPUT_MODE"] = "both"
os.environ["OMNIVOICE_MCP_BASE_PATH"] = "/tmp/mcp_outputs"
audio_bytes, file_note = generate_speech(
text="Good morning!",
voice_id="en_us_001",
engine="elevenlabs"
)
print(f"Returned {len(audio_bytes)} bytes")
print(f"File location: {file_note}")
Security and Path Validation
When operating in files or both mode, generate_speech strictly validates that any generated file path resides within OMNIVOICE_MCP_BASE_PATH. This prevents directory traversal attacks that could attempt to write audio files outside the designated output directory. The validation logic in backend/mcp_server.py checks path containment before any filesystem write operations occur.
Summary
- The generate_speech tool in
backend/mcp_server.pyprovides text-to-speech synthesis for the VoiceStudio MCP server. - Behavior is controlled by
OMNIVOICE_MCP_OUTPUT_MODE,OMNIVOICE_MCP_BASE_PATH, andOMNIVOICE_MCP_TIMEOUT_S. - Three output modes exist:
resources(raw bytes),files(saved WAV), andboth(bytes plus file path). - The tool enforces path validation to prevent writes outside the configured base directory.
- Access the tool via the OpenAI-compatible
/v1/audio/speechendpoint inapi/routers/speech.pyor import it directly for programmatic use.
Frequently Asked Questions
What audio format does generate_speech return?
The tool returns raw PCM 16-bit audio sampled at 16 kHz. When saving to disk in files or both mode, the output is wrapped in a WAV container to preserve the sample rate and bit depth metadata.
How do I switch between output modes?
Set the OMNIVOICE_MCP_OUTPUT_MODE environment variable to resources, files, or both before starting the server. The variable is evaluated once per request at runtime, so changes take effect immediately for subsequent calls without requiring a server restart.
What happens if OMNIVOICE_MCP_BASE_PATH is not set?
If you request files or both output mode without configuring OMNIVOICE_MCP_BASE_PATH, the generate_speech function raises a ValueError indicating that the base path is required. This check ensures the server never attempts to write audio files to an undefined or potentially insecure location.
Where is the OpenAI-compatible endpoint implemented?
The HTTP endpoint that wraps generate_speech is implemented in api/routers/speech.py. This router handles POST requests to /v1/audio/speech, validates incoming JSON payloads, and forwards parameters to the MCP server backend before returning the synthesized audio in the configured format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →