VoiceStudio MCP Server: Complete Tool Reference and API Guide
The VoiceStudio MCP server exposes seven JSON-RPC tools—generate_speech, clone_voice, transcribe, list_voices, list_languages, list_personalities, and check_health—that enable AI agents to synthesize speech, clone voices, and transcribe audio via the /mcp HTTP endpoint.
The VoiceStudio MCP server implements the Model Context Protocol (MCP) to bridge AI agents with voice synthesis capabilities. Bundled with the VoiceStudio repository, this server mounts automatically on the backend at the /mcp endpoint, providing a JSON-RPC/HTTP interface for text-to-speech generation, voice cloning, and audio transcription as implemented in backend/mcp_server.py.
Available Tools on the VoiceStudio MCP Server
The server defines seven primary tool functions that agents can invoke via JSON-RPC payloads. These are implemented in backend/mcp_server.py (lines 8-15) and documented in docs/mcp.md (lines 11-16).
generate_speech
The generate_speech tool converts text or voice-design prompts into WAV audio clips. It accepts a text parameter and optionally a voice_profile to control synthesis characteristics. Depending on the OMNIVOICE_MCP_OUTPUT_MODE environment variable, it returns either Base64-encoded audio (resources), a file URL (files), or both (both).
clone_voice
The clone_voice tool creates new voice profiles from reference audio. It accepts ref_audio_base64 (a Base64-encoded audio clip) and returns a unique profile_id that can be passed to subsequent generate_speech calls. This enables personalized voice synthesis without pre-training.
transcribe
The transcribe tool performs speech-to-text conversion, supporting 646 languages. It accepts either audio_path (relative to OMNIVOICE_MCP_BASE_PATH) or audio_base64 for direct input. The tool returns plain-text transcription of the provided audio content.
list_voices
The list_voices tool enumerates all stored voice profiles in the system. It returns an array of voice-profile metadata including identifiers and configuration parameters. Agents can query this to present voice selection options to users.
list_languages
The list_languages tool retrieves all TTS languages supported by the backend synthesis engine. It returns an array of language identifiers that can be used when configuring transcription or generation parameters.
list_personalities
The list_personalities tool exposes preset voice personalities configured in the system. It returns an array of personality descriptors that agents can reference to apply predefined speaking styles to generated speech.
check_health
The check_health tool reports backend health status and active GPU availability. It returns JSON fields indicating service readiness and hardware utilization, useful for agent orchestration and monitoring.
Environment Configuration and Security
The VoiceStudio MCP server respects two critical environment variables defined in backend/mcp_server.py that control output handling and file system security.
OMNIVOICE_MCP_OUTPUT_MODE controls audio response format:
resources: Returns Base64-encoded WAV data inline (default)files: Returns a URL and file path to the generated audioboth: Returns both Base64 data and file references
OMNIVOICE_MCP_BASE_PATH defines the security boundary for file operations. Agents may only read from or write to paths within this directory. This prevents unauthorized file system access when processing audio_path arguments.
How to Invoke VoiceStudio MCP Server Tools
Tools are invoked via HTTP POST requests to http://localhost:3900/mcp with JSON-RPC 2.0 payloads.
Generate speech with default resources mode:
curl -X POST http://localhost:3900/mcp \
-H "Content-Type: application/json" \
-d '{
"jsonrpc":"2.0",
"method":"generate_speech",
"params":{"text":"Hello, world!"},
"id":1
}'
Configure files mode and generate with URL output:
export OMNIVOICE_MCP_OUTPUT_MODE=files
export OMNIVOICE_MCP_BASE_PATH=/tmp/mcp_outputs
curl -X POST http://localhost:3900/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","method":"generate_speech","params":{"text":"Hello"},"id":1}'
Clone a voice from base64 audio:
curl -X POST http://localhost:3900/mcp \
-H "Content-Type: application/json" \
-d '{
"jsonrpc":"2.0",
"method":"clone_voice",
"params":{"ref_audio_base64":"<base64-data>"},
"id":2
}'
Transcribe with file path (requires configured base path):
export OMNIVOICE_MCP_BASE_PATH=/tmp/mcp_inputs
# Ensure /tmp/mcp_inputs/sample.wav exists
curl -X POST http://localhost:3900/mcp \
-H "Content-Type: application/json" \
-d '{
"jsonrpc":"2.0",
"method":"transcribe",
"params":{"audio_path":"sample.wav"},
"id":3
}'
Alternative REST endpoints exist for specific queries:
# List available voices
curl http://localhost:3900/api/mcp/voices
# Check server health
curl http://localhost:3900/api/mcp/health
Frontend Integration Components
The VoiceStudio frontend provides UI components for managing MCP agent bindings. The MCPBindingsPanel.jsx component located at frontend/src/components/settings/MCPBindingsPanel.jsx allows users to bind specific agent client IDs to voice profiles. This panel integrates into the main Settings view through frontend/src/pages/Settings.jsx, creating a complete interface for configuring how AI agents interact with the MCP server.
Summary
- The VoiceStudio MCP server exposes seven JSON-RPC tools via the
/mcpendpoint:generate_speech,clone_voice,transcribe,list_voices,list_languages,list_personalities, andcheck_health. - Tool implementations reside in
backend/mcp_server.pywith user documentation indocs/mcp.md. - Environment variables
OMNIVOICE_MCP_OUTPUT_MODEandOMNIVOICE_MCP_BASE_PATHcontrol response formatting and file system security boundaries. - AI agents interact via standard HTTP POST requests with JSON-RPC 2.0 payloads, with alternative REST endpoints available for voice listing and health checks.
- The React frontend provides configuration panels in
MCPBindingsPanel.jsxfor managing agent-to-voice bindings.
Frequently Asked Questions
What is the VoiceStudio MCP server and where is it mounted?
The VoiceStudio MCP server is a Model Context Protocol implementation bundled with the VoiceStudio repository. It mounts automatically on the backend at the /mcp HTTP endpoint, providing JSON-RPC access to voice synthesis and transcription capabilities as defined in backend/mcp_server.py.
How do I switch between Base64 and file URL audio outputs?
Set the OMNIVOICE_MCP_OUTPUT_MODE environment variable to resources for Base64-encoded audio, files for URL/path references, or both for simultaneous return. This variable controls how the generate_speech tool structures its response payload.
What security measures protect file system access in the VoiceStudio MCP server?
The server uses OMNIVOICE_MCP_BASE_PATH as a mandatory security boundary. All file operations using audio_path parameters must reference files within this directory. The server rejects paths attempting to traverse outside this base directory, preventing unauthorized file access.
Can the VoiceStudio MCP server clone voices from reference audio?
Yes. The clone_voice tool accepts Base64-encoded reference audio via the ref_audio_base64 parameter and returns a unique profile_id. This identifier can then be passed to generate_speech requests to synthesize speech using the cloned voice characteristics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →