How to Integrate with the VoiceStudio MCP Server: A Complete Guide

To integrate with the VoiceStudio MCP server, start the standalone process with python -m backend.mcp_server or embed it into a FastAPI app using mount_mcp(), configure the OMNIVOICE_* environment variables, and invoke voice synthesis or transcription tools via JSON-RPC.

The VoiceStudio MCP (Multi-Client Protocol) server is a FastMCP-based service from the debpalash/VoiceStudio repository that exposes voice generation and transcription capabilities to AI agents. This lightweight JSON-RPC transport enables programmatic access to features like text-to-speech, voice cloning, and audio transcription through a standardized interface. Whether you run it as a standalone service or embed it within a larger application, integrating with the VoiceStudio MCP server requires understanding its startup modes, configuration options, and tool invocation patterns.

Starting the VoiceStudio MCP Server

The server provides two deployment patterns depending on your infrastructure requirements.

Standalone Mode

Run the MCP server as an independent process using the stdio transport. This mode is ideal for command-line agents or local development workflows.

python -m backend.mcp_server

In this mode, defined in backend/mcp_server.py within the main() function (lines 68-96), the server listens on standard input/output and processes JSON-RPC messages directly. Agents communicate through the provided stdio shim located in backend/mcp_shim/__main__.py.

Embedded Mode with FastAPI

Mount the MCP server as a sub-application within an existing FastAPI service. This approach consolidates your voice API with other HTTP endpoints.

from fastapi import FastAPI
from backend.mcp_server import mount_mcp

app = FastAPI()
if mount_mcp(app):
    print("MCP successfully mounted at /mcp")
else:
    print("MCP disabled – fallback to normal API")

The mount_mcp(app) function in backend/mcp_server.py (lines 47-66) attaches the FastMCP HTTP sub-application at the /mcp endpoint and stores the session manager on app.state. The core server logic is constructed in create_mcp_server() (lines 50-60), which initializes the FastMCP instance with proper tool registration.

Configuring Environment Variables

The VoiceStudio MCP server behavior is controlled through the OMNIVOICE_* environment variable family, defined in the configuration section of backend/mcp_server.py.

Output Mode Settings

Set OMNIVOICE_MCP_OUTPUT_MODE to determine how audio data returns to clients (lines 99-115):

  • resources: Returns base64-encoded WAV data directly in the JSON response
  • files: Returns a URL reference to the generated audio file
  • both: Returns both base64 data and the file URL

Secure File Paths

Define OMNIVOICE_MCP_BASE_PATH to establish a secure directory boundary for file-based audio inputs and outputs (lines 17-24). This path restriction ensures agents can only read or write audio files within the specified directory, preventing directory traversal attacks when using the audio_path parameter in transcription or voice cloning operations.

Additional optional variables include:

  • OMNIVOICE_MCP_TIMEOUT_S: Request timeout in seconds
  • OMNIVOICE_API_URL: Backend base URL for voice services
  • OMNIVOICE_MCP_ALLOWED_HOSTS: Transport-level allow-list for security

Calling Voice Tools via JSON-RPC

The server registers FastMCP tools using the @mcp.tool() decorator. Clients invoke these through JSON-RPC requests.

generate_speech

Synthesize speech from text input. Located in backend/mcp_server.py (lines 24-33), this tool accepts parameters including text, language, profile_id, speed, and steps.

import json, subprocess

payload = {
    "jsonrpc": "2.0",
    "id": 1,
    "method": "generate_speech",
    "params": {
        "text": "Hello, world!",
        "language": "en",
        "profile_id": None,
        "speed": 1.0,
        "steps": 16
    }
}

proc = subprocess.Popen(
    ["python", "-m", "backend.mcp_server"],
    stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True
)
out, _ = proc.communicate(json.dumps(payload) + "\n")
result = json.loads(out)

# Result contains audio_id, timing, and either base64 data or URL depending on OUTPUT_MODE

transcribe

Convert audio to text, supporting both base64 blobs and file paths relative to OMNIVOICE_MCP_BASE_PATH (lines 17-28).

import os, json, subprocess

os.environ["OMNIVOICE_MCP_BASE_PATH"] = "/tmp/mcp-data"

payload = {
    "jsonrpc": "2.0",
    "id": 2,
    "method": "transcribe",
    "params": {"audio_path": "sample.wav"}  # Relative to base path

}

proc = subprocess.Popen(
    ["python", "-m", "backend.mcp_server"],
    stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True
)
out, _ = proc.communicate(json.dumps(payload) + "\n")
print(json.loads(out))  # Returns: {"text": "...", "language": "...", "duration": ...}

clone_voice

Create new voice profiles from reference audio (lines 77-86). This tool accepts name and audio input via either audio_base64 or audio_path parameters, following the same security model as transcribe.

Voice Profile Resolution

When agents invoke voice tools without explicit profile identifiers, the system resolves the appropriate voice through backend/services/mcp_bindings.py. The resolve_voice(client_id, explicit_profile_id) function (lines 104-126) applies a precedence rule:

  1. Explicit profile_id argument if provided
  2. Client-specific binding from the session manager
  3. Global default profile

The function returns a dictionary containing profile_id, default_engine, and a source tag for diagnostic logging, ensuring consistent voice selection across MCP sessions.

Summary

  • Deployment flexibility: Run the VoiceStudio MCP server standalone via python -m backend.mcp_server or embed it using mount_mcp(app) in FastAPI
  • Configuration: Control behavior through OMNIVOICE_MCP_OUTPUT_MODE, OMNIVOICE_MCP_BASE_PATH, and related environment variables
  • Security: File path operations are restricted to OMNIVOICE_MCP_BASE_PATH to prevent unauthorized filesystem access
  • Toolset: Access generate_speech, transcribe, and clone_voice through standard JSON-RPC messages
  • Resolution: Voice profiles are automatically resolved in backend/services/mcp_bindings.py using client-specific bindings or global defaults

Frequently Asked Questions

What transport protocols does the VoiceStudio MCP server support?

The server supports stdio transport for standalone mode and HTTP/SSE when embedded via mount_mcp(). The stdio mode uses the shim in backend/mcp_shim/__main__.py to translate between JSON-RPC messages and the FastMCP HTTP endpoint, while embedded mode exposes the native FastMCP sub-application directly at the /mcp route.

How do I secure file access when using the MCP server?

Set the OMNIVOICE_MCP_BASE_PATH environment variable to a dedicated directory with restricted permissions. This variable acts as a chroot boundary; any audio_path parameters in transcribe or clone_voice calls must resolve within this directory. Attempts to access paths outside this boundary are rejected by the server logic in backend/mcp_server.py.

Can I embed the MCP server in an existing FastAPI application?

Yes. Import mount_mcp from backend.mcp_server and call mount_mcp(app) on your FastAPI instance. This function returns a boolean indicating success and attaches the MCP endpoints at /mcp. The session manager is stored in app.state for persistence across requests, allowing seamless integration with existing authentication or middleware stacks.

What is the difference between resources and files output modes?

OMNIVOICE_MCP_OUTPUT_MODE=resources returns base64-encoded audio data directly in the JSON response, suitable for immediate processing without filesystem access. files mode returns a URL pointer to the generated audio file, reducing payload size for large audio segments. Setting both returns both formats, providing flexibility for different client capabilities as implemented in the output handling logic of backend/mcp_server.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →