# How to Integrate with the VoiceStudio MCP Server: A Complete Guide

> Integrate with the VoiceStudio MCP server easily. Start the standalone process or embed into FastAPI. Configure environment variables and invoke voice synthesis or transcription via JSON-RPC.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-09

---

**To integrate with the VoiceStudio MCP server, start the standalone process with `python -m backend.mcp_server` or embed it into a FastAPI app using `mount_mcp()`, configure the `OMNIVOICE_*` environment variables, and invoke voice synthesis or transcription tools via JSON-RPC.**

The VoiceStudio MCP (Multi-Client Protocol) server is a FastMCP-based service from the [debpalash/VoiceStudio](https://github.com/debpalash/VoiceStudio) repository that exposes voice generation and transcription capabilities to AI agents. This lightweight JSON-RPC transport enables programmatic access to features like text-to-speech, voice cloning, and audio transcription through a standardized interface. Whether you run it as a standalone service or embed it within a larger application, integrating with the VoiceStudio MCP server requires understanding its startup modes, configuration options, and tool invocation patterns.

## Starting the VoiceStudio MCP Server

The server provides two deployment patterns depending on your infrastructure requirements.

### Standalone Mode

Run the MCP server as an independent process using the stdio transport. This mode is ideal for command-line agents or local development workflows.

```bash
python -m backend.mcp_server

```

In this mode, defined in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) within the `main()` function (lines 68-96), the server listens on standard input/output and processes JSON-RPC messages directly. Agents communicate through the provided stdio shim located in [`backend/mcp_shim/__main__.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_shim/__main__.py).

### Embedded Mode with FastAPI

Mount the MCP server as a sub-application within an existing FastAPI service. This approach consolidates your voice API with other HTTP endpoints.

```python
from fastapi import FastAPI
from backend.mcp_server import mount_mcp

app = FastAPI()
if mount_mcp(app):
    print("MCP successfully mounted at /mcp")
else:
    print("MCP disabled – fallback to normal API")

```

The `mount_mcp(app)` function in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) (lines 47-66) attaches the FastMCP HTTP sub-application at the `/mcp` endpoint and stores the session manager on `app.state`. The core server logic is constructed in `create_mcp_server()` (lines 50-60), which initializes the FastMCP instance with proper tool registration.

## Configuring Environment Variables

The VoiceStudio MCP server behavior is controlled through the `OMNIVOICE_*` environment variable family, defined in the configuration section of [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py).

### Output Mode Settings

Set `OMNIVOICE_MCP_OUTPUT_MODE` to determine how audio data returns to clients (lines 99-115):

- **`resources`**: Returns base64-encoded WAV data directly in the JSON response
- **`files`**: Returns a URL reference to the generated audio file
- **`both`**: Returns both base64 data and the file URL

### Secure File Paths

Define `OMNIVOICE_MCP_BASE_PATH` to establish a secure directory boundary for file-based audio inputs and outputs (lines 17-24). This path restriction ensures agents can only read or write audio files within the specified directory, preventing directory traversal attacks when using the `audio_path` parameter in transcription or voice cloning operations.

Additional optional variables include:
- `OMNIVOICE_MCP_TIMEOUT_S`: Request timeout in seconds
- `OMNIVOICE_API_URL`: Backend base URL for voice services
- `OMNIVOICE_MCP_ALLOWED_HOSTS`: Transport-level allow-list for security

## Calling Voice Tools via JSON-RPC

The server registers FastMCP tools using the `@mcp.tool()` decorator. Clients invoke these through JSON-RPC requests.

### generate_speech

Synthesize speech from text input. Located in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) (lines 24-33), this tool accepts parameters including `text`, `language`, `profile_id`, `speed`, and `steps`.

```python
import json, subprocess

payload = {
    "jsonrpc": "2.0",
    "id": 1,
    "method": "generate_speech",
    "params": {
        "text": "Hello, world!",
        "language": "en",
        "profile_id": None,
        "speed": 1.0,
        "steps": 16
    }
}

proc = subprocess.Popen(
    ["python", "-m", "backend.mcp_server"],
    stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True
)
out, _ = proc.communicate(json.dumps(payload) + "\n")
result = json.loads(out)

# Result contains audio_id, timing, and either base64 data or URL depending on OUTPUT_MODE

```

### transcribe

Convert audio to text, supporting both base64 blobs and file paths relative to `OMNIVOICE_MCP_BASE_PATH` (lines 17-28).

```python
import os, json, subprocess

os.environ["OMNIVOICE_MCP_BASE_PATH"] = "/tmp/mcp-data"

payload = {
    "jsonrpc": "2.0",
    "id": 2,
    "method": "transcribe",
    "params": {"audio_path": "sample.wav"}  # Relative to base path

}

proc = subprocess.Popen(
    ["python", "-m", "backend.mcp_server"],
    stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True
)
out, _ = proc.communicate(json.dumps(payload) + "\n")
print(json.loads(out))  # Returns: {"text": "...", "language": "...", "duration": ...}

```

### clone_voice

Create new voice profiles from reference audio (lines 77-86). This tool accepts `name` and audio input via either `audio_base64` or `audio_path` parameters, following the same security model as `transcribe`.

## Voice Profile Resolution

When agents invoke voice tools without explicit profile identifiers, the system resolves the appropriate voice through [`backend/services/mcp_bindings.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/mcp_bindings.py). The `resolve_voice(client_id, explicit_profile_id)` function (lines 104-126) applies a precedence rule:

1. Explicit `profile_id` argument if provided
2. Client-specific binding from the session manager
3. Global default profile

The function returns a dictionary containing `profile_id`, `default_engine`, and a `source` tag for diagnostic logging, ensuring consistent voice selection across MCP sessions.

## Summary

- **Deployment flexibility**: Run the VoiceStudio MCP server standalone via `python -m backend.mcp_server` or embed it using `mount_mcp(app)` in FastAPI
- **Configuration**: Control behavior through `OMNIVOICE_MCP_OUTPUT_MODE`, `OMNIVOICE_MCP_BASE_PATH`, and related environment variables
- **Security**: File path operations are restricted to `OMNIVOICE_MCP_BASE_PATH` to prevent unauthorized filesystem access
- **Toolset**: Access `generate_speech`, `transcribe`, and `clone_voice` through standard JSON-RPC messages
- **Resolution**: Voice profiles are automatically resolved in [`backend/services/mcp_bindings.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/mcp_bindings.py) using client-specific bindings or global defaults

## Frequently Asked Questions

### What transport protocols does the VoiceStudio MCP server support?

The server supports stdio transport for standalone mode and HTTP/SSE when embedded via `mount_mcp()`. The stdio mode uses the shim in [`backend/mcp_shim/__main__.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_shim/__main__.py) to translate between JSON-RPC messages and the FastMCP HTTP endpoint, while embedded mode exposes the native FastMCP sub-application directly at the `/mcp` route.

### How do I secure file access when using the MCP server?

Set the `OMNIVOICE_MCP_BASE_PATH` environment variable to a dedicated directory with restricted permissions. This variable acts as a chroot boundary; any `audio_path` parameters in `transcribe` or `clone_voice` calls must resolve within this directory. Attempts to access paths outside this boundary are rejected by the server logic in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py).

### Can I embed the MCP server in an existing FastAPI application?

Yes. Import `mount_mcp` from `backend.mcp_server` and call `mount_mcp(app)` on your FastAPI instance. This function returns a boolean indicating success and attaches the MCP endpoints at `/mcp`. The session manager is stored in `app.state` for persistence across requests, allowing seamless integration with existing authentication or middleware stacks.

### What is the difference between resources and files output modes?

`OMNIVOICE_MCP_OUTPUT_MODE=resources` returns base64-encoded audio data directly in the JSON response, suitable for immediate processing without filesystem access. `files` mode returns a URL pointer to the generated audio file, reducing payload size for large audio segments. Setting `both` returns both formats, providing flexibility for different client capabilities as implemented in the output handling logic of [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py).