# VoiceStudio MCP Server: Complete Tool Reference and API Guide

> Explore the VoiceStudio MCP server's seven JSON-RPC tools for speech synthesis, voice cloning, and transcription. Access the complete tool reference and API guide for AI agent integration.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: api-reference
- Published: 2026-09-09

---

**The VoiceStudio MCP server exposes seven JSON-RPC tools—`generate_speech`, `clone_voice`, `transcribe`, `list_voices`, `list_languages`, `list_personalities`, and `check_health`—that enable AI agents to synthesize speech, clone voices, and transcribe audio via the `/mcp` HTTP endpoint.**

The VoiceStudio MCP server implements the **Model Context Protocol (MCP)** to bridge AI agents with voice synthesis capabilities. Bundled with the VoiceStudio repository, this server mounts automatically on the backend at the `/mcp` endpoint, providing a JSON-RPC/HTTP interface for text-to-speech generation, voice cloning, and audio transcription as implemented in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py).

## Available Tools on the VoiceStudio MCP Server

The server defines seven primary tool functions that agents can invoke via JSON-RPC payloads. These are implemented in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) (lines 8-15) and documented in [`docs/mcp.md`](https://github.com/debpalash/VoiceStudio/blob/main/docs/mcp.md) (lines 11-16).

### generate_speech

The `generate_speech` tool converts text or voice-design prompts into WAV audio clips. It accepts a `text` parameter and optionally a `voice_profile` to control synthesis characteristics. Depending on the `OMNIVOICE_MCP_OUTPUT_MODE` environment variable, it returns either Base64-encoded audio (`resources`), a file URL (`files`), or both (`both`).

### clone_voice

The `clone_voice` tool creates new voice profiles from reference audio. It accepts `ref_audio_base64` (a Base64-encoded audio clip) and returns a unique `profile_id` that can be passed to subsequent `generate_speech` calls. This enables personalized voice synthesis without pre-training.

### transcribe

The `transcribe` tool performs speech-to-text conversion, supporting 646 languages. It accepts either `audio_path` (relative to `OMNIVOICE_MCP_BASE_PATH`) or `audio_base64` for direct input. The tool returns plain-text transcription of the provided audio content.

### list_voices

The `list_voices` tool enumerates all stored voice profiles in the system. It returns an array of voice-profile metadata including identifiers and configuration parameters. Agents can query this to present voice selection options to users.

### list_languages

The `list_languages` tool retrieves all TTS languages supported by the backend synthesis engine. It returns an array of language identifiers that can be used when configuring transcription or generation parameters.

### list_personalities

The `list_personalities` tool exposes preset voice personalities configured in the system. It returns an array of personality descriptors that agents can reference to apply predefined speaking styles to generated speech.

### check_health

The `check_health` tool reports backend health status and active GPU availability. It returns JSON fields indicating service readiness and hardware utilization, useful for agent orchestration and monitoring.

## Environment Configuration and Security

The VoiceStudio MCP server respects two critical environment variables defined in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) that control output handling and file system security.

**OMNIVOICE_MCP_OUTPUT_MODE** controls audio response format:

- `resources`: Returns Base64-encoded WAV data inline (default)
- `files`: Returns a URL and file path to the generated audio
- `both`: Returns both Base64 data and file references

**OMNIVOICE_MCP_BASE_PATH** defines the security boundary for file operations. Agents may only read from or write to paths within this directory. This prevents unauthorized file system access when processing `audio_path` arguments.

## How to Invoke VoiceStudio MCP Server Tools

Tools are invoked via HTTP POST requests to `http://localhost:3900/mcp` with JSON-RPC 2.0 payloads.

Generate speech with default resources mode:

```bash
curl -X POST http://localhost:3900/mcp \
  -H "Content-Type: application/json" \
  -d '{
        "jsonrpc":"2.0",
        "method":"generate_speech",
        "params":{"text":"Hello, world!"},
        "id":1
      }'

```

Configure files mode and generate with URL output:

```bash
export OMNIVOICE_MCP_OUTPUT_MODE=files
export OMNIVOICE_MCP_BASE_PATH=/tmp/mcp_outputs
curl -X POST http://localhost:3900/mcp \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","method":"generate_speech","params":{"text":"Hello"},"id":1}'

```

Clone a voice from base64 audio:

```bash
curl -X POST http://localhost:3900/mcp \
  -H "Content-Type: application/json" \
  -d '{
        "jsonrpc":"2.0",
        "method":"clone_voice",
        "params":{"ref_audio_base64":"<base64-data>"},
        "id":2
      }'

```

Transcribe with file path (requires configured base path):

```bash
export OMNIVOICE_MCP_BASE_PATH=/tmp/mcp_inputs

# Ensure /tmp/mcp_inputs/sample.wav exists

curl -X POST http://localhost:3900/mcp \
  -H "Content-Type: application/json" \
  -d '{
        "jsonrpc":"2.0",
        "method":"transcribe",
        "params":{"audio_path":"sample.wav"},
        "id":3
      }'

```

Alternative REST endpoints exist for specific queries:

```bash

# List available voices

curl http://localhost:3900/api/mcp/voices

# Check server health

curl http://localhost:3900/api/mcp/health

```

## Frontend Integration Components

The VoiceStudio frontend provides UI components for managing MCP agent bindings. The [`MCPBindingsPanel.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/MCPBindingsPanel.jsx) component located at [`frontend/src/components/settings/MCPBindingsPanel.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/frontend/src/components/settings/MCPBindingsPanel.jsx) allows users to bind specific agent client IDs to voice profiles. This panel integrates into the main Settings view through [`frontend/src/pages/Settings.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/frontend/src/pages/Settings.jsx), creating a complete interface for configuring how AI agents interact with the MCP server.

## Summary

- The VoiceStudio MCP server exposes seven JSON-RPC tools via the `/mcp` endpoint: `generate_speech`, `clone_voice`, `transcribe`, `list_voices`, `list_languages`, `list_personalities`, and `check_health`.
- Tool implementations reside in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py) with user documentation in [`docs/mcp.md`](https://github.com/debpalash/VoiceStudio/blob/main/docs/mcp.md).
- Environment variables `OMNIVOICE_MCP_OUTPUT_MODE` and `OMNIVOICE_MCP_BASE_PATH` control response formatting and file system security boundaries.
- AI agents interact via standard HTTP POST requests with JSON-RPC 2.0 payloads, with alternative REST endpoints available for voice listing and health checks.
- The React frontend provides configuration panels in [`MCPBindingsPanel.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/MCPBindingsPanel.jsx) for managing agent-to-voice bindings.

## Frequently Asked Questions

### What is the VoiceStudio MCP server and where is it mounted?

The VoiceStudio MCP server is a Model Context Protocol implementation bundled with the VoiceStudio repository. It mounts automatically on the backend at the `/mcp` HTTP endpoint, providing JSON-RPC access to voice synthesis and transcription capabilities as defined in [`backend/mcp_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/mcp_server.py).

### How do I switch between Base64 and file URL audio outputs?

Set the `OMNIVOICE_MCP_OUTPUT_MODE` environment variable to `resources` for Base64-encoded audio, `files` for URL/path references, or `both` for simultaneous return. This variable controls how the `generate_speech` tool structures its response payload.

### What security measures protect file system access in the VoiceStudio MCP server?

The server uses `OMNIVOICE_MCP_BASE_PATH` as a mandatory security boundary. All file operations using `audio_path` parameters must reference files within this directory. The server rejects paths attempting to traverse outside this base directory, preventing unauthorized file access.

### Can the VoiceStudio MCP server clone voices from reference audio?

Yes. The `clone_voice` tool accepts Base64-encoded reference audio via the `ref_audio_base64` parameter and returns a unique `profile_id`. This identifier can then be passed to `generate_speech` requests to synthesize speech using the cloned voice characteristics.