# How `pipecat_minimal.py` Integrates Pipecat with the VoiceStudio MCP Server

> Discover how pipecat_minimal.py integrates with the VoiceStudio MCP server. Learn to connect Pipecat AI pipelines to your local voice processing backend for seamless STT and TTS.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-13

---

**The [`pipecat_minimal.py`](https://github.com/debpalash/VoiceStudio/blob/main/pipecat_minimal.py) example connects Pipecat AI pipelines to VoiceStudio's MCP server by configuring OpenAI-compatible STT and TTS services that point to `http://localhost:3900/v1`, enabling local voice processing through a FastAPI backend.**

VoiceStudio is an open-source voice AI platform that exposes speech-to-text and text-to-speech capabilities through a local MCP (Multi-Channel Provider) server. The [`examples/agentic/pipecat_minimal.py`](https://github.com/debpalash/VoiceStudio/blob/main/examples/agentic/pipecat_minimal.py) script demonstrates the minimal boilerplate required to bridge Pipecat's conversational AI framework with VoiceStudio's self-hosted voice infrastructure, eliminating dependency on external cloud APIs.

## Understanding the VoiceStudio MCP Server Architecture

The VoiceStudio MCP server runs locally as a FastAPI application, exposing OpenAI-compatible HTTP endpoints under the `/v1` namespace. This design allows Pipecat's existing OpenAI service adapters to communicate with VoiceStudio without protocol modifications.

By default, the server listens on port **3900** and implements authentication via an optional `Authorization: Bearer <key>` header. The backend routes requests to internal voice synthesis and speech-to-text engines, returning raw PCM audio at 24 kHz for TTS operations and JSON transcription objects for STT operations.

Key backend files defining this behavior include [`backend/api.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api.py) for the FastAPI application entry point, [`backend/routes/tts.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/tts.py) for `/v1/audio/speech`, and [`backend/routes/stt.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/stt.py) for `/v1/audio/transcriptions`.

## Configuring the Pipecat Integration Endpoint

The integration begins with environment variable configuration in [`examples/agentic/pipecat_minimal.py`](https://github.com/debpalash/VoiceStudio/blob/main/examples/agentic/pipecat_minimal.py). The script reads `OMNIVOICE_API_URL` and `OMNIVOICE_API_KEY` to construct the base URL for all MCP API requests.

```python
import os

OMNIVOICE_BASE_URL = os.environ.get("OMNIVOICE_API_URL", "http://localhost:3900") + "/v1"
OMNIVOICE_API_KEY = os.environ.get("OMNIVOICE_API_KEY", "not-needed-locally")

```

This configuration points Pipecat's service classes to the local VoiceStudio instance. The default key `not-needed-locally` indicates that local deployments skip authentication, while production deployments can enforceBearer token validation through the same header mechanism.

## Constructing OpenAI-Compatible Services

Inside the `build_services()` function, the example lazily imports Pipecat's OpenAI adapters and instantiates two service objects that map Pipecat's internal methods to VoiceStudio's MCP endpoints.

```python
from pipecat.services.openai.stt import OpenAISTTService
from pipecat.services.openai.tts import OpenAITTSService

OMNIVOICE_VOICE = os.environ.get("OMNIVOICE_VOICE", "default")

stt = OpenAISTTService(base_url=OMNIVOICE_BASE_URL, api_key=OMNIVOICE_API_KEY)
tts = OpenAITTSService(
    base_url=OMNIVOICE_BASE_URL,
    api_key=OMNIVOICE_API_KEY,
    voice=OMNIVOICE_VOICE,
    model="omnivoice",
    sample_rate=24000,
)

```

The **`OpenAISTTService`** translates Pipecat's `recognize` calls into `POST` requests to `/v1/audio/transcriptions`, while **`OpenAITTSService`** handles `speak` calls via `POST /v1/audio/speech`. The `model="omnivoice"` parameter identifies the VoiceStudio backend to the service, and `sample_rate=24000` ensures the raw PCM audio stream matches Pipecat's expected format without requiring resampling.

## MCP Server Endpoints and Data Flow

The VoiceStudio MCP server implements specific routes that the Pipecat services consume. Understanding these endpoints clarifies how voice data flows between the pipeline and the local server.

**Speech-to-Text Processing:**
- **Endpoint**: `POST /v1/audio/transcriptions` (defined in [`backend/routes/stt.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/stt.py))
- **Input**: Raw audio bytes or multipart form data
- **Output**: JSON transcription following OpenAI's schema

**Text-to-Speech Synthesis:**
- **Endpoint**: `POST /v1/audio/speech` (defined in [`backend/routes/tts.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/tts.py))
- **Parameters**: `voice` (profile ID), `model`, `input` text
- **Output**: Raw PCM audio stream at 24 kHz

**Voice Enumeration:**
- **Endpoint**: `GET /v1/audio/voices` (defined in [`backend/routes/voices.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/voices.py))
- **Purpose**: Returns available voice profile IDs for the `OMNIVOICE_VOICE` environment variable

The server returns binary audio data directly for TTS operations, which Pipecat processes without additional conversion layers. This zero-copy approach minimizes latency in conversational pipelines.

## Extending the Minimal Example

The [`pipecat_minimal.py`](https://github.com/debpalash/VoiceStudio/blob/main/pipecat_minimal.py) script intentionally stops after constructing the `stt` and `tts` objects, printing diagnostic messages that confirm successful initialization against `OMNIVOICE_BASE_URL`.

```python
print("VoiceStudio STT + TTS services constructed against", OMNIVOICE_BASE_URL)
print("Wire `stt` and `tts` into your pipecat Pipeline with a transport")
print("and an LLM service. See docs/agentic-voice.md.")

```

To build a complete conversational agent, developers must inject these services into a Pipecat `Pipeline` alongside a transport layer (WebSocket, LiveKit, or daily-webrtc) and an LLM service. The minimal example isolates the **MCP integration** concerns, allowing developers to verify VoiceStudio connectivity before adding complexity.

## Summary

- **[`pipecat_minimal.py`](https://github.com/debpalash/VoiceStudio/blob/main/pipecat_minimal.py)** demonstrates the simplest valid integration between Pipecat and VoiceStudio's MCP server using OpenAI-compatible service adapters.
- **Environment variables** `OMNIVOICE_API_URL` and `OMNIVOICE_API_KEY` configure the connection to the local FastAPI backend running on port 3900.
- **Service objects** `OpenAISTTService` and `OpenAITTSService` map Pipecat pipeline methods to VoiceStudio's `/v1/audio/transcriptions` and `/v1/audio/speech` endpoints.
- **Raw PCM audio** at 24 kHz flows directly from VoiceStudio to Pipecat without format conversion, handled by the `sample_rate=24000` parameter.
- **Backend routes** in [`backend/routes/stt.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/stt.py), [`backend/routes/tts.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/tts.py), and [`backend/routes/voices.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/voices.py) implement the OpenAI-compatible schema that enables this integration.

## Frequently Asked Questions

### Does VoiceStudio require an API key for local development?

No. When running VoiceStudio locally, the default configuration accepts the placeholder key `not-needed-locally` or any string value. The MCP server validates Bearer tokens only when explicitly configured for production deployments, making local development frictionless.

### Why does the example use OpenAI service classes to connect to VoiceStudio?

VoiceStudio implements an OpenAI-compatible REST API under the `/v1` namespace, allowing Pipecat's existing `OpenAISTTService` and `OpenAITTSService` to communicate with the local server without custom adapters. This compatibility layer means any Pipecat pipeline built for OpenAI's voice APIs can switch to VoiceStudio by changing only the `base_url` parameter.

### How do I select a specific voice profile in the integration?

Set the `OMNIVOICE_VOICE` environment variable to a voice ID returned by the `GET /v1/audio/voices` endpoint. The `OpenAITTSService` passes this value as the `voice` parameter in TTS requests to VoiceStudio's `/v1/audio/speech` endpoint, which routes the request to the appropriate synthesis engine profile defined in [`backend/routes/voices.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/routes/voices.py).

### Can I use this example with a remote VoiceStudio deployment?

Yes. Replace `http://localhost:3900` in `OMNIVOICE_API_URL` with your remote instance's address and provide the corresponding `OMNIVOICE_API_KEY`. The Pipecat services communicate via standard HTTP, so network topology does not affect the integration pattern, though latency considerations may influence transport layer selection for real-time conversational agents.