How to Integrate VoiceStudio Backend with Other Services: REST, WebSocket, and Custom Router Guide

Integrate VoiceStudio backend with other services by calling its FastAPI REST endpoints at localhost:3900, streaming real-time audio over WebSocket at /ws/tts, or mounting custom routers that inherit shared security dependencies like require_loopback and require_admin.

VoiceStudio is an open-source AI voice and language toolkit built on FastAPI that exposes TTS, LLM, and ASR capabilities through standard HTTP and WebSocket interfaces. To integrate VoiceStudio backend with other services, you can consume the public API directly, extend the server with proprietary endpoints, or configure external LLM providers via environment variables. This guide leverages the actual source code from debpalash/VoiceStudio to demonstrate authentication patterns, request routing, and concrete integration implementations.

Architecture Overview

VoiceStudio's backend follows a modular FastAPI architecture where the main application aggregates multiple routers and shared dependencies. The entry point in [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py) bootstraps the server by configuring lifespan management and mounting API routers. Security is enforced through reusable dependencies defined in [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py), which provide require_loopback and require_admin guards to protect privileged routes.

The system separates concerns into distinct layers:

Component Purpose Key Source File
FastAPI Application Server bootstrap and global middleware [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py)
Security Dependencies API key validation and host restriction [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py)
TTS Streaming Real-time audio delivery via WebSocket [backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py)
LLM Backend Abstraction for language model providers [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py)
Engine Routing GPU/CPU resource allocation logic [backend/services/engine_routing.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/engine_routing.py)

Integration Methods

External services can integrate with VoiceStudio through three primary interfaces: synchronous REST calls, persistent WebSocket streams, or server-side router extensions.

REST API Consumption

The backend exposes standard HTTP endpoints for batch processing under /tts and related paths. These routes accept JSON payloads and return synthesized audio metadata or LLM completions. All production routes inherit security dependencies that validate API keys or restrict access to loopback addresses.

WebSocket Streaming for Real-Time TTS

For low-latency applications, the /ws/tts endpoint in [backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py) delivers raw PCM audio chunks as they are generated. This endpoint requires a persistent WebSocket connection and accepts JSON control messages to specify voice parameters and text input.

Custom Router Mounting

To add proprietary endpoints or bridge external protocols, developers can create FastAPI routers and mount them onto the main application. Custom routers automatically gain access to VoiceStudio's dependency injection system, including the require_admin guard for protected operations.

Quick Start Integration Steps

Follow these steps to integrate VoiceStudio backend with other services:

  1. Start the Backend Server – Run uvicorn backend.main:app --host 0.0.0.0 --port 3900 to initialize the FastAPI application defined in [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py).

  2. Configure Environment Variables – Set OMNIVOICE_LLM_BACKEND, TRANSLATE_BASE_URL, and TRANSLATE_API_KEY to connect external LLM services through [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py).

  3. Consume the API – Use HTTP clients for REST endpoints or WebSocket clients for real-time streaming, respecting the require_loopback and require_admin dependencies from [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py).

  4. Extend with Custom Routers – Mount additional FastAPI routers to app in backend/main.py to expose proprietary endpoints that participate in VoiceStudio's dependency injection system.

Integration Code Examples

Synchronous TTS via REST

Use standard HTTP clients to call the /tts endpoint for batch synthesis. The endpoint accepts JSON payloads containing the text, voice selection, and optional engine overrides.

import requests

BASE_URL = "http://127.0.0.1:3900"

def synthesize_speech(text: str, voice: str = "default"):
    payload = {
        "text": text,
        "voice": voice,
        "speed": 1.0,
        "engine": None
    }
    response = requests.post(f"{BASE_URL}/tts", json=payload)
    response.raise_for_status()
    return response.json()  # Returns audio URL and metadata

audio_info = synthesize_speech("Hello world", voice="en_us_female")
print(audio_info["audio_url"])

Real-Time Audio Streaming via WebSocket

Connect to ws://localhost:3900/ws/tts to receive PCM16 audio chunks at 24kHz in real-time. The protocol requires an initial JSON configuration message followed by binary audio frames.

import asyncio
import json
import websockets

async def stream_tts(text: str):
    uri = "ws://127.0.0.1:3900/ws/tts"
    async with websockets.connect(uri) as websocket:
        await websocket.send(json.dumps({
            "text": text,
            "voice": "en_us_female"
        }))
        
        async for message in websocket:
            if isinstance(message, bytes):
                # Handle PCM16 audio chunk

                with open("output.pcm", "ab") as f:
                    f.write(message)
            else:
                data = json.loads(message)
                if data.get("type") == "done":
                    break
                elif data.get("type") == "error":
                    raise RuntimeError(data["detail"])

asyncio.run(stream_tts("Streaming voice synthesis"))

Configuring External LLM Providers

VoiceStudio's [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py) implements an OpenAICompatBackend class that routes requests to external OpenAI-compatible endpoints. No code changes are required to switch providers; set the environment variables before starting the server:

export OMNIVOICE_LLM_BACKEND=custom
export TRANSLATE_BASE_URL="https://api.external-llm.com/v1"
export TRANSLATE_API_KEY="your-secret-key"

Extending the Backend with Custom Routers

Create additional endpoints by defining FastAPI routers and mounting them in backend/main.py. This approach inherits VoiceStudio's authentication dependencies while allowing custom business logic.


# my_integration.py

from fastapi import APIRouter, Depends
from backend.api.dependencies import require_admin

router = APIRouter()

@router.get("/external-status")
def check_external_service(admin=Depends(require_admin)):
    return {"status": "connected", "service": "external-crm"}

# In backend/main.py, add:

# from my_integration import router as ext_router

# app.include_router(ext_router, prefix="/integrations")

Key Source Files for Integration

Understanding these specific files is essential when you integrate VoiceStudio backend with other services:

Summary

  • Start the server with uvicorn backend.main:app to expose the FastAPI application on port 3900.
  • Call REST endpoints at /tts using standard HTTP clients, passing JSON payloads for batch processing.
  • Stream audio by connecting to the /ws/tts WebSocket endpoint for real-time PCM16 audio delivery.
  • Configure external LLMs via TRANSLATE_BASE_URL and TRANSLATE_API_KEY environment variables to route AI requests to external providers.
  • Extend functionality by mounting custom FastAPI routers that inherit require_admin and require_loopback dependencies from [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py).
  • Secure integrations by respecting the dependency guards that restrict administrative actions to loopback addresses or valid API keys.

Frequently Asked Questions

How do I authenticate requests when integrating with VoiceStudio?

VoiceStudio uses FastAPI dependencies defined in [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py) to enforce security. The require_loopback dependency restricts sensitive endpoints to localhost connections, while require_admin validates API keys against configured secrets. When calling from external services, ensure your requests include the Authorization header with a valid admin key, or deploy a custom router that relaxes these restrictions for specific trusted networks.

Can I use VoiceStudio with my own LLM provider instead of the default backend?

Yes. The [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py) file implements the OpenAICompatBackend class, which supports any OpenAI-compatible API. Set the OMNIVOICE_LLM_BACKEND environment variable to custom, provide the TRANSLATE_BASE_URL pointing to your LLM endpoint, and set TRANSLATE_API_KEY for authentication. VoiceStudio will automatically route generation requests to your external provider without requiring code changes.

What is the difference between the REST and WebSocket TTS endpoints?

The REST endpoint at /tts processes complete text inputs and returns audio metadata or file URLs synchronously. The WebSocket endpoint at /ws/tts defined in [backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py) streams raw PCM16 audio chunks at 24kHz as they are synthesized, enabling real-time applications with low latency requirements. Use REST for batch processing and WebSocket for interactive voice assistants.

How do I add a custom endpoint to the VoiceStudio backend?

Create a new FastAPI router in a Python file, define your routes using standard FastAPI decorators, and import the router into [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py). Use app.include_router() to mount your routes at your preferred prefix. Your custom endpoints can import and use the shared dependencies like require_admin from [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py) to maintain consistent security policies with the core VoiceStudio routes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →