How to Integrate VoiceStudio Backend with Other Services: REST, WebSocket, and Custom Router Guide
Integrate VoiceStudio backend with other services by calling its FastAPI REST endpoints at localhost:3900, streaming real-time audio over WebSocket at /ws/tts, or mounting custom routers that inherit shared security dependencies like require_loopback and require_admin.
VoiceStudio is an open-source AI voice and language toolkit built on FastAPI that exposes TTS, LLM, and ASR capabilities through standard HTTP and WebSocket interfaces. To integrate VoiceStudio backend with other services, you can consume the public API directly, extend the server with proprietary endpoints, or configure external LLM providers via environment variables. This guide leverages the actual source code from debpalash/VoiceStudio to demonstrate authentication patterns, request routing, and concrete integration implementations.
Architecture Overview
VoiceStudio's backend follows a modular FastAPI architecture where the main application aggregates multiple routers and shared dependencies. The entry point in [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py) bootstraps the server by configuring lifespan management and mounting API routers. Security is enforced through reusable dependencies defined in [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py), which provide require_loopback and require_admin guards to protect privileged routes.
The system separates concerns into distinct layers:
| Component | Purpose | Key Source File |
|---|---|---|
| FastAPI Application | Server bootstrap and global middleware | [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py) |
| Security Dependencies | API key validation and host restriction | [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py) |
| TTS Streaming | Real-time audio delivery via WebSocket | [backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py) |
| LLM Backend | Abstraction for language model providers | [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py) |
| Engine Routing | GPU/CPU resource allocation logic | [backend/services/engine_routing.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/engine_routing.py) |
Integration Methods
External services can integrate with VoiceStudio through three primary interfaces: synchronous REST calls, persistent WebSocket streams, or server-side router extensions.
REST API Consumption
The backend exposes standard HTTP endpoints for batch processing under /tts and related paths. These routes accept JSON payloads and return synthesized audio metadata or LLM completions. All production routes inherit security dependencies that validate API keys or restrict access to loopback addresses.
WebSocket Streaming for Real-Time TTS
For low-latency applications, the /ws/tts endpoint in [backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py) delivers raw PCM audio chunks as they are generated. This endpoint requires a persistent WebSocket connection and accepts JSON control messages to specify voice parameters and text input.
Custom Router Mounting
To add proprietary endpoints or bridge external protocols, developers can create FastAPI routers and mount them onto the main application. Custom routers automatically gain access to VoiceStudio's dependency injection system, including the require_admin guard for protected operations.
Quick Start Integration Steps
Follow these steps to integrate VoiceStudio backend with other services:
-
Start the Backend Server – Run
uvicorn backend.main:app --host 0.0.0.0 --port 3900to initialize the FastAPI application defined in [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py). -
Configure Environment Variables – Set
OMNIVOICE_LLM_BACKEND,TRANSLATE_BASE_URL, andTRANSLATE_API_KEYto connect external LLM services through [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py). -
Consume the API – Use HTTP clients for REST endpoints or WebSocket clients for real-time streaming, respecting the
require_loopbackandrequire_admindependencies from [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py). -
Extend with Custom Routers – Mount additional FastAPI routers to
appinbackend/main.pyto expose proprietary endpoints that participate in VoiceStudio's dependency injection system.
Integration Code Examples
Synchronous TTS via REST
Use standard HTTP clients to call the /tts endpoint for batch synthesis. The endpoint accepts JSON payloads containing the text, voice selection, and optional engine overrides.
import requests
BASE_URL = "http://127.0.0.1:3900"
def synthesize_speech(text: str, voice: str = "default"):
payload = {
"text": text,
"voice": voice,
"speed": 1.0,
"engine": None
}
response = requests.post(f"{BASE_URL}/tts", json=payload)
response.raise_for_status()
return response.json() # Returns audio URL and metadata
audio_info = synthesize_speech("Hello world", voice="en_us_female")
print(audio_info["audio_url"])
Real-Time Audio Streaming via WebSocket
Connect to ws://localhost:3900/ws/tts to receive PCM16 audio chunks at 24kHz in real-time. The protocol requires an initial JSON configuration message followed by binary audio frames.
import asyncio
import json
import websockets
async def stream_tts(text: str):
uri = "ws://127.0.0.1:3900/ws/tts"
async with websockets.connect(uri) as websocket:
await websocket.send(json.dumps({
"text": text,
"voice": "en_us_female"
}))
async for message in websocket:
if isinstance(message, bytes):
# Handle PCM16 audio chunk
with open("output.pcm", "ab") as f:
f.write(message)
else:
data = json.loads(message)
if data.get("type") == "done":
break
elif data.get("type") == "error":
raise RuntimeError(data["detail"])
asyncio.run(stream_tts("Streaming voice synthesis"))
Configuring External LLM Providers
VoiceStudio's [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py) implements an OpenAICompatBackend class that routes requests to external OpenAI-compatible endpoints. No code changes are required to switch providers; set the environment variables before starting the server:
export OMNIVOICE_LLM_BACKEND=custom
export TRANSLATE_BASE_URL="https://api.external-llm.com/v1"
export TRANSLATE_API_KEY="your-secret-key"
Extending the Backend with Custom Routers
Create additional endpoints by defining FastAPI routers and mounting them in backend/main.py. This approach inherits VoiceStudio's authentication dependencies while allowing custom business logic.
# my_integration.py
from fastapi import APIRouter, Depends
from backend.api.dependencies import require_admin
router = APIRouter()
@router.get("/external-status")
def check_external_service(admin=Depends(require_admin)):
return {"status": "connected", "service": "external-crm"}
# In backend/main.py, add:
# from my_integration import router as ext_router
# app.include_router(ext_router, prefix="/integrations")
Key Source Files for Integration
Understanding these specific files is essential when you integrate VoiceStudio backend with other services:
- [
backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py) – Application factory that creates the FastAPI instance and mounts all routers. - [
backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py) – Definesrequire_loopbackandrequire_adminfor route protection. - [
backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py) – WebSocket endpoint implementation for streaming TTS. - [
backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py) – Backend abstraction and provider selection logic. - [
backend/services/engine_routing.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/engine_routing.py) – Determines whether requests execute on local GPU, remote GPU, or CPU.
Summary
- Start the server with
uvicorn backend.main:appto expose the FastAPI application on port 3900. - Call REST endpoints at
/ttsusing standard HTTP clients, passing JSON payloads for batch processing. - Stream audio by connecting to the
/ws/ttsWebSocket endpoint for real-time PCM16 audio delivery. - Configure external LLMs via
TRANSLATE_BASE_URLandTRANSLATE_API_KEYenvironment variables to route AI requests to external providers. - Extend functionality by mounting custom FastAPI routers that inherit
require_adminandrequire_loopbackdependencies from [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py). - Secure integrations by respecting the dependency guards that restrict administrative actions to loopback addresses or valid API keys.
Frequently Asked Questions
How do I authenticate requests when integrating with VoiceStudio?
VoiceStudio uses FastAPI dependencies defined in [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py) to enforce security. The require_loopback dependency restricts sensitive endpoints to localhost connections, while require_admin validates API keys against configured secrets. When calling from external services, ensure your requests include the Authorization header with a valid admin key, or deploy a custom router that relaxes these restrictions for specific trusted networks.
Can I use VoiceStudio with my own LLM provider instead of the default backend?
Yes. The [backend/services/llm_backend.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/llm_backend.py) file implements the OpenAICompatBackend class, which supports any OpenAI-compatible API. Set the OMNIVOICE_LLM_BACKEND environment variable to custom, provide the TRANSLATE_BASE_URL pointing to your LLM endpoint, and set TRANSLATE_API_KEY for authentication. VoiceStudio will automatically route generation requests to your external provider without requiring code changes.
What is the difference between the REST and WebSocket TTS endpoints?
The REST endpoint at /tts processes complete text inputs and returns audio metadata or file URLs synchronously. The WebSocket endpoint at /ws/tts defined in [backend/api/routers/tts_stream.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/tts_stream.py) streams raw PCM16 audio chunks at 24kHz as they are synthesized, enabling real-time applications with low latency requirements. Use REST for batch processing and WebSocket for interactive voice assistants.
How do I add a custom endpoint to the VoiceStudio backend?
Create a new FastAPI router in a Python file, define your routes using standard FastAPI decorators, and import the router into [backend/main.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/main.py). Use app.include_router() to mount your routes at your preferred prefix. Your custom endpoints can import and use the shared dependencies like require_admin from [backend/api/dependencies.py](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/dependencies.py) to maintain consistent security policies with the core VoiceStudio routes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →