How TTS/ASR Integration Works with VoxCPM2 and FunASR in OpenMAIC
TLDR: OpenMAIC implements speech processing as modular provider plugins, where VoxCPM2 handles neural text-to-speech with voice cloning capabilities and FunASR manages real-time speech recognition via WebSocket streaming, both coordinated through a centralized workbench synchronization layer.
The OpenMAIC framework (THU-MAIC/OpenMAIC) abstracts Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) as pluggable backend services rather than monolithic features. This architecture enables seamless integration of VoxCPM2 for high-fidelity voice synthesis and FunASR for streaming transcription through a unified provider registry and stage synchronization system.
Provider Architecture and Registration
Factory Pattern and Provider IDs
In lib/workbench/provider-registry.ts, each speech service registers a unique identifier during application initialization. VoxCPM2 registers as voxcpm-tts while FunASR registers as fun-asr, each exposing a factory function that instantiates a protocol-specific client bound to user-configured endpoints. This registration pattern allows the workbench to instantiate the correct client implementation at runtime without hardcoding backend-specific logic.
Configuration Interface
The React component at components/Settings/ProviderConfig.tsx renders backend selection controls for both TTS and ASR modules. Users select the backend type—such as Python API, vLLM-Omni, or REST—and supply the Base URL (e.g., http://localhost:8000 for VoxCPM2 or http://host.docker.internal:8188 for FunASR). The component persists these values to the global settingsStore, which tts-stage-sync.ts and asr-stage-sync.ts query to determine active providers.
VoxCPM2 Text-to-Speech Integration
Backend Deployment Options
VoxCPM2 supports three serving strategies compatible with OpenMAIC: the official Python FastAPI server, vLLM-Omni, and Nano-vLLM lightweight instances. A typical production deployment launches the model endpoint using:
python -m vllm_omni.server --model openbmb/VoxCPM2 --port 8000
The OpenMAIC client expects the /tts/upload endpoint to be available at this base URL for voice registration and subsequent synthesis requests.
Voice Cloning and Reference Management
The lib/workbench/voxcpm-voices.ts module implements the voice cloning pipeline. When a user records or uploads a reference clip (constrained to ≤60 seconds and ≤10 MiB), the client uploads the audio to /tts/upload and caches the returned voiceId in IndexedDB via ReferenceClipStore. This identifier is then attached to every synthesis request, enabling zero-shot voice cloning without repeated uploads.
Synthesis Request Flow
The tts-stage-sync.ts orchestrator queries settingsStore.ttsProviderId to resolve the active TTS backend. It constructs a request payload containing the target text, cached voiceId, and selected style parameter (one of natural, soft, or authoritative), then forwards it to the VoxCPM2 client. The client returns an audio Blob, which the workbench dispatches to the playback engine.
// Register a VoxCPM2 voice (one-time setup)
import { registerVoxCPMVoice } from '@/lib/workbench/voxcpm-voices';
await registerVoxCPMVoice({
baseUrl: 'http://localhost:8000',
referenceAudioBase64: btoa(await fetch('ref.wav').then(r => r.arrayBuffer())),
});
// Synthesize speech with the selected voice
import { synthesize } from '@/lib/workbench/tts-stage-sync';
const audioBlob = await synthesize({
text: 'Welcome to the classroom!',
voiceId: 'teacher-voice',
style: 'soft',
});
playAudioBlob(audioBlob);
FunASR Speech-to-Text Integration
WebSocket Streaming Protocol
The lib/workbench/fun-asr.ts module implements a WebSocket client that connects to the /asr/stream endpoint. It transmits raw PCM audio chunks and receives JSON messages structured as {type: 'partial'|'final', text: string}. This streaming architecture eliminates polling overhead and enables real-time transcription display as the user speaks.
Audio Chunking and Latency Management
In lib/workbench/asr-stage-sync.ts, the system captures microphone input via the Web Audio API and buffers audio into 20ms chunks before transmission. This default configuration maintains end-to-end latency ≤150ms, with the chunkMs parameter exposed for adjustment based on network conditions or specific microphone hardware characteristics.
// Start a FunASR session
import { startFunASR } from '@/lib/workbench/fun-asr';
const asr = await startFunASR({
baseUrl: 'http://host.docker.internal:8188',
language: 'en-US',
});
asr.on('partial', txt => console.log('Partial:', txt));
asr.on('final', txt => console.log('Final:', txt));
asr.start(); // begins microphone capture
State Synchronization Across the Workbench
Both TTS and ASR pipelines update lib/workbench/session-store.ts to broadcast audio and transcription states. This central store guarantees that the message composer, replay engine, and lesson interface maintain consistency with ongoing speech synthesis and recognition events, ensuring that voice-cloned TTS output and live transcription remain synchronized with the active session context.
Summary
- OpenMAIC uses a provider registry pattern in
provider-registry.tsto abstract VoxCPM2 and FunASR as interchangeable backends - VoxCPM2 integration supports voice cloning through reference clip uploads to
/tts/uploadand offers three voice styles: natural, soft, and authoritative - FunASR employs WebSocket streaming via
fun-asr.tswith configurable 20ms chunk buffering to maintain sub-150ms latency - The
tts-stage-sync.tsandasr-stage-sync.tsmodules route requests and manage audio state transitions - Centralized session store synchronization keeps speech data consistent across the entire workbench interface
Frequently Asked Questions
What audio format and size limits apply to VoxCPM2 voice cloning?
VoxCPM2 accepts reference clips up to 60 seconds in duration and 10 MiB in file size. The voxcpm-voices.ts module handles base64 encoding of the audio data and persists the resulting voiceId in IndexedDB via ReferenceClipStore for reuse across synthesis sessions.
How does OpenMAIC handle FunASR connection interruptions?
The fun-asr.ts client implements exponential backoff retry logic. When the WebSocket returns errors such as token limits or network timeouts, the client automatically attempts reconnection and surfaces error notifications as toast messages in the UI without crashing the transcription session.
Can I switch between different TTS backends without restarting OpenMAIC?
Yes. The system reads settingsStore.ttsProviderId at runtime through the tts-stage-sync.ts router, allowing users to toggle between VoxCPM2, vLLM-Omni, or generic REST API backends via the ProviderConfig.tsx interface without requiring an application restart.
What is the expected latency for FunASR real-time transcription?
The default configuration maintains end-to-end latency under 150 milliseconds by transmitting 20-millisecond audio chunks. You can adjust the chunkMs parameter in asr-stage-sync.ts to optimize the trade-off between network overhead and transcription responsiveness for specific deployment environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →