# How TTS/ASR Integration Works with VoxCPM2 and FunASR in OpenMAIC

> Learn how OpenMAIC integrates TTS/ASR using VoxCPM2 for voice cloning and FunASR for real-time recognition via WebSocket. Discover modular speech processing in action.

- Repository: [MAIC/OpenMAIC](https://github.com/THU-MAIC/OpenMAIC)
- Tags: deep-dive
- Published: 2026-09-13

---

**TLDR:** OpenMAIC implements speech processing as modular provider plugins, where VoxCPM2 handles neural text-to-speech with voice cloning capabilities and FunASR manages real-time speech recognition via WebSocket streaming, both coordinated through a centralized workbench synchronization layer.

The OpenMAIC framework (THU-MAIC/OpenMAIC) abstracts Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) as pluggable backend services rather than monolithic features. This architecture enables seamless integration of VoxCPM2 for high-fidelity voice synthesis and FunASR for streaming transcription through a unified provider registry and stage synchronization system.

## Provider Architecture and Registration

### Factory Pattern and Provider IDs

In [`lib/workbench/provider-registry.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/provider-registry.ts), each speech service registers a unique identifier during application initialization. VoxCPM2 registers as `voxcpm-tts` while FunASR registers as `fun-asr`, each exposing a factory function that instantiates a protocol-specific client bound to user-configured endpoints. This registration pattern allows the workbench to instantiate the correct client implementation at runtime without hardcoding backend-specific logic.

### Configuration Interface

The React component at [`components/Settings/ProviderConfig.tsx`](https://github.com/THU-MAIC/OpenMAIC/blob/main/components/Settings/ProviderConfig.tsx) renders backend selection controls for both TTS and ASR modules. Users select the backend type—such as *Python API*, *vLLM-Omni*, or *REST*—and supply the **Base URL** (e.g., `http://localhost:8000` for VoxCPM2 or `http://host.docker.internal:8188` for FunASR). The component persists these values to the global `settingsStore`, which [`tts-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tts-stage-sync.ts) and [`asr-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/asr-stage-sync.ts) query to determine active providers.

## VoxCPM2 Text-to-Speech Integration

### Backend Deployment Options

VoxCPM2 supports three serving strategies compatible with OpenMAIC: the official Python FastAPI server, *vLLM-Omni*, and *Nano-vLLM* lightweight instances. A typical production deployment launches the model endpoint using:

```bash
python -m vllm_omni.server --model openbmb/VoxCPM2 --port 8000

```

The OpenMAIC client expects the `/tts/upload` endpoint to be available at this base URL for voice registration and subsequent synthesis requests.

### Voice Cloning and Reference Management

The [`lib/workbench/voxcpm-voices.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/voxcpm-voices.ts) module implements the voice cloning pipeline. When a user records or uploads a reference clip (constrained to ≤60 seconds and ≤10 MiB), the client uploads the audio to `/tts/upload` and caches the returned `voiceId` in IndexedDB via `ReferenceClipStore`. This identifier is then attached to every synthesis request, enabling zero-shot voice cloning without repeated uploads.

### Synthesis Request Flow

The [`tts-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tts-stage-sync.ts) orchestrator queries `settingsStore.ttsProviderId` to resolve the active TTS backend. It constructs a request payload containing the target `text`, cached `voiceId`, and selected `style` parameter (one of *natural*, *soft*, or *authoritative*), then forwards it to the VoxCPM2 client. The client returns an audio **Blob**, which the workbench dispatches to the playback engine.

```typescript
// Register a VoxCPM2 voice (one-time setup)
import { registerVoxCPMVoice } from '@/lib/workbench/voxcpm-voices';

await registerVoxCPMVoice({
  baseUrl: 'http://localhost:8000',
  referenceAudioBase64: btoa(await fetch('ref.wav').then(r => r.arrayBuffer())),
});

```

```typescript
// Synthesize speech with the selected voice
import { synthesize } from '@/lib/workbench/tts-stage-sync';

const audioBlob = await synthesize({
  text: 'Welcome to the classroom!',
  voiceId: 'teacher-voice',
  style: 'soft',
});
playAudioBlob(audioBlob);

```

## FunASR Speech-to-Text Integration

### WebSocket Streaming Protocol

The [`lib/workbench/fun-asr.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/fun-asr.ts) module implements a WebSocket client that connects to the `/asr/stream` endpoint. It transmits raw PCM audio chunks and receives JSON messages structured as `{type: 'partial'|'final', text: string}`. This streaming architecture eliminates polling overhead and enables real-time transcription display as the user speaks.

### Audio Chunking and Latency Management

In [`lib/workbench/asr-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/asr-stage-sync.ts), the system captures microphone input via the Web Audio API and buffers audio into **20ms chunks** before transmission. This default configuration maintains end-to-end latency **≤150ms**, with the `chunkMs` parameter exposed for adjustment based on network conditions or specific microphone hardware characteristics.

```typescript
// Start a FunASR session
import { startFunASR } from '@/lib/workbench/fun-asr';

const asr = await startFunASR({
  baseUrl: 'http://host.docker.internal:8188',
  language: 'en-US',
});

asr.on('partial', txt => console.log('Partial:', txt));
asr.on('final', txt => console.log('Final:', txt));
asr.start();   // begins microphone capture

```

## State Synchronization Across the Workbench

Both TTS and ASR pipelines update [`lib/workbench/session-store.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/session-store.ts) to broadcast audio and transcription states. This central store guarantees that the message composer, replay engine, and lesson interface maintain consistency with ongoing speech synthesis and recognition events, ensuring that voice-cloned TTS output and live transcription remain synchronized with the active session context.

## Summary

- OpenMAIC uses a provider registry pattern in [`provider-registry.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/provider-registry.ts) to abstract VoxCPM2 and FunASR as interchangeable backends
- VoxCPM2 integration supports voice cloning through reference clip uploads to `/tts/upload` and offers three voice styles: natural, soft, and authoritative
- FunASR employs WebSocket streaming via [`fun-asr.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/fun-asr.ts) with configurable 20ms chunk buffering to maintain sub-150ms latency
- The [`tts-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tts-stage-sync.ts) and [`asr-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/asr-stage-sync.ts) modules route requests and manage audio state transitions
- Centralized session store synchronization keeps speech data consistent across the entire workbench interface

## Frequently Asked Questions

### What audio format and size limits apply to VoxCPM2 voice cloning?

VoxCPM2 accepts reference clips up to 60 seconds in duration and 10 MiB in file size. The [`voxcpm-voices.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/voxcpm-voices.ts) module handles base64 encoding of the audio data and persists the resulting `voiceId` in IndexedDB via `ReferenceClipStore` for reuse across synthesis sessions.

### How does OpenMAIC handle FunASR connection interruptions?

The [`fun-asr.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/fun-asr.ts) client implements exponential backoff retry logic. When the WebSocket returns errors such as token limits or network timeouts, the client automatically attempts reconnection and surfaces error notifications as toast messages in the UI without crashing the transcription session.

### Can I switch between different TTS backends without restarting OpenMAIC?

Yes. The system reads `settingsStore.ttsProviderId` at runtime through the [`tts-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tts-stage-sync.ts) router, allowing users to toggle between VoxCPM2, vLLM-Omni, or generic REST API backends via the [`ProviderConfig.tsx`](https://github.com/THU-MAIC/OpenMAIC/blob/main/ProviderConfig.tsx) interface without requiring an application restart.

### What is the expected latency for FunASR real-time transcription?

The default configuration maintains end-to-end latency under 150 milliseconds by transmitting 20-millisecond audio chunks. You can adjust the `chunkMs` parameter in [`asr-stage-sync.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/asr-stage-sync.ts) to optimize the trade-off between network overhead and transcription responsiveness for specific deployment environments.