Where to Find the Engine Capability Matrix in VoiceStudio: Documentation and API Guide
The engine capability matrix is documented in docs/specs/01-expressive-tts.md (section 4.4) and programmatically accessible via the list_backends() API endpoint.
VoiceStudio is an open-source expressive text-to-speech platform that maintains a detailed engine capability matrix to track TTS backend features. This matrix maps which expression mechanisms, emotion controls, and reference-clip capabilities are available across supported engines like omnivoice, cosyvoice, indextts2, and voxcpm2.
Location of the Engine Capability Matrix in VoiceStudio Documentation
The authoritative documentation for the engine capability matrix lives in two key locations within the repository.
Primary Documentation in Expressive TTS Specs
The most detailed specification resides in docs/specs/01-expressive-tts.md under subsection "4.4 Per-engine capability matrix (expression.lower)". This table lists each TTS engine alongside its supported expression mechanisms, intensity handling, reference-clip support, and how the lower-casing (lower()) conversion maps those capabilities. According to the VoiceStudio source code, this markdown file serves as the single source of truth for backend capability definitions.
Roadmap References
A high-level overview also appears in docs/ROADMAP.md where the matrix is referenced as part of the "per-engine capability gates" discussion. This document provides strategic context about how capability matrices influence feature rollout and engine integration priorities.
Accessing the Engine Capability Matrix Programmatically
VoiceStudio ensures the UI and documentation stay synchronized by rendering capability data directly from the backend API. The runtime source lives in backend/tts_backend.py, specifically within the list_backends() function.
Python Backend API
You can retrieve the capability matrix directly from Python using the backend module:
from voice_studio.backends import list_backends
# The dict key `expression_caps` holds the per‑engine matrix
engine_caps = list_backends()["expression_caps"]
for engine_id, caps in engine_caps.items():
print(f"{engine_id}:")
print(f" mechanism: {caps['mechanism']}")
print(f" emotion support: {caps['emotion']}")
print(f" intensity: {caps['intensity']}")
print(f" emo‑ref: {caps['emo_ref']}")
This returns a dictionary keyed by engine ID, where each value contains the specific capability flags defined in the specification.
Frontend JavaScript Integration
For React-based frontend applications, fetch the same data from the REST endpoint:
import { useEffect, useState } from "react";
import axios from "axios";
function EngineCapabilityTable() {
const [caps, setCaps] = useState({});
useEffect(() => {
axios.get("/api/engines/tts").then((res) => {
setCaps(res.data.expression_caps);
});
}, []);
return (
<table>
<thead>
<tr>
<th>Engine</th><th>Mechanism</th><th>Emotion</th><th>Intensity</th><th>Emo‑Ref</th>
</tr>
</thead>
<tbody>
{Object.entries(caps).map(([id, v]) => (
<tr key={id}>
<td>{id}</td><td>{v.mechanism}</td><td>{v.emotion}</td>
<td>{v.intensity}</td><td>{v.emo_ref}</td>
</tr>
))}
</tbody>
</table>
);
}
Understanding the Matrix Structure
The capability matrix standardizes five critical dimensions across all supported engines:
- mechanism: The underlying TTS synthesis method employed by the engine
- emotion: Boolean flag indicating native emotional expression support
- intensity: Capability for controlling expression strength levels
- emo_ref: Support for reference audio clips to guide emotional style transfer
This structure ensures consistent feature gating across the omnivoice, cosyvoice, indextts2, and voxcpm2 backends implemented in VoiceStudio.
Summary
- The authoritative engine capability matrix is defined in
docs/specs/01-expressive-tts.mdsection 4.4 - Strategic context appears in
docs/ROADMAP.mdas part of capability gates planning - Runtime access is provided via
list_backends()['expression_caps']inbackend/tts_backend.py - Both Python and JavaScript APIs expose identical capability data to maintain documentation-backend parity
- The matrix tracks mechanism, emotion, intensity, and reference-clip support across all integrated TTS engines
Frequently Asked Questions
Where is the engine capability matrix documented in VoiceStudio?
The matrix is formally documented in docs/specs/01-expressive-tts.md under section 4.4, "Per-engine capability matrix (expression.lower)". This file contains the complete specification table mapping each TTS engine to its supported expression features.
How do I retrieve engine capabilities via the VoiceStudio API?
Import list_backends from voice_studio.backends and access the expression_caps key. This returns a dictionary containing the capability matrix for all configured engines, ensuring your application receives the same data used by the VoiceStudio UI.
What TTS engines are supported in the VoiceStudio capability matrix?
According to the specification, the matrix includes commercial and open-source engines such as omnivoice, cosyvoice, indextts2, and voxcpm2. Each entry details that engine's specific support for mechanisms, emotions, intensity controls, and reference-clip processing.
Does the VoiceStudio UI use the same capability data as the documentation?
Yes. The UI renders the capability matrix programmatically from list_backends()['expression_caps'], ensuring the interface displays capabilities that exactly match the docs/specs/01-expressive-tts.md specification. This synchronization prevents feature mismatches between documentation and runtime behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →