How to Create a Voice Profile Using the VoiceStudio API: Complete Implementation Guide
You create a voice profile by sending a multipart POST request to the /profiles endpoint with a name, kind parameter set to either "clone" or "design", and type-specific data such as ref_audio for voice cloning or vd_states for designed voices.
Voice profiles act as reusable voice-identity containers within the VoiceStudio ecosystem, enabling synthesis, cloning, and design operations through a RESTful interface. To create a voice profile using the VoiceStudio API, you interact with the FastAPI backend implemented in the debpalash/VoiceStudio repository, which validates requests, stores metadata in SQLite, and manages audio assets in structured directories. The system distinguishes between clone profiles (derived from reference recordings) and design profiles (synthesized from archetype parameters), each requiring specific payload structures.
Voice Profile Types and Use Cases
VoiceStudio supports two distinct profile architectures that determine how the voice identity is established and rendered.
Clone Profiles
Clone profiles replicate an existing voice from uploaded audio samples. When kind is set to clone, the API requires a ref_audio field containing a WAV or MP3 file. The system stores this reference audio in the VOICES_DIR directory (configured in core/config.py) using a filename derived from a short UUID (profile_id.wav). This approach captures the acoustic characteristics of a specific speaker for later synthesis operations.
Design Profiles
Design profiles construct voices from parametric descriptions rather than recordings. When kind is set to design, you must provide vd_states as a JSON object describing design-category selections (such as pitch, speed, or tone). According to the implementation in backend/api/routers/profiles.py (lines 18-27), the server optionally renders a deterministic sample WAV using seed 42 via the _render_archetype_wav function from backend/api/routers/archetypes.py. If the voice engine is not ready, the sample generation is deferred and persisted later.
POST /profiles Endpoint Specification
The profile creation endpoint is defined in backend/api/routers/profiles.py at lines 43-53. It expects multipart/form-data rather than JSON to accommodate binary audio uploads.
Required and Optional Parameters
The endpoint accepts the following fields:
name(required): Human-readable identifier for the profile.kind(required): Enum value of eithercloneordesign.ref_audio(required ifkind=clone): Reference recording file (WAV or MP3 format).vd_states(required ifkind=design): JSON object describing voice design parameters (e.g.,{"pitch":"high","speed":"fast"}).ref_text(optional): Transcript of the reference audio or custom sample script for design profiles.instruct(optional): Instruction string describing desired characteristics (e.g., "female, calm"). The system processes this throughheal_design_instructandsanitize_instructhelpers inomnivoice/utils/voice_design.py.language(optional): Target language code; defaults to "Auto".seed(optional): Integer for deterministic voice rendering.personality(optional): Preset personality tag for the voice.
Validation and Error Handling
The router enforces strict validation rules (lines 55-63). If kind is clone but ref_audio is missing, the API returns a 422 Unprocessable Entity error. Similarly, design profiles without vd_states trigger validation failures. The validation occurs before database insertion via the db_conn() context manager defined in core/db.py.
Implementation Details from Source Code
Understanding the backend flow helps optimize your API usage and debug issues effectively.
Database Storage Flow
Upon successful validation, the profile is inserted into the SQLite voice_profiles table using the shared database connection from core/db.py. The insertion occurs within an async context manager that ensures proper connection handling. The generated profile ID uses a short UUID format that also serves as the basename for stored audio files.
Audio File Management
For clone profiles, the uploaded ref_audio is persisted to the filesystem location defined by VOICES_DIR (from core/config.py) with the filename pattern profile_id.wav. This deterministic naming convention allows the GET /profiles/{id}/audio endpoint to retrieve files without database lookups for path resolution.
Design Profile Rendering Pipeline
When creating design profiles, the system optionally invokes the archetype renderer immediately. The _render_archetype_wav function (sourced from backend/api/routers/archetypes.py) generates a deterministic sample using the provided vd_states and seed value. If the voice engine is unavailable during creation, the profile is created without the sample, and the audio is generated asynchronously when the engine becomes ready.
Step-by-Step Code Examples
Creating a Clone Profile with Python
Use the requests library to upload reference audio with metadata:
import requests
url = "http://localhost:8000/profiles"
# Open the reference audio file in binary mode
files = {"ref_audio": open("my_voice.wav", "rb")}
data = {
"name": "My Clone",
"kind": "clone",
"language": "en",
"instruct": "female, calm",
}
resp = requests.post(url, files=files, data=data)
print(resp.json())
# Output: {"id": "a1b2c3d4", "name": "My Clone", "kind": "clone"}
The API returns a JSON payload containing the new profile's id, name, and kind, which you should store for subsequent operations.
Creating a Design Profile with cURL
For parametric voice design without reference audio:
curl -X POST http://localhost:8000/profiles \
-F "name=My Design" \
-F "kind=design" \
-F 'vd_states={"pitch":"high","speed":"fast"}' \
-F "instruct=female, energetic"
Note that vd_states must be valid JSON passed as a form field, not a file upload.
Retrieving Profile Metadata
After creation, fetch the profile details using the returned ID:
profile_id = "a1b2c3d4"
resp = requests.get(f"http://localhost:8000/profiles/{profile_id}")
print(resp.json())
This returns the complete profile record including stored parameters and processing status.
Downloading Profile Audio
To retrieve the stored reference audio or generated design sample:
audio_resp = requests.get(f"http://localhost:8000/profiles/{profile_id}/audio")
with open("profile_audio.wav", "wb") as f:
f.write(audio_resp.content)
This endpoint streams the audio file directly from VOICES_DIR using the profile ID as the filename lookup key.
Key Source Files and Architecture
The profile creation system spans several modules in the debpalash/VoiceStudio repository:
backend/api/routers/profiles.py– Implements the POST/profilesendpoint, request validation, and database insertion logic at lines 43-63.core/db.py– Provides thedb_conn()context manager for atomic SQLite transactions.core/config.py– DefinesVOICES_DIRandOUTPUTS_DIRconstants governing audio file storage locations.omnivoice/utils/voice_design.py– Contains text sanitization functionsheal_design_instructandsanitize_instructthat clean instruction strings before persistence.backend/api/routers/archetypes.py– Supplies the_render_archetype_wavfunction used for deterministic design-profile sample generation.
Summary
- VoiceStudio profiles are created via POST
/profilesusing multipart form data, not JSON. - The
kindparameter determines whether you provideref_audio(for clones) orvd_states(for design). - Clone audio is stored in
VOICES_DIRwith a UUID filename derived from the profile ID. - The FastAPI router in
backend/api/routers/profiles.pyhandles validation, database insertion viacore/db.py, and optional archetype rendering. - Successful creation returns the profile
id,name, andkindfor immediate use in synthesis operations.
Frequently Asked Questions
What is the difference between clone and design profiles in VoiceStudio?
Clone profiles replicate existing voices from uploaded audio files (ref_audio), capturing specific speaker characteristics for accurate voice cloning. Design profiles synthesize voices from parametric JSON descriptions (vd_states) without requiring reference recordings, allowing you to construct voices from descriptive attributes like pitch and speed. The API routes both types through the same endpoint but enforces different required fields for each.
What audio formats are accepted for clone profiles?
The API accepts standard audio formats including WAV and MP3 for the ref_audio field. The system stores the uploaded file in the VOICES_DIR directory (defined in core/config.py) using the profile's short UUID as the filename with a .wav extension. Ensure your reference audio is clear and representative of the target voice for optimal cloning results.
How is the reference audio stored server-side?
According to the implementation in backend/api/routers/profiles.py, uploaded audio is saved to the filesystem path specified by VOICES_DIR with a filename derived from the profile's short UUID (profile_id.wav). The metadata, including the file reference and profile parameters, is stored in the SQLite voice_profiles table via the db_conn() context manager from core/db.py.
Can I update a voice profile after creation?
Yes, the API supports modifying existing profiles via PUT /profiles/{id}. You can update fields such as name, instruct, or ref_text without creating a new profile. However, changing fundamental characteristics like kind (switching from clone to design) typically requires creating a new profile, as the underlying data structures and validation requirements differ significantly between the two profile types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →