How to Create a Voice Profile Using the VoiceStudio API: Complete Implementation Guide

You create a voice profile by sending a multipart POST request to the /profiles endpoint with a name, kind parameter set to either "clone" or "design", and type-specific data such as ref_audio for voice cloning or vd_states for designed voices.

Voice profiles act as reusable voice-identity containers within the VoiceStudio ecosystem, enabling synthesis, cloning, and design operations through a RESTful interface. To create a voice profile using the VoiceStudio API, you interact with the FastAPI backend implemented in the debpalash/VoiceStudio repository, which validates requests, stores metadata in SQLite, and manages audio assets in structured directories. The system distinguishes between clone profiles (derived from reference recordings) and design profiles (synthesized from archetype parameters), each requiring specific payload structures.

Voice Profile Types and Use Cases

VoiceStudio supports two distinct profile architectures that determine how the voice identity is established and rendered.

Clone Profiles

Clone profiles replicate an existing voice from uploaded audio samples. When kind is set to clone, the API requires a ref_audio field containing a WAV or MP3 file. The system stores this reference audio in the VOICES_DIR directory (configured in core/config.py) using a filename derived from a short UUID (profile_id.wav). This approach captures the acoustic characteristics of a specific speaker for later synthesis operations.

Design Profiles

Design profiles construct voices from parametric descriptions rather than recordings. When kind is set to design, you must provide vd_states as a JSON object describing design-category selections (such as pitch, speed, or tone). According to the implementation in backend/api/routers/profiles.py (lines 18-27), the server optionally renders a deterministic sample WAV using seed 42 via the _render_archetype_wav function from backend/api/routers/archetypes.py. If the voice engine is not ready, the sample generation is deferred and persisted later.

POST /profiles Endpoint Specification

The profile creation endpoint is defined in backend/api/routers/profiles.py at lines 43-53. It expects multipart/form-data rather than JSON to accommodate binary audio uploads.

Required and Optional Parameters

The endpoint accepts the following fields:

  • name (required): Human-readable identifier for the profile.
  • kind (required): Enum value of either clone or design.
  • ref_audio (required if kind=clone): Reference recording file (WAV or MP3 format).
  • vd_states (required if kind=design): JSON object describing voice design parameters (e.g., {"pitch":"high","speed":"fast"}).
  • ref_text (optional): Transcript of the reference audio or custom sample script for design profiles.
  • instruct (optional): Instruction string describing desired characteristics (e.g., "female, calm"). The system processes this through heal_design_instruct and sanitize_instruct helpers in omnivoice/utils/voice_design.py.
  • language (optional): Target language code; defaults to "Auto".
  • seed (optional): Integer for deterministic voice rendering.
  • personality (optional): Preset personality tag for the voice.

Validation and Error Handling

The router enforces strict validation rules (lines 55-63). If kind is clone but ref_audio is missing, the API returns a 422 Unprocessable Entity error. Similarly, design profiles without vd_states trigger validation failures. The validation occurs before database insertion via the db_conn() context manager defined in core/db.py.

Implementation Details from Source Code

Understanding the backend flow helps optimize your API usage and debug issues effectively.

Database Storage Flow

Upon successful validation, the profile is inserted into the SQLite voice_profiles table using the shared database connection from core/db.py. The insertion occurs within an async context manager that ensures proper connection handling. The generated profile ID uses a short UUID format that also serves as the basename for stored audio files.

Audio File Management

For clone profiles, the uploaded ref_audio is persisted to the filesystem location defined by VOICES_DIR (from core/config.py) with the filename pattern profile_id.wav. This deterministic naming convention allows the GET /profiles/{id}/audio endpoint to retrieve files without database lookups for path resolution.

Design Profile Rendering Pipeline

When creating design profiles, the system optionally invokes the archetype renderer immediately. The _render_archetype_wav function (sourced from backend/api/routers/archetypes.py) generates a deterministic sample using the provided vd_states and seed value. If the voice engine is unavailable during creation, the profile is created without the sample, and the audio is generated asynchronously when the engine becomes ready.

Step-by-Step Code Examples

Creating a Clone Profile with Python

Use the requests library to upload reference audio with metadata:

import requests

url = "http://localhost:8000/profiles"

# Open the reference audio file in binary mode

files = {"ref_audio": open("my_voice.wav", "rb")}
data = {
    "name": "My Clone",
    "kind": "clone",
    "language": "en",
    "instruct": "female, calm",
}

resp = requests.post(url, files=files, data=data)
print(resp.json())

# Output: {"id": "a1b2c3d4", "name": "My Clone", "kind": "clone"}

The API returns a JSON payload containing the new profile's id, name, and kind, which you should store for subsequent operations.

Creating a Design Profile with cURL

For parametric voice design without reference audio:

curl -X POST http://localhost:8000/profiles \
  -F "name=My Design" \
  -F "kind=design" \
  -F 'vd_states={"pitch":"high","speed":"fast"}' \
  -F "instruct=female, energetic"

Note that vd_states must be valid JSON passed as a form field, not a file upload.

Retrieving Profile Metadata

After creation, fetch the profile details using the returned ID:

profile_id = "a1b2c3d4"
resp = requests.get(f"http://localhost:8000/profiles/{profile_id}")
print(resp.json())

This returns the complete profile record including stored parameters and processing status.

Downloading Profile Audio

To retrieve the stored reference audio or generated design sample:

audio_resp = requests.get(f"http://localhost:8000/profiles/{profile_id}/audio")
with open("profile_audio.wav", "wb") as f:
    f.write(audio_resp.content)

This endpoint streams the audio file directly from VOICES_DIR using the profile ID as the filename lookup key.

Key Source Files and Architecture

The profile creation system spans several modules in the debpalash/VoiceStudio repository:

  • backend/api/routers/profiles.py – Implements the POST /profiles endpoint, request validation, and database insertion logic at lines 43-63.
  • core/db.py – Provides the db_conn() context manager for atomic SQLite transactions.
  • core/config.py – Defines VOICES_DIR and OUTPUTS_DIR constants governing audio file storage locations.
  • omnivoice/utils/voice_design.py – Contains text sanitization functions heal_design_instruct and sanitize_instruct that clean instruction strings before persistence.
  • backend/api/routers/archetypes.py – Supplies the _render_archetype_wav function used for deterministic design-profile sample generation.

Summary

  • VoiceStudio profiles are created via POST /profiles using multipart form data, not JSON.
  • The kind parameter determines whether you provide ref_audio (for clones) or vd_states (for design).
  • Clone audio is stored in VOICES_DIR with a UUID filename derived from the profile ID.
  • The FastAPI router in backend/api/routers/profiles.py handles validation, database insertion via core/db.py, and optional archetype rendering.
  • Successful creation returns the profile id, name, and kind for immediate use in synthesis operations.

Frequently Asked Questions

What is the difference between clone and design profiles in VoiceStudio?

Clone profiles replicate existing voices from uploaded audio files (ref_audio), capturing specific speaker characteristics for accurate voice cloning. Design profiles synthesize voices from parametric JSON descriptions (vd_states) without requiring reference recordings, allowing you to construct voices from descriptive attributes like pitch and speed. The API routes both types through the same endpoint but enforces different required fields for each.

What audio formats are accepted for clone profiles?

The API accepts standard audio formats including WAV and MP3 for the ref_audio field. The system stores the uploaded file in the VOICES_DIR directory (defined in core/config.py) using the profile's short UUID as the filename with a .wav extension. Ensure your reference audio is clear and representative of the target voice for optimal cloning results.

How is the reference audio stored server-side?

According to the implementation in backend/api/routers/profiles.py, uploaded audio is saved to the filesystem path specified by VOICES_DIR with a filename derived from the profile's short UUID (profile_id.wav). The metadata, including the file reference and profile parameters, is stored in the SQLite voice_profiles table via the db_conn() context manager from core/db.py.

Can I update a voice profile after creation?

Yes, the API supports modifying existing profiles via PUT /profiles/{id}. You can update fields such as name, instruct, or ref_text without creating a new profile. However, changing fundamental characteristics like kind (switching from clone to design) typically requires creating a new profile, as the underlying data structures and validation requirements differ significantly between the two profile types.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →