# How to Create a Voice Profile Using the VoiceStudio API: Complete Implementation Guide

> Learn to create a voice profile using the VoiceStudio API. Implement voice cloning or design voices via multipart POST requests to the /profiles endpoint. Get the complete guide.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-09

---

**You create a voice profile by sending a multipart POST request to the `/profiles` endpoint with a `name`, `kind` parameter set to either "clone" or "design", and type-specific data such as `ref_audio` for voice cloning or `vd_states` for designed voices.**

Voice profiles act as reusable voice-identity containers within the VoiceStudio ecosystem, enabling synthesis, cloning, and design operations through a RESTful interface. To create a voice profile using the VoiceStudio API, you interact with the FastAPI backend implemented in the `debpalash/VoiceStudio` repository, which validates requests, stores metadata in SQLite, and manages audio assets in structured directories. The system distinguishes between **clone profiles** (derived from reference recordings) and **design profiles** (synthesized from archetype parameters), each requiring specific payload structures.

## Voice Profile Types and Use Cases

VoiceStudio supports two distinct profile architectures that determine how the voice identity is established and rendered.

### Clone Profiles

**Clone profiles** replicate an existing voice from uploaded audio samples. When `kind` is set to `clone`, the API requires a `ref_audio` field containing a WAV or MP3 file. The system stores this reference audio in the `VOICES_DIR` directory (configured in [`core/config.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/config.py)) using a filename derived from a short UUID (`profile_id.wav`). This approach captures the acoustic characteristics of a specific speaker for later synthesis operations.

### Design Profiles

**Design profiles** construct voices from parametric descriptions rather than recordings. When `kind` is set to `design`, you must provide `vd_states` as a JSON object describing design-category selections (such as pitch, speed, or tone). According to the implementation in [`backend/api/routers/profiles.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/profiles.py) (lines 18-27), the server optionally renders a deterministic sample WAV using seed 42 via the `_render_archetype_wav` function from [`backend/api/routers/archetypes.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/archetypes.py). If the voice engine is not ready, the sample generation is deferred and persisted later.

## POST /profiles Endpoint Specification

The profile creation endpoint is defined in [`backend/api/routers/profiles.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/profiles.py) at lines 43-53. It expects **multipart/form-data** rather than JSON to accommodate binary audio uploads.

### Required and Optional Parameters

The endpoint accepts the following fields:

- **`name`** (required): Human-readable identifier for the profile.
- **`kind`** (required): Enum value of either `clone` or `design`.
- **`ref_audio`** (required if `kind=clone`): Reference recording file (WAV or MP3 format).
- **`vd_states`** (required if `kind=design`): JSON object describing voice design parameters (e.g., `{"pitch":"high","speed":"fast"}`).
- **`ref_text`** (optional): Transcript of the reference audio or custom sample script for design profiles.
- **`instruct`** (optional): Instruction string describing desired characteristics (e.g., "female, calm"). The system processes this through `heal_design_instruct` and `sanitize_instruct` helpers in [`omnivoice/utils/voice_design.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/utils/voice_design.py).
- **`language`** (optional): Target language code; defaults to "Auto".
- **`seed`** (optional): Integer for deterministic voice rendering.
- **`personality`** (optional): Preset personality tag for the voice.

### Validation and Error Handling

The router enforces strict validation rules (lines 55-63). If `kind` is `clone` but `ref_audio` is missing, the API returns a **422 Unprocessable Entity** error. Similarly, design profiles without `vd_states` trigger validation failures. The validation occurs before database insertion via the `db_conn()` context manager defined in [`core/db.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/db.py).

## Implementation Details from Source Code

Understanding the backend flow helps optimize your API usage and debug issues effectively.

### Database Storage Flow

Upon successful validation, the profile is inserted into the SQLite `voice_profiles` table using the shared database connection from [`core/db.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/db.py). The insertion occurs within an async context manager that ensures proper connection handling. The generated profile ID uses a short UUID format that also serves as the basename for stored audio files.

### Audio File Management

For clone profiles, the uploaded `ref_audio` is persisted to the filesystem location defined by `VOICES_DIR` (from [`core/config.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/config.py)) with the filename pattern `profile_id.wav`. This deterministic naming convention allows the `GET /profiles/{id}/audio` endpoint to retrieve files without database lookups for path resolution.

### Design Profile Rendering Pipeline

When creating design profiles, the system optionally invokes the archetype renderer immediately. The `_render_archetype_wav` function (sourced from [`backend/api/routers/archetypes.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/archetypes.py)) generates a deterministic sample using the provided `vd_states` and seed value. If the voice engine is unavailable during creation, the profile is created without the sample, and the audio is generated asynchronously when the engine becomes ready.

## Step-by-Step Code Examples

### Creating a Clone Profile with Python

Use the `requests` library to upload reference audio with metadata:

```python
import requests

url = "http://localhost:8000/profiles"

# Open the reference audio file in binary mode

files = {"ref_audio": open("my_voice.wav", "rb")}
data = {
    "name": "My Clone",
    "kind": "clone",
    "language": "en",
    "instruct": "female, calm",
}

resp = requests.post(url, files=files, data=data)
print(resp.json())

# Output: {"id": "a1b2c3d4", "name": "My Clone", "kind": "clone"}

```

The API returns a JSON payload containing the new profile's `id`, `name`, and `kind`, which you should store for subsequent operations.

### Creating a Design Profile with cURL

For parametric voice design without reference audio:

```bash
curl -X POST http://localhost:8000/profiles \
  -F "name=My Design" \
  -F "kind=design" \
  -F 'vd_states={"pitch":"high","speed":"fast"}' \
  -F "instruct=female, energetic"

```

Note that `vd_states` must be valid JSON passed as a form field, not a file upload.

### Retrieving Profile Metadata

After creation, fetch the profile details using the returned ID:

```python
profile_id = "a1b2c3d4"
resp = requests.get(f"http://localhost:8000/profiles/{profile_id}")
print(resp.json())

```

This returns the complete profile record including stored parameters and processing status.

### Downloading Profile Audio

To retrieve the stored reference audio or generated design sample:

```python
audio_resp = requests.get(f"http://localhost:8000/profiles/{profile_id}/audio")
with open("profile_audio.wav", "wb") as f:
    f.write(audio_resp.content)

```

This endpoint streams the audio file directly from `VOICES_DIR` using the profile ID as the filename lookup key.

## Key Source Files and Architecture

The profile creation system spans several modules in the `debpalash/VoiceStudio` repository:

- **[`backend/api/routers/profiles.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/profiles.py)** – Implements the POST `/profiles` endpoint, request validation, and database insertion logic at lines 43-63.
- **[`core/db.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/db.py)** – Provides the `db_conn()` context manager for atomic SQLite transactions.
- **[`core/config.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/config.py)** – Defines `VOICES_DIR` and `OUTPUTS_DIR` constants governing audio file storage locations.
- **[`omnivoice/utils/voice_design.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/utils/voice_design.py)** – Contains text sanitization functions `heal_design_instruct` and `sanitize_instruct` that clean instruction strings before persistence.
- **[`backend/api/routers/archetypes.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/archetypes.py)** – Supplies the `_render_archetype_wav` function used for deterministic design-profile sample generation.

## Summary

- VoiceStudio profiles are created via **POST `/profiles`** using multipart form data, not JSON.
- The **`kind`** parameter determines whether you provide `ref_audio` (for clones) or `vd_states` (for design).
- Clone audio is stored in `VOICES_DIR` with a UUID filename derived from the profile ID.
- The FastAPI router in [`backend/api/routers/profiles.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/profiles.py) handles validation, database insertion via [`core/db.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/db.py), and optional archetype rendering.
- Successful creation returns the profile `id`, `name`, and `kind` for immediate use in synthesis operations.

## Frequently Asked Questions

### What is the difference between clone and design profiles in VoiceStudio?

**Clone profiles** replicate existing voices from uploaded audio files (`ref_audio`), capturing specific speaker characteristics for accurate voice cloning. **Design profiles** synthesize voices from parametric JSON descriptions (`vd_states`) without requiring reference recordings, allowing you to construct voices from descriptive attributes like pitch and speed. The API routes both types through the same endpoint but enforces different required fields for each.

### What audio formats are accepted for clone profiles?

The API accepts standard audio formats including **WAV and MP3** for the `ref_audio` field. The system stores the uploaded file in the `VOICES_DIR` directory (defined in [`core/config.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/config.py)) using the profile's short UUID as the filename with a `.wav` extension. Ensure your reference audio is clear and representative of the target voice for optimal cloning results.

### How is the reference audio stored server-side?

According to the implementation in [`backend/api/routers/profiles.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/profiles.py), uploaded audio is saved to the filesystem path specified by `VOICES_DIR` with a filename derived from the profile's short UUID (`profile_id.wav`). The metadata, including the file reference and profile parameters, is stored in the SQLite `voice_profiles` table via the `db_conn()` context manager from [`core/db.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/db.py).

### Can I update a voice profile after creation?

Yes, the API supports modifying existing profiles via **PUT `/profiles/{id}`**. You can update fields such as `name`, `instruct`, or `ref_text` without creating a new profile. However, changing fundamental characteristics like `kind` (switching from clone to design) typically requires creating a new profile, as the underlying data structures and validation requirements differ significantly between the two profile types.