# How to Generate Speech in a Cloned Voice Using the VoiceStudio API

> Learn how to generate speech in a cloned voice with the VoiceStudio API. Upload audio, submit a clone task, and retrieve synthesized speech for your projects.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-09

---

**You can generate speech in a cloned voice using the VoiceStudio API by uploading a reference audio file to the `/upload` endpoint, submitting a generation task with `"operation": "clone"` to the `/generate` endpoint, and retrieving the synthesized audio from the resulting output artifact.**

Voice cloning enables developers to synthesize speech that mimics a specific speaker from a short audio sample. The `debpalash/VoiceStudio` repository implements this capability through an asynchronous, engine-agnostic architecture that supports multiple TTS backends. This guide explains the complete workflow to generate speech in a cloned voice, referencing the actual source implementation and API contracts.

## Prerequisites: Engine Support for Cloning

Before initiating a clone request, verify that your configured TTS engine supports voice cloning. According to [`api/routers/generation.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/generation.py), the system validates that `engine.can_clone == True` before accepting clone tasks. The engine registry in [`services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/tts_backend.py) (lines 150–165) defines which backends advertise the `"clone"` capability.

## Step-by-Step Workflow

The VoiceStudio API implements voice cloning through a three-phase asynchronous process.

### Step 1: Upload the Reference Audio File

First, upload the raw WAV or MP3 file containing the target voice. Send a multipart form POST request to the `/upload` endpoint with the audio file assigned to the `file` field.

```python
import requests

files = {"file": open("my_voice.wav", "rb")}
resp = requests.post("https://api.voicestudio.example.com/upload", files=files)
artifact_id = resp.json()["artifact_id"]  # Returns path like "inputs/voice.wav"

```

The server stores the uploaded file under `omnivoice_data/voices/` and returns an `artifact_id` that uniquely identifies the reference audio. This upload flow is validated in [`tests/test_worker_upload_server.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_upload_server.py) (lines 1494–1517).

### Step 2: Submit a Clone Generation Task

Create a generation task specifying the clone operation. POST to `/generate` with a JSON payload containing the engine identifier, operation type, and parameters object.

```python
payload = {
    "engine": "indextts",
    "operation": "clone",
    "params": {
        "text": "Hello, this is my cloned voice speaking.",
        "ref_audio": artifact_id
    }
}
resp = requests.post("https://api.voicestudio.example.com/generate", json=payload)
task_id = resp.json()["task_id"]

```

The `operation` field must be explicitly set to `"clone"`, and the `ref_audio` parameter must match the `artifact_id` from Step 1. The generation router in [`api/routers/generation.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/generation.py) validates these parameters and creates a `Task` object with `operation="clone"`. This specific payload structure is demonstrated in [`tests/test_worker_inputs.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_inputs.py) (lines 85–94) and [`tests/test_worker_service_api.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_service_api.py) (lines 755–767).

### Step 3: Retrieve the Synthesized Audio

After the worker processes the task, retrieve the generated audio. Poll the task status endpoint until the state indicates completion, then download the result from the artifact store.

```python
while True:
    status = requests.get(f"https://api.voicestudio.example.com/task/{task_id}").json()
    if status["state"] == "completed":
        break

audio_id = status["output_artifact_id"]
audio = requests.get(f"https://api.voicestudio.example.com/artifact/{audio_id}").content

with open("cloned_output.wav", "wb") as f:
    f.write(audio)

```

The synthesized waveform is stored as a new artifact in the content-addressed cache. The verification logic for this retrieval process appears in [`tests/test_worker_inputs.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_inputs.py) (lines 184–196).

## Implementation Architecture

Understanding the backend flow clarifies how reference audio is handled securely and efficiently.

### Generation Router and Validation

The [`api/routers/generation.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/generation.py) module handles incoming `/generate` requests (source lines 34–50). It inspects the requested engine's capabilities and rejects clone requests for engines where `can_clone` is false, preventing invalid task creation early in the pipeline.

### Task Staging and Reference Audio Storage

The [`services/task_store.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/task_store.py) component manages reference audio staging (lines 70–90). When a clone task is created, the system copies the `ref_audio` file into a content-addressed cache location. This staging step ensures that distributed workers can access the reference audio without exposing internal control-plane file paths or requiring direct filesystem access to the upload directory. The staging logic is exercised in [`tests/test_worker_inputs.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_inputs.py) (lines 202–214).

### Worker Execution and Speech Synthesis

The [`services/worker.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/worker.py) module executes clone tasks by loading the appropriate TTS backend via `tts_backend.get_engine_instance`. For clone operations, the worker invokes the engine's `clone_speech(ref_audio, text)` method, which generates a new waveform matching the reference voice's acoustic characteristics. The capability discovery mechanism is validated in [`tests/test_worker_service_api.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_service_api.py) (lines 191–203).

## Summary

- **Upload reference audio** to `/upload` to receive an `artifact_id` for your voice sample.
- **Submit a clone task** via `/generate` with `"operation": "clone"` and the reference audio ID in `params.ref_audio`.
- **Retrieve results** by polling the task status and downloading from `/artifact/{artifact_id}`.
- **Engine validation** occurs in [`api/routers/generation.py`](https://github.com/debpalash/VoiceStudio/blob/main/api/routers/generation.py) through the `can_clone` property check.
- **Secure staging** happens in [`services/task_store.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/task_store.py) to enable distributed worker access without exposing upload paths.

## Frequently Asked Questions

### What audio formats does VoiceStudio accept for voice cloning?

VoiceStudio accepts standard audio formats including WAV and MP3 for the reference upload. The system stores these files in `omnivoice_data/voices/` before staging them for processing. Ensure your reference recording is clear and contains only the target speaker's voice for optimal synthesis quality.

### How does the VoiceStudio API handle concurrent clone requests?

The API uses an asynchronous task queue managed by [`services/task_store.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/task_store.py) and processed by [`services/worker.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/worker.py). Each clone request creates an independent `Task` object. The content-addressed cache allows multiple generation tasks to reference the same uploaded voice sample without creating duplicate stored copies.

### Which TTS engines support the clone operation?

Support depends on the specific engine configuration in [`services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/tts_backend.py). Engines must explicitly declare the `"clone"` capability in their supported operations list. The test suite verifies this capability discovery mechanism in [`tests/test_worker_service_api.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_worker_service_api.py) (lines 191–203). The `indextts` engine is commonly used for cloning, but availability depends on your deployment configuration.

### Can I use the same reference audio for multiple generation tasks?

Yes. Once uploaded, the reference audio remains available in the artifact store under its assigned `artifact_id`. You can reference this same ID in multiple `/generate` requests to create different speech content using the same cloned voice without re-uploading the original sample.