How to Generate Speech in a Cloned Voice Using the VoiceStudio API

You can generate speech in a cloned voice using the VoiceStudio API by uploading a reference audio file to the /upload endpoint, submitting a generation task with "operation": "clone" to the /generate endpoint, and retrieving the synthesized audio from the resulting output artifact.

Voice cloning enables developers to synthesize speech that mimics a specific speaker from a short audio sample. The debpalash/VoiceStudio repository implements this capability through an asynchronous, engine-agnostic architecture that supports multiple TTS backends. This guide explains the complete workflow to generate speech in a cloned voice, referencing the actual source implementation and API contracts.

Prerequisites: Engine Support for Cloning

Before initiating a clone request, verify that your configured TTS engine supports voice cloning. According to api/routers/generation.py, the system validates that engine.can_clone == True before accepting clone tasks. The engine registry in services/tts_backend.py (lines 150–165) defines which backends advertise the "clone" capability.

Step-by-Step Workflow

The VoiceStudio API implements voice cloning through a three-phase asynchronous process.

Step 1: Upload the Reference Audio File

First, upload the raw WAV or MP3 file containing the target voice. Send a multipart form POST request to the /upload endpoint with the audio file assigned to the file field.

import requests

files = {"file": open("my_voice.wav", "rb")}
resp = requests.post("https://api.voicestudio.example.com/upload", files=files)
artifact_id = resp.json()["artifact_id"]  # Returns path like "inputs/voice.wav"

The server stores the uploaded file under omnivoice_data/voices/ and returns an artifact_id that uniquely identifies the reference audio. This upload flow is validated in tests/test_worker_upload_server.py (lines 1494–1517).

Step 2: Submit a Clone Generation Task

Create a generation task specifying the clone operation. POST to /generate with a JSON payload containing the engine identifier, operation type, and parameters object.

payload = {
    "engine": "indextts",
    "operation": "clone",
    "params": {
        "text": "Hello, this is my cloned voice speaking.",
        "ref_audio": artifact_id
    }
}
resp = requests.post("https://api.voicestudio.example.com/generate", json=payload)
task_id = resp.json()["task_id"]

The operation field must be explicitly set to "clone", and the ref_audio parameter must match the artifact_id from Step 1. The generation router in api/routers/generation.py validates these parameters and creates a Task object with operation="clone". This specific payload structure is demonstrated in tests/test_worker_inputs.py (lines 85–94) and tests/test_worker_service_api.py (lines 755–767).

Step 3: Retrieve the Synthesized Audio

After the worker processes the task, retrieve the generated audio. Poll the task status endpoint until the state indicates completion, then download the result from the artifact store.

while True:
    status = requests.get(f"https://api.voicestudio.example.com/task/{task_id}").json()
    if status["state"] == "completed":
        break

audio_id = status["output_artifact_id"]
audio = requests.get(f"https://api.voicestudio.example.com/artifact/{audio_id}").content

with open("cloned_output.wav", "wb") as f:
    f.write(audio)

The synthesized waveform is stored as a new artifact in the content-addressed cache. The verification logic for this retrieval process appears in tests/test_worker_inputs.py (lines 184–196).

Implementation Architecture

Understanding the backend flow clarifies how reference audio is handled securely and efficiently.

Generation Router and Validation

The api/routers/generation.py module handles incoming /generate requests (source lines 34–50). It inspects the requested engine's capabilities and rejects clone requests for engines where can_clone is false, preventing invalid task creation early in the pipeline.

Task Staging and Reference Audio Storage

The services/task_store.py component manages reference audio staging (lines 70–90). When a clone task is created, the system copies the ref_audio file into a content-addressed cache location. This staging step ensures that distributed workers can access the reference audio without exposing internal control-plane file paths or requiring direct filesystem access to the upload directory. The staging logic is exercised in tests/test_worker_inputs.py (lines 202–214).

Worker Execution and Speech Synthesis

The services/worker.py module executes clone tasks by loading the appropriate TTS backend via tts_backend.get_engine_instance. For clone operations, the worker invokes the engine's clone_speech(ref_audio, text) method, which generates a new waveform matching the reference voice's acoustic characteristics. The capability discovery mechanism is validated in tests/test_worker_service_api.py (lines 191–203).

Summary

  • Upload reference audio to /upload to receive an artifact_id for your voice sample.
  • Submit a clone task via /generate with "operation": "clone" and the reference audio ID in params.ref_audio.
  • Retrieve results by polling the task status and downloading from /artifact/{artifact_id}.
  • Engine validation occurs in api/routers/generation.py through the can_clone property check.
  • Secure staging happens in services/task_store.py to enable distributed worker access without exposing upload paths.

Frequently Asked Questions

What audio formats does VoiceStudio accept for voice cloning?

VoiceStudio accepts standard audio formats including WAV and MP3 for the reference upload. The system stores these files in omnivoice_data/voices/ before staging them for processing. Ensure your reference recording is clear and contains only the target speaker's voice for optimal synthesis quality.

How does the VoiceStudio API handle concurrent clone requests?

The API uses an asynchronous task queue managed by services/task_store.py and processed by services/worker.py. Each clone request creates an independent Task object. The content-addressed cache allows multiple generation tasks to reference the same uploaded voice sample without creating duplicate stored copies.

Which TTS engines support the clone operation?

Support depends on the specific engine configuration in services/tts_backend.py. Engines must explicitly declare the "clone" capability in their supported operations list. The test suite verifies this capability discovery mechanism in tests/test_worker_service_api.py (lines 191–203). The indextts engine is commonly used for cloning, but availability depends on your deployment configuration.

Can I use the same reference audio for multiple generation tasks?

Yes. Once uploaded, the reference audio remains available in the artifact store under its assigned artifact_id. You can reference this same ID in multiple /generate requests to create different speech content using the same cloned voice without re-uploading the original sample.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →