Guidelines for Reference Audio in Voice Cloning: VoiceStudio Requirements and Best Practices

VoiceStudio requires reference audio between 5–15 seconds (minimum 3 seconds, maximum 20 seconds) in 16 kHz mono format with minimal background noise and matching target language to successfully clone a voice.

VoiceStudio (debpalash/VoiceStudio) performs voice cloning by analyzing a short reference clip to capture speaker characteristics. The system enforces strict validation rules in omnivoice/utils/audio.py to ensure quality, rejecting clips that violate technical constraints with specific error markers surfaced to the UI.

Technical Requirements for Reference Audio

VoiceStudio validates all incoming reference audio against hard constraints before processing. Violations trigger immediate failures with descriptive error messages.

Duration Constraints and Limits

The duration requirements follow a three-tier structure enforced during preprocessing:

  • Minimum: approximately 3 seconds of non-silent audio (below this threshold, the clone lacks sufficient phonetic diversity)
  • Optimal: 5–15 seconds for best quality results
  • Maximum: strictly capped at 20 seconds (exceeding this triggers CLONE_REF_TOO_LONG_MARKER)

As implemented in omnivoice/utils/audio.py lines 148–199, the system automatically removes leading and trailing silence before measuring duration. If silence removal leaves the clip empty, the operation fails with "Reference audio is empty after silence removal".

Audio Format Specifications

VoiceStudio accepts multiple input formats but normalizes them internally to 16 kHz mono PCM. The supported formats include:

  • WAV (recommended)
  • MP3
  • FLAC
  • Base64-encoded data-URI (for API submissions)

All stereo inputs are mixed down to mono during preprocessing. Higher sampling rates are automatically down-sampled to 16 kHz, but supplying native 16 kHz mono files reduces processing overhead.

Content Quality Guidelines

Beyond technical specifications, the acoustic content of the reference clip significantly impacts cloning fidelity.

Background Noise and Audio Environment

The reference recording must contain minimal background noise, no music, and no ambient sounds. The system detects usable audio segments, and excessive noise can trigger CLONE_REF_UNUSABLE_MARKER, indicating no usable sound remains after preprocessing. Record in quiet environments using close-mic techniques to isolate the target speaker.

Language Matching and Speaker Consistency

The spoken language in the reference audio must match the language of the text you intend to synthesize. VoiceStudio infers language characteristics from the reference clip, and mismatched language input degrades voice quality and pronunciation accuracy. The clip must feature a single, consistent speaker without overlapping voices or conversation.

Volume Normalization and Content Selection

Optimal reference audio maintains normalized loudness around ‑23 LUFS ( Loudness Units relative to Full Scale). Clips that are too quiet risk "no usable sound" errors, while clipped or distorted audio degrades timbre capture. Select neutral script content—such as a short paragraph read in a natural, conversational tone—rather than emotionally extreme or whispered speech. Expressive styles can be added later via instruct strings during generation.

Error Handling and Validation

VoiceStudio surfaces validation failures through specific error markers defined in the audio processing pipeline:

  • CLONE_REF_UNUSABLE_MARKER: Reference audio has no usable sound after silence removal or noise filtering
  • CLONE_REF_TOO_LONG_MARKER: Audio exceeds the 20‑second limit

These markers propagate to the frontend UI, triggering user-facing toast messages that guide corrective action, as tested in frontend/src/test/errorToastUserFixable.test.jsx lines 25–50. Users receive immediate feedback to re-record or trim their samples rather than waiting for processing completion.

Practical Implementation Examples

CLI Usage with OmniVoice

Clone a voice from a local file using the command-line interface:

omnivoice clone --ref-audio path/to/voice.wav \
                --text "Hello, this is my cloned voice." \
                --output cloned.wav

The CLI automatically validates duration and format before submitting to the backend.

Python API Integration

Submit base64-encoded reference audio to the MCP clone_voice endpoint:

import requests
import base64

# Load and encode reference audio

with open("voice.wav", "rb") as f:
    b64_audio = base64.b64encode(f.read()).decode()

payload = {
    "ref_audio": f"data:audio/wav;base64,{b64_audio}",
    "text": "Welcome to VoiceStudio!",
    "language": "en"
}

resp = requests.post(
    "http://localhost:3900/v1/clone_voice",
    json=payload,
    timeout=30
)
resp.raise_for_status()
profile_id = resp.json()["profile_id"]

Generate speech from the created profile:

gen_payload = {
    "profile_id": profile_id,
    "text": "This is the synthesized output using the cloned voice.",
    "language": "en"
}

gen_resp = requests.post(
    "http://localhost:3900/v1/generate_speech",
    json=gen_payload
)

with open("output.wav", "wb") as out:
    out.write(gen_resp.content)

The API accepts the same audio formats as the CLI, encoding them as data-URIs for HTTP transport as documented in docs/mcp.md lines 14–16.

Summary

  • Duration limits: 3 s minimum, 5–15 s optimal, 20 s maximum enforced in omnivoice/utils/audio.py
  • Format requirements: 16 kHz mono PCM preferred; WAV, MP3, FLAC, or base64 data-URI accepted
  • Audio quality: Clean, single-speaker recordings with minimal noise and target language matching
  • Volume standards: Normalized to approximately ‑23 LUFS to avoid usability errors
  • Error markers: CLONE_REF_UNUSABLE_MARKER and CLONE_REF_TOO_LONG_MARKER provide specific failure reasons
  • Implementation: Both CLI and HTTP API enforce these guidelines before cloning

Frequently Asked Questions

What is the minimum length for reference audio in VoiceStudio?

VoiceStudio requires a minimum of approximately 3 seconds of non-silent audio. Clips shorter than this lack sufficient phonetic diversity to capture the speaker's characteristics accurately. For optimal results, the README.md recommends 5–15 seconds of clear speech.

What file formats are supported for voice cloning?

VoiceStudio accepts WAV, MP3, and FLAC files directly. When using the MCP API endpoint at /v1/clone_voice, you can also submit audio as a base64-encoded data-URI. All formats are normalized internally to 16 kHz mono PCM regardless of input specification.

Why does VoiceStudio reject my reference audio as "empty"?

This error occurs when silence removal filters eliminate all audio content, leaving no usable sound. Common causes include extremely quiet recordings, files containing only noise, or clips with corrupted headers. Ensure your input has clear speech at normalized volume levels (around ‑23 LUFS) and minimal background noise to avoid triggering CLONE_REF_UNUSABLE_MARKER.

Can I use a video file or noisy recording for voice cloning?

No. Video files must be extracted to audio (WAV/MP3/FLAC) before submission. Noisy recordings with music, ambient sounds, or multiple speakers typically fail validation or produce poor clones. The system enforces strict single-speaker, low-noise requirements because background audio interferes with the model's ability to isolate and replicate the target voice's timbre.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →