Guidelines for Reference Audio in Voice Cloning: VoiceStudio Requirements and Best Practices
VoiceStudio requires reference audio between 5–15 seconds (minimum 3 seconds, maximum 20 seconds) in 16 kHz mono format with minimal background noise and matching target language to successfully clone a voice.
VoiceStudio (debpalash/VoiceStudio) performs voice cloning by analyzing a short reference clip to capture speaker characteristics. The system enforces strict validation rules in omnivoice/utils/audio.py to ensure quality, rejecting clips that violate technical constraints with specific error markers surfaced to the UI.
Technical Requirements for Reference Audio
VoiceStudio validates all incoming reference audio against hard constraints before processing. Violations trigger immediate failures with descriptive error messages.
Duration Constraints and Limits
The duration requirements follow a three-tier structure enforced during preprocessing:
- Minimum: approximately 3 seconds of non-silent audio (below this threshold, the clone lacks sufficient phonetic diversity)
- Optimal: 5–15 seconds for best quality results
- Maximum: strictly capped at 20 seconds (exceeding this triggers
CLONE_REF_TOO_LONG_MARKER)
As implemented in omnivoice/utils/audio.py lines 148–199, the system automatically removes leading and trailing silence before measuring duration. If silence removal leaves the clip empty, the operation fails with "Reference audio is empty after silence removal".
Audio Format Specifications
VoiceStudio accepts multiple input formats but normalizes them internally to 16 kHz mono PCM. The supported formats include:
- WAV (recommended)
- MP3
- FLAC
- Base64-encoded data-URI (for API submissions)
All stereo inputs are mixed down to mono during preprocessing. Higher sampling rates are automatically down-sampled to 16 kHz, but supplying native 16 kHz mono files reduces processing overhead.
Content Quality Guidelines
Beyond technical specifications, the acoustic content of the reference clip significantly impacts cloning fidelity.
Background Noise and Audio Environment
The reference recording must contain minimal background noise, no music, and no ambient sounds. The system detects usable audio segments, and excessive noise can trigger CLONE_REF_UNUSABLE_MARKER, indicating no usable sound remains after preprocessing. Record in quiet environments using close-mic techniques to isolate the target speaker.
Language Matching and Speaker Consistency
The spoken language in the reference audio must match the language of the text you intend to synthesize. VoiceStudio infers language characteristics from the reference clip, and mismatched language input degrades voice quality and pronunciation accuracy. The clip must feature a single, consistent speaker without overlapping voices or conversation.
Volume Normalization and Content Selection
Optimal reference audio maintains normalized loudness around ‑23 LUFS ( Loudness Units relative to Full Scale). Clips that are too quiet risk "no usable sound" errors, while clipped or distorted audio degrades timbre capture. Select neutral script content—such as a short paragraph read in a natural, conversational tone—rather than emotionally extreme or whispered speech. Expressive styles can be added later via instruct strings during generation.
Error Handling and Validation
VoiceStudio surfaces validation failures through specific error markers defined in the audio processing pipeline:
CLONE_REF_UNUSABLE_MARKER: Reference audio has no usable sound after silence removal or noise filteringCLONE_REF_TOO_LONG_MARKER: Audio exceeds the 20‑second limit
These markers propagate to the frontend UI, triggering user-facing toast messages that guide corrective action, as tested in frontend/src/test/errorToastUserFixable.test.jsx lines 25–50. Users receive immediate feedback to re-record or trim their samples rather than waiting for processing completion.
Practical Implementation Examples
CLI Usage with OmniVoice
Clone a voice from a local file using the command-line interface:
omnivoice clone --ref-audio path/to/voice.wav \
--text "Hello, this is my cloned voice." \
--output cloned.wav
The CLI automatically validates duration and format before submitting to the backend.
Python API Integration
Submit base64-encoded reference audio to the MCP clone_voice endpoint:
import requests
import base64
# Load and encode reference audio
with open("voice.wav", "rb") as f:
b64_audio = base64.b64encode(f.read()).decode()
payload = {
"ref_audio": f"data:audio/wav;base64,{b64_audio}",
"text": "Welcome to VoiceStudio!",
"language": "en"
}
resp = requests.post(
"http://localhost:3900/v1/clone_voice",
json=payload,
timeout=30
)
resp.raise_for_status()
profile_id = resp.json()["profile_id"]
Generate speech from the created profile:
gen_payload = {
"profile_id": profile_id,
"text": "This is the synthesized output using the cloned voice.",
"language": "en"
}
gen_resp = requests.post(
"http://localhost:3900/v1/generate_speech",
json=gen_payload
)
with open("output.wav", "wb") as out:
out.write(gen_resp.content)
The API accepts the same audio formats as the CLI, encoding them as data-URIs for HTTP transport as documented in docs/mcp.md lines 14–16.
Summary
- Duration limits: 3 s minimum, 5–15 s optimal, 20 s maximum enforced in
omnivoice/utils/audio.py - Format requirements: 16 kHz mono PCM preferred; WAV, MP3, FLAC, or base64 data-URI accepted
- Audio quality: Clean, single-speaker recordings with minimal noise and target language matching
- Volume standards: Normalized to approximately ‑23 LUFS to avoid usability errors
- Error markers:
CLONE_REF_UNUSABLE_MARKERandCLONE_REF_TOO_LONG_MARKERprovide specific failure reasons - Implementation: Both CLI and HTTP API enforce these guidelines before cloning
Frequently Asked Questions
What is the minimum length for reference audio in VoiceStudio?
VoiceStudio requires a minimum of approximately 3 seconds of non-silent audio. Clips shorter than this lack sufficient phonetic diversity to capture the speaker's characteristics accurately. For optimal results, the README.md recommends 5–15 seconds of clear speech.
What file formats are supported for voice cloning?
VoiceStudio accepts WAV, MP3, and FLAC files directly. When using the MCP API endpoint at /v1/clone_voice, you can also submit audio as a base64-encoded data-URI. All formats are normalized internally to 16 kHz mono PCM regardless of input specification.
Why does VoiceStudio reject my reference audio as "empty"?
This error occurs when silence removal filters eliminate all audio content, leaving no usable sound. Common causes include extremely quiet recordings, files containing only noise, or clips with corrupted headers. Ensure your input has clear speech at normalized volume levels (around ‑23 LUFS) and minimal background noise to avoid triggering CLONE_REF_UNUSABLE_MARKER.
Can I use a video file or noisy recording for voice cloning?
No. Video files must be extracted to audio (WAV/MP3/FLAC) before submission. Noisy recordings with music, ambient sounds, or multiple speakers typically fail validation or produce poor clones. The system enforces strict single-speaker, low-noise requirements because background audio interferes with the model's ability to isolate and replicate the target voice's timbre.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →