How to Fine-Tune OmniVoice Using the `omnivoice/training/` Pipeline: Complete Data Preparation and Execution Guide
Fine-tuning OmniVoice requires a JSONL manifest with 16‑kHz mono audio paths, transcripts, and speaker IDs, followed by running omnivoice/training/run_fine_tune.py with your base checkpoint and hyper-parameter configuration.
The debpalash/VoiceStudio repository embeds OmniVoice as a Git submodule under omnivoice/, providing a dedicated training pipeline for adapting the text-to-speech model to new speakers or domains. This guide walks through the exact data preparation requirements and execution steps based on the source code structure.
Data Preparation Requirements for OmniVoice Fine-Tuning
OmniVoice's training pipeline in omnivoice/training/ expects strictly formatted inputs. The run_fine_tune.py script parses a metadata manifest to locate audio files and associated labels.
Required Dataset Format
| Component | Specification |
|---|---|
| Audio files | 16‑kHz mono WAV or lossless FLAC; single utterance per file; minimal background noise |
| Transcripts | UTF‑8 plain text matching spoken content; punctuation optional |
| Speaker ID | Alphanumeric identifier shared across all utterances from the same voice |
| Metadata manifest | JSONL file with fields: audio_filepath, text, speaker_id, duration |
Manifest Structure
Each line in train_manifest.jsonl contains a JSON object:
{"audio_filepath": "wav/spk01_001.wav", "text": "Hello, this is a sample utterance.", "speaker_id": "spk01", "duration": 3.42}
The duration field is optional—OmniVoice computes it during preprocessing if omitted.
Optional Speaker Embeddings
For improved speaker consistency, pre-computed speaker embeddings may be included:
{"audio_filepath": "wav/spk01_001.wav", "text": "Hello world.", "speaker_id": "spk01", "speaker_embedding": "embeddings/spk01_001.npy"}
These embeddings typically come from a speaker-verification model like ECAPA-TDNN.
Recommended Directory Layout
omnivoice/
└── training/
├── data/
│ ├── wav/
│ │ ├── spk01_001.wav
│ │ └── spk01_002.wav
│ ├── embeddings/ # Optional
│ │ └── spk01_001.npy
│ └── train_manifest.jsonl
└── configs/
└── fine_tune.yaml
All paths in the manifest are relative to the data/ folder or the working directory from which you launch training.
Installation and Environment Setup
The OmniVoice training dependencies are managed through the VoiceStudio repository's submodule structure and UV lock file.
Clone with Submodules
git clone --recurse-submodules https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
The --recurse-submodules flag ensures omnivoice/ is populated with the training code at omnivoice/training/.
Install Dependencies
# Sync Python environment using UV
uv sync
# Install training-specific extras
pip install "omnivoice[train]"
The uv.lock file in the repository root pins reproducible dependency versions.
Selecting a Base Checkpoint
Before fine-tuning, choose a pre-trained OmniVoice model from omnivoice/models/. The helper function in omnivoice/models/__init__.py resolves checkpoint identifiers:
from omnivoice.models import resolve_omnivoice_checkpoint
# Returns path to default checkpoint
base_ckpt = resolve_omnivoice_checkpoint()
# Or specify a variant
base_ckpt = resolve_omnivoice_checkpoint("omnivoice:large")
Common identifiers include omnivoice:default, omnivoice:large, and custom paths.
Running the Fine-Tuning Script
The primary entry point is omnivoice/training/run_fine_tune.py, invoked as a module with CLI arguments.
Basic Execution Command
python -m omnivoice.training.run_fine_tune \
--base-checkpoint omnivoice:default \
--manifest omnivoice/training/data/train_manifest.jsonl \
--output-dir ./fine_tuned_model \
--config omnivoice/training/configs/fine_tune.yaml \
--batch-size 32 \
--epochs 30 \
--learning-rate 5e-5
Core CLI Arguments
| Argument | Description |
|---|---|
--base-checkpoint |
Pretrained model identifier or path |
--manifest |
Path to JSONL metadata file |
--output-dir |
Destination for checkpoints and logs |
--config |
YAML hyper-parameter configuration |
--batch-size |
Training batch size (default: 16) |
--epochs |
Total training epochs |
--learning-rate |
Initial learning rate for optimizer |
--gradient-accumulation-steps |
Effective batch size multiplier |
--warmup-steps |
Linear warmup for learning rate schedule |
Configuration File (fine_tune.yaml)
The YAML at omnivoice/training/configs/fine_tune.yaml controls:
optimizer:
type: AdamW
betas: [0.9, 0.999]
weight_decay: 0.01
scheduler:
type: cosine_with_warmup
warmup_ratio: 0.1
training:
max_grad_norm: 1.0
fp16: true
dataloader_num_workers: 4
Override any value via CLI flags or by editing the YAML directly.
Monitoring Training Progress
The run_fine_tune.py script outputs:
- Console logs: Loss values, learning rate, and step timing to
stdout - TensorBoard events: Written to
output-dir/logs/for visualization - Checkpointing: Models saved every epoch and at best validation loss
View training curves:
tensorboard --logdir ./fine_tuned_model/logs/
Validating the Fine-Tuned Model
After training completes, verify speaker fidelity with inference:
from omnivoice import OmniVoice
# Load from local checkpoint
model = OmniVoice(model_id="./fine_tuned_model", download_if_missing=False)
audio = model.generate(
text="This is a validation test of the fine-tuned voice.",
speaker_id="spk01"
)
audio.save("validation_output.wav")
Listen for natural prosody and accurate speaker timbre matching your training data.
Deploying to VoiceStudio
Integrate the fine-tuned model into the VoiceStudio engine:
File Placement
mkdir -p omnivoice_data/models/
cp -r ./fine_tuned_model omnivoice_data/models/my_fine_tuned
Registration via API
curl -X POST http://localhost:8000/engines/register \
-H "Content-Type: application/json" \
-d '{"engine": "omnivoice", "model_id": "my_fine_tuned"}'
The model becomes selectable in the VoiceStudio UI under OmniVoice → My Fine-Tuned.
Key Source Files Reference
| File | Purpose |
|---|---|
omnivoice/training/run_fine_tune.py |
Main training script entry point |
omnivoice/training/configs/fine_tune.yaml |
Default hyper-parameter configuration |
omnivoice/models/__init__.py |
Checkpoint resolution utilities (resolve_omnivoice_checkpoint) |
omnivoice/training/data/ |
Example manifests and dataset templates |
tests/test_vendored_omnivoice_imports_1417.py |
Submodule import verification (diagnostic utility) |
These paths reflect the repository structure as implemented in debpalash/VoiceStudio, with OmniVoice vendored as a submodule.
Summary
- Data preparation: Create 16‑kHz mono audio files and a JSONL manifest with
audio_filepath,text,speaker_id, and optionaldurationorspeaker_embeddingfields - Environment: Clone with
--recurse-submodules, runuv sync, and installomnivoice[train] - Execution: Call
python -m omnivoice.training.run_fine_tunewith--base-checkpoint,--manifest,--config, and training hyper-parameters - Validation: Generate test audio using
OmniVoice(model_id=...)withdownload_if_missing=False - Deployment: Copy checkpoint to
omnivoice_data/models/and register via the VoiceStudio API
Frequently Asked Questions
What audio format does OmniVoice training require?
OmniVoice expects 16‑kHz mono WAV or FLAC files. Each file should contain a single clean utterance. The training pipeline resamples and validates audio automatically, but source quality directly impacts fine-tuning results.
How much training data is needed for speaker fine-tuning?
Based on typical TTS fine-tuning practices and OmniVoice's architecture, 10–30 minutes of clean speech per speaker yields usable results. For production-quality voice cloning, 1–2 hours provides better prosodic consistency. The manifest must include at least several hundred utterances for stable convergence.
Can I fine-tune on multiple speakers simultaneously?
Yes. Include all speakers in a single train_manifest.jsonl with distinct speaker_id values. OmniVoice's conditioning mechanism separates speakers via ID lookup, and the model learns a shared acoustic space. Batch composition automatically mixes speakers during training.
Where are training checkpoints saved and how do I resume?
Checkpoints write to the --output-dir path, with subdirectories checkpoint-*/ containing pytorch_model.bin and training_args.json. To resume, pass the checkpoint directory to --base-checkpoint:
python -m omnivoice.training.run_fine_tune \
--base-checkpoint ./fine_tuned_model/checkpoint-1500 \
--manifest ... # remaining arguments
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →