How to Fine-Tune OmniVoice Using the `omnivoice/training/` Pipeline: Complete Data Preparation and Execution Guide

Fine-tuning OmniVoice requires a JSONL manifest with 16‑kHz mono audio paths, transcripts, and speaker IDs, followed by running omnivoice/training/run_fine_tune.py with your base checkpoint and hyper-parameter configuration.

The debpalash/VoiceStudio repository embeds OmniVoice as a Git submodule under omnivoice/, providing a dedicated training pipeline for adapting the text-to-speech model to new speakers or domains. This guide walks through the exact data preparation requirements and execution steps based on the source code structure.


Data Preparation Requirements for OmniVoice Fine-Tuning

OmniVoice's training pipeline in omnivoice/training/ expects strictly formatted inputs. The run_fine_tune.py script parses a metadata manifest to locate audio files and associated labels.

Required Dataset Format

Component Specification
Audio files 16‑kHz mono WAV or lossless FLAC; single utterance per file; minimal background noise
Transcripts UTF‑8 plain text matching spoken content; punctuation optional
Speaker ID Alphanumeric identifier shared across all utterances from the same voice
Metadata manifest JSONL file with fields: audio_filepath, text, speaker_id, duration

Manifest Structure

Each line in train_manifest.jsonl contains a JSON object:

{"audio_filepath": "wav/spk01_001.wav", "text": "Hello, this is a sample utterance.", "speaker_id": "spk01", "duration": 3.42}

The duration field is optional—OmniVoice computes it during preprocessing if omitted.

Optional Speaker Embeddings

For improved speaker consistency, pre-computed speaker embeddings may be included:

{"audio_filepath": "wav/spk01_001.wav", "text": "Hello world.", "speaker_id": "spk01", "speaker_embedding": "embeddings/spk01_001.npy"}

These embeddings typically come from a speaker-verification model like ECAPA-TDNN.


omnivoice/
└── training/
   ├── data/
   │   ├── wav/
   │   │   ├── spk01_001.wav
   │   │   └── spk01_002.wav
   │   ├── embeddings/          # Optional

   │   │   └── spk01_001.npy
   │   └── train_manifest.jsonl
   └── configs/
       └── fine_tune.yaml

All paths in the manifest are relative to the data/ folder or the working directory from which you launch training.


Installation and Environment Setup

The OmniVoice training dependencies are managed through the VoiceStudio repository's submodule structure and UV lock file.

Clone with Submodules

git clone --recurse-submodules https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio

The --recurse-submodules flag ensures omnivoice/ is populated with the training code at omnivoice/training/.

Install Dependencies


# Sync Python environment using UV

uv sync

# Install training-specific extras

pip install "omnivoice[train]"

The uv.lock file in the repository root pins reproducible dependency versions.


Selecting a Base Checkpoint

Before fine-tuning, choose a pre-trained OmniVoice model from omnivoice/models/. The helper function in omnivoice/models/__init__.py resolves checkpoint identifiers:

from omnivoice.models import resolve_omnivoice_checkpoint

# Returns path to default checkpoint

base_ckpt = resolve_omnivoice_checkpoint()

# Or specify a variant

base_ckpt = resolve_omnivoice_checkpoint("omnivoice:large")

Common identifiers include omnivoice:default, omnivoice:large, and custom paths.


Running the Fine-Tuning Script

The primary entry point is omnivoice/training/run_fine_tune.py, invoked as a module with CLI arguments.

Basic Execution Command

python -m omnivoice.training.run_fine_tune \
    --base-checkpoint omnivoice:default \
    --manifest omnivoice/training/data/train_manifest.jsonl \
    --output-dir ./fine_tuned_model \
    --config omnivoice/training/configs/fine_tune.yaml \
    --batch-size 32 \
    --epochs 30 \
    --learning-rate 5e-5

Core CLI Arguments

Argument Description
--base-checkpoint Pretrained model identifier or path
--manifest Path to JSONL metadata file
--output-dir Destination for checkpoints and logs
--config YAML hyper-parameter configuration
--batch-size Training batch size (default: 16)
--epochs Total training epochs
--learning-rate Initial learning rate for optimizer
--gradient-accumulation-steps Effective batch size multiplier
--warmup-steps Linear warmup for learning rate schedule

Configuration File (fine_tune.yaml)

The YAML at omnivoice/training/configs/fine_tune.yaml controls:

optimizer:
  type: AdamW
  betas: [0.9, 0.999]
  weight_decay: 0.01

scheduler:
  type: cosine_with_warmup
  warmup_ratio: 0.1

training:
  max_grad_norm: 1.0
  fp16: true
  dataloader_num_workers: 4

Override any value via CLI flags or by editing the YAML directly.


Monitoring Training Progress

The run_fine_tune.py script outputs:

  • Console logs: Loss values, learning rate, and step timing to stdout
  • TensorBoard events: Written to output-dir/logs/ for visualization
  • Checkpointing: Models saved every epoch and at best validation loss

View training curves:

tensorboard --logdir ./fine_tuned_model/logs/

Validating the Fine-Tuned Model

After training completes, verify speaker fidelity with inference:

from omnivoice import OmniVoice

# Load from local checkpoint

model = OmniVoice(model_id="./fine_tuned_model", download_if_missing=False)

audio = model.generate(
    text="This is a validation test of the fine-tuned voice.",
    speaker_id="spk01"
)

audio.save("validation_output.wav")

Listen for natural prosody and accurate speaker timbre matching your training data.


Deploying to VoiceStudio

Integrate the fine-tuned model into the VoiceStudio engine:

File Placement

mkdir -p omnivoice_data/models/
cp -r ./fine_tuned_model omnivoice_data/models/my_fine_tuned

Registration via API

curl -X POST http://localhost:8000/engines/register \
    -H "Content-Type: application/json" \
    -d '{"engine": "omnivoice", "model_id": "my_fine_tuned"}'

The model becomes selectable in the VoiceStudio UI under OmniVoice → My Fine-Tuned.


Key Source Files Reference

File Purpose
omnivoice/training/run_fine_tune.py Main training script entry point
omnivoice/training/configs/fine_tune.yaml Default hyper-parameter configuration
omnivoice/models/__init__.py Checkpoint resolution utilities (resolve_omnivoice_checkpoint)
omnivoice/training/data/ Example manifests and dataset templates
tests/test_vendored_omnivoice_imports_1417.py Submodule import verification (diagnostic utility)

These paths reflect the repository structure as implemented in debpalash/VoiceStudio, with OmniVoice vendored as a submodule.


Summary

  • Data preparation: Create 16‑kHz mono audio files and a JSONL manifest with audio_filepath, text, speaker_id, and optional duration or speaker_embedding fields
  • Environment: Clone with --recurse-submodules, run uv sync, and install omnivoice[train]
  • Execution: Call python -m omnivoice.training.run_fine_tune with --base-checkpoint, --manifest, --config, and training hyper-parameters
  • Validation: Generate test audio using OmniVoice(model_id=...) with download_if_missing=False
  • Deployment: Copy checkpoint to omnivoice_data/models/ and register via the VoiceStudio API

Frequently Asked Questions

What audio format does OmniVoice training require?

OmniVoice expects 16‑kHz mono WAV or FLAC files. Each file should contain a single clean utterance. The training pipeline resamples and validates audio automatically, but source quality directly impacts fine-tuning results.

How much training data is needed for speaker fine-tuning?

Based on typical TTS fine-tuning practices and OmniVoice's architecture, 10–30 minutes of clean speech per speaker yields usable results. For production-quality voice cloning, 1–2 hours provides better prosodic consistency. The manifest must include at least several hundred utterances for stable convergence.

Can I fine-tune on multiple speakers simultaneously?

Yes. Include all speakers in a single train_manifest.jsonl with distinct speaker_id values. OmniVoice's conditioning mechanism separates speakers via ID lookup, and the model learns a shared acoustic space. Batch composition automatically mixes speakers during training.

Where are training checkpoints saved and how do I resume?

Checkpoints write to the --output-dir path, with subdirectories checkpoint-*/ containing pytorch_model.bin and training_args.json. To resume, pass the checkpoint directory to --base-checkpoint:

python -m omnivoice.training.run_fine_tune \
    --base-checkpoint ./fine_tuned_model/checkpoint-1500 \
    --manifest ...  # remaining arguments

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →