# How to Fine-Tune OmniVoice Using the `omnivoice/training/` Pipeline: Complete Data Preparation and Execution Guide

> Learn how to fine-tune OmniVoice with this guide. Prepare your data with a JSONL manifest and run the training pipeline for custom voice model creation. Get step-by-step instructions.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Fine-tuning OmniVoice requires a JSONL manifest with 16‑kHz mono audio paths, transcripts, and speaker IDs, followed by running [`omnivoice/training/run_fine_tune.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/training/run_fine_tune.py) with your base checkpoint and hyper-parameter configuration.**

The `debpalash/VoiceStudio` repository embeds OmniVoice as a Git submodule under `omnivoice/`, providing a dedicated training pipeline for adapting the text-to-speech model to new speakers or domains. This guide walks through the exact data preparation requirements and execution steps based on the source code structure.

---

## Data Preparation Requirements for OmniVoice Fine-Tuning

OmniVoice's training pipeline in `omnivoice/training/` expects strictly formatted inputs. The [`run_fine_tune.py`](https://github.com/debpalash/VoiceStudio/blob/main/run_fine_tune.py) script parses a **metadata manifest** to locate audio files and associated labels.

### Required Dataset Format

| Component | Specification |
|-----------|---------------|
| **Audio files** | 16‑kHz mono WAV or lossless FLAC; single utterance per file; minimal background noise |
| **Transcripts** | UTF‑8 plain text matching spoken content; punctuation optional |
| **Speaker ID** | Alphanumeric identifier shared across all utterances from the same voice |
| **Metadata manifest** | JSONL file with fields: `audio_filepath`, `text`, `speaker_id`, `duration` |

### Manifest Structure

Each line in `train_manifest.jsonl` contains a JSON object:

```json
{"audio_filepath": "wav/spk01_001.wav", "text": "Hello, this is a sample utterance.", "speaker_id": "spk01", "duration": 3.42}

```

The `duration` field is optional—OmniVoice computes it during preprocessing if omitted.

### Optional Speaker Embeddings

For improved speaker consistency, pre-computed speaker embeddings may be included:

```json
{"audio_filepath": "wav/spk01_001.wav", "text": "Hello world.", "speaker_id": "spk01", "speaker_embedding": "embeddings/spk01_001.npy"}

```

These embeddings typically come from a speaker-verification model like ECAPA-TDNN.

### Recommended Directory Layout

```

omnivoice/
└── training/
   ├── data/
   │   ├── wav/
   │   │   ├── spk01_001.wav
   │   │   └── spk01_002.wav
   │   ├── embeddings/          # Optional

   │   │   └── spk01_001.npy
   │   └── train_manifest.jsonl
   └── configs/
       └── fine_tune.yaml

```

All paths in the manifest are **relative to the `data/` folder** or the working directory from which you launch training.

---

## Installation and Environment Setup

The OmniVoice training dependencies are managed through the VoiceStudio repository's submodule structure and UV lock file.

### Clone with Submodules

```bash
git clone --recurse-submodules https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio

```

The `--recurse-submodules` flag ensures `omnivoice/` is populated with the training code at `omnivoice/training/`.

### Install Dependencies

```bash

# Sync Python environment using UV

uv sync

# Install training-specific extras

pip install "omnivoice[train]"

```

The `uv.lock` file in the repository root pins reproducible dependency versions.

---

## Selecting a Base Checkpoint

Before fine-tuning, choose a pre-trained OmniVoice model from `omnivoice/models/`. The helper function in [`omnivoice/models/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/models/__init__.py) resolves checkpoint identifiers:

```python
from omnivoice.models import resolve_omnivoice_checkpoint

# Returns path to default checkpoint

base_ckpt = resolve_omnivoice_checkpoint()

# Or specify a variant

base_ckpt = resolve_omnivoice_checkpoint("omnivoice:large")

```

Common identifiers include `omnivoice:default`, `omnivoice:large`, and custom paths.

---

## Running the Fine-Tuning Script

The primary entry point is [`omnivoice/training/run_fine_tune.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/training/run_fine_tune.py), invoked as a module with CLI arguments.

### Basic Execution Command

```bash
python -m omnivoice.training.run_fine_tune \
    --base-checkpoint omnivoice:default \
    --manifest omnivoice/training/data/train_manifest.jsonl \
    --output-dir ./fine_tuned_model \
    --config omnivoice/training/configs/fine_tune.yaml \
    --batch-size 32 \
    --epochs 30 \
    --learning-rate 5e-5

```

### Core CLI Arguments

| Argument | Description |
|----------|-------------|
| `--base-checkpoint` | Pretrained model identifier or path |
| `--manifest` | Path to JSONL metadata file |
| `--output-dir` | Destination for checkpoints and logs |
| `--config` | YAML hyper-parameter configuration |
| `--batch-size` | Training batch size (default: 16) |
| `--epochs` | Total training epochs |
| `--learning-rate` | Initial learning rate for optimizer |
| `--gradient-accumulation-steps` | Effective batch size multiplier |
| `--warmup-steps` | Linear warmup for learning rate schedule |

### Configuration File ([`fine_tune.yaml`](https://github.com/debpalash/VoiceStudio/blob/main/fine_tune.yaml))

The YAML at [`omnivoice/training/configs/fine_tune.yaml`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/training/configs/fine_tune.yaml) controls:

```yaml
optimizer:
  type: AdamW
  betas: [0.9, 0.999]
  weight_decay: 0.01

scheduler:
  type: cosine_with_warmup
  warmup_ratio: 0.1

training:
  max_grad_norm: 1.0
  fp16: true
  dataloader_num_workers: 4

```

Override any value via CLI flags or by editing the YAML directly.

---

## Monitoring Training Progress

The [`run_fine_tune.py`](https://github.com/debpalash/VoiceStudio/blob/main/run_fine_tune.py) script outputs:

- **Console logs**: Loss values, learning rate, and step timing to `stdout`
- **TensorBoard events**: Written to `output-dir/logs/` for visualization
- **Checkpointing**: Models saved every epoch and at best validation loss

View training curves:

```bash
tensorboard --logdir ./fine_tuned_model/logs/

```

---

## Validating the Fine-Tuned Model

After training completes, verify speaker fidelity with inference:

```python
from omnivoice import OmniVoice

# Load from local checkpoint

model = OmniVoice(model_id="./fine_tuned_model", download_if_missing=False)

audio = model.generate(
    text="This is a validation test of the fine-tuned voice.",
    speaker_id="spk01"
)

audio.save("validation_output.wav")

```

Listen for natural prosody and accurate speaker timbre matching your training data.

---

## Deploying to VoiceStudio

Integrate the fine-tuned model into the VoiceStudio engine:

### File Placement

```bash
mkdir -p omnivoice_data/models/
cp -r ./fine_tuned_model omnivoice_data/models/my_fine_tuned

```

### Registration via API

```bash
curl -X POST http://localhost:8000/engines/register \
    -H "Content-Type: application/json" \
    -d '{"engine": "omnivoice", "model_id": "my_fine_tuned"}'

```

The model becomes selectable in the VoiceStudio UI under **OmniVoice → My Fine-Tuned**.

---

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`omnivoice/training/run_fine_tune.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/training/run_fine_tune.py) | Main training script entry point |
| [`omnivoice/training/configs/fine_tune.yaml`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/training/configs/fine_tune.yaml) | Default hyper-parameter configuration |
| [`omnivoice/models/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/models/__init__.py) | Checkpoint resolution utilities (`resolve_omnivoice_checkpoint`) |
| `omnivoice/training/data/` | Example manifests and dataset templates |
| [`tests/test_vendored_omnivoice_imports_1417.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_vendored_omnivoice_imports_1417.py) | Submodule import verification (diagnostic utility) |

These paths reflect the repository structure as implemented in `debpalash/VoiceStudio`, with OmniVoice vendored as a submodule.

---

## Summary

- **Data preparation**: Create 16‑kHz mono audio files and a JSONL manifest with `audio_filepath`, `text`, `speaker_id`, and optional `duration` or `speaker_embedding` fields
- **Environment**: Clone with `--recurse-submodules`, run `uv sync`, and install `omnivoice[train]`
- **Execution**: Call `python -m omnivoice.training.run_fine_tune` with `--base-checkpoint`, `--manifest`, `--config`, and training hyper-parameters
- **Validation**: Generate test audio using `OmniVoice(model_id=...)` with `download_if_missing=False`
- **Deployment**: Copy checkpoint to `omnivoice_data/models/` and register via the VoiceStudio API

---

## Frequently Asked Questions

### What audio format does OmniVoice training require?

OmniVoice expects **16‑kHz mono WAV or FLAC files**. Each file should contain a single clean utterance. The training pipeline resamples and validates audio automatically, but source quality directly impacts fine-tuning results.

### How much training data is needed for speaker fine-tuning?

Based on typical TTS fine-tuning practices and OmniVoice's architecture, **10–30 minutes of clean speech per speaker** yields usable results. For production-quality voice cloning, 1–2 hours provides better prosodic consistency. The manifest must include at least several hundred utterances for stable convergence.

### Can I fine-tune on multiple speakers simultaneously?

Yes. Include all speakers in a single `train_manifest.jsonl` with distinct `speaker_id` values. OmniVoice's conditioning mechanism separates speakers via ID lookup, and the model learns a shared acoustic space. Batch composition automatically mixes speakers during training.

### Where are training checkpoints saved and how do I resume?

Checkpoints write to the `--output-dir` path, with subdirectories `checkpoint-*/` containing `pytorch_model.bin` and [`training_args.json`](https://github.com/debpalash/VoiceStudio/blob/main/training_args.json). To resume, pass the checkpoint directory to `--base-checkpoint`:

```bash
python -m omnivoice.training.run_fine_tune \
    --base-checkpoint ./fine_tuned_model/checkpoint-1500 \
    --manifest ...  # remaining arguments

```