# Supported LLM Fine-Tuning Tasks in Soup CLI: A Complete Guide

> Explore 12+ LLM fine-tuning tasks in Soup CLI including SFT DPO ORPO PPO KTO GRPO unlearning text-to-speech distillation and MoE routing. Master your models with our comprehensive guide.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**The Soup CLI supports 12+ distinct LLM fine-tuning tasks including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), reinforcement learning variants (ORPO, PPO, KTO, GRPO), unlearning (NPO/RMU), text-to-speech, distillation, and MoE routing, all accessible via the unified `soup train --task <name>` interface.**

Soup provides a modular, extensible framework for adapting large language models to specific use cases. Every task is driven by the `soup train` command and configured through a centralized Pydantic v2 schema, enabling researchers to switch between training paradigms without rewriting data pipelines or training loops.

## Core Supervised and Preference-Based Tasks

### Supervised Fine-Tuning (SFT)

**Supervised Fine-Tuning** is the default behavior when running `soup train`. In [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py), the configuration schema defines `training.task = "sft"`, which triggers the standard next-token prediction loss using a causal language modeling data collator. This path is optimized for instruction-following datasets and supports LoRA adapters out of the box.

```bash
soup train --config soup.yaml  # Implicitly uses task="sft"

```

### Direct Preference Optimization (DPO)

**Direct Preference Optimization** eliminates the need for a separate reward model by optimizing directly against preference pairs. When invoked with `--task dpo`, the training loop imports the loss implementation from [`src/soup_cli/training/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/dpo.py) and swaps the standard cross-entropy loss for the DPO objective.

```bash
soup train --config soup.yaml --task dpo --lr 5e-5 --epochs 3

```

### Advanced Preference Optimization Methods

For reinforcement learning style alignment, Soup implements several policy-gradient and KL-regularized objectives selected via `--task <name>`:

- **Ortho-Policy-Optimization (ORPO)** – Optimizes the policy while maintaining orthogonality constraints.
- **Proximal Policy Optimization (PPO)** – Standard RLHF implementation with clipped objectives.
- **K-to-One (KTO)** – Simplifies preference learning by contrasting against a single negative.
- **Simulated Preference Optimization (SIMPO)** – Simulates preference pairs on-the-fly.
- **Generalized Reinforcement-Learning Preference Optimization (GRPO)** – Group-relative policy optimization for multi-turn reasoning.

Each implementation resides in its own module under `src/soup_cli/training/` and registers itself with the task registry in [`src/soup_cli/training/registry.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/registry.py).

## Specialized Fine-Tuning Modes

### Pre-training

The **pre-training** task (`--task pretrain`) uses the same data pipeline as SFT but disables LoRA adapter-only mode to allow full-model weight updates. This is intended for continued pre-training on large, raw corpora rather than instruction datasets.

### Classification, Reranking, and Cross-Encoder Tasks

Soup supports head-only fine-tuning for retrieval-augmented generation workflows through three specific task types:

- `--task classifier` – Standard multi-class or binary classification.
- `--task reranker` – Pairwise ranking loss for reordering retrieved documents.
- `--task cross_encoder` – Full cross-attention scoring between query and document pairs.

These tasks freeze the transformer backbone and train only a classification head using cross-entropy loss over task-specific labels.

### Text-to-Speech (TTS)

The **TTS** task fine-tunes models for audio token generation. When running `--task tts`, Soup loads a codec-token data collator from the data pipeline and applies a regression loss on audio embeddings defined in [`src/soup_cli/training/tts.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/tts.py).

```bash
soup train --config soup.yaml --task tts \
           --audio-dir ./audio \
           --model whisper-base

```

### Unlearning and Model Editing

Soup provides **unlearning** capabilities through the `unlearn` task, which uses negative preference optimization (NPO), SimNPO, or regularized memory updates (RMU) to remove memorized information without full retraining. The logic is implemented in [`src/soup_cli/training/unlearn.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/unlearn.py).

```bash
soup train --config soup.yaml --task unlearn \
           --unlearn-strategy npo \
           --data forget_set.jsonl

```

## Advanced Architectural Patterns

### Knowledge Distillation

The **distillation** task (`--task distill`) supports both token-level and sequence-level knowledge transfer from a teacher model. Set `distill_mode=token` for KL-divergence on per-token logits or `distill_mode=sequence` for sequence-level distribution matching. The training loop routes data through the teacher model for loss computation while updating the student.

### Mixture-of-Experts (MoE) and MoLE

For sparse expert architectures, Soup supports:

- `--task moe_lora_routing` – Trains only the router gating parameters while locking expert weights.
- `--task mole` – Implements "Mixture-of-LoRA-Experts" by fine-tuning the gating network over a set of LoRA adapters.

Both modes are handled in [`src/soup_cli/training/moe.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/moe.py) and enable efficient fine-tuning of massive MoE backbones on consumer hardware.

### Retriever-Augmented Generation (RA-DIT)

The **RA-DIT** workflow (`soup ra-dit`) executes a two-stage training process where a retriever and generator are trained jointly. This command spawns a retriever-training run followed by a generator-training run, automatically wiring retrieval results into the generator's context window according to the protocol defined in [`docs/commands.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/commands.md).

## Workflow Automation and Specialized Inference

The CLI integrates **synthetic data workflows** that auto-generate training data before fine-tuning. Commands like `soup data best-of-n` and `soup data evolve` internally invoke `soup train` with the appropriate `--task` flag after data preparation steps, as documented in [`docs/commands.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/commands.md) at lines 100–108.

Additionally, Soup ships with an **Audio-ASR** inference task (`soup infer --task asr`) for Whisper-style models, implying a matching training task exists for automatic speech recognition fine-tuning.

## Architectural Implementation Details

Soup's task system relies on three core architectural decisions:

**Unified Config Schema** – All tasks share the Pydantic v2 schema in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py). The `training.task` field serves as the single source of truth, with validation ensuring task-specific hyper-parameters are only accepted when the relevant task is active.

**Modular Loss Registry** – Each task registers its loss function, collator, and trainer via [`src/soup_cli/training/registry.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/registry.py). The main training loop fetches implementations by `task_name`, enabling plug-and-play addition of new objectives.

**Lazy Dependency Loading** – Heavy libraries including `torch`, `transformers`, and `peft` are imported inside task functions rather than at module load time, keeping the CLI lightweight when executing non-training commands.

## Summary

- Soup CLI supports **12+ LLM fine-tuning tasks** ranging from standard SFT to advanced RL methods like GRPO and ORPO.
- All tasks are invoked via `soup train --task <name>` with configuration centralized in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py).
- **Preference optimization** tasks (DPO, KTO, SIMPO) reside in [`src/soup_cli/training/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/dpo.py) and related modules.
- **Specialized modes** include TTS audio fine-tuning, unlearning (NPO/RMU), distillation, and MoE routing.
- **Classification tasks** (classifier, reranker, cross-encoder) use head-only training for RAG applications.
- The **registry pattern** in [`src/soup_cli/training/registry.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/registry.py) enables modular loss functions and data collators per task.

## Frequently Asked Questions

### What LLM fine-tuning tasks does Soup CLI support by default?

Soup CLI supports Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), ORPO, PPO, KTO, SIMPO, GRPO, pre-training, classification, reranking, cross-encoding, TTS, unlearning, distillation, MoE routing, and RA-DIT workflows. Each task is exposed through the `soup train --task <name>` argument as documented in [`docs/commands.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/commands.md).

### How does Soup CLI switch between different training objectives?

The CLI uses a unified configuration schema in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) where the `training.task` field determines which loss function and data collator to load from [`src/soup_cli/training/registry.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/registry.py). Changing the task name automatically configures the training loop for that specific objective without requiring changes to the config file structure.

### Can I add custom fine-tuning tasks to Soup CLI?

Yes. The modular registry system in [`src/soup_cli/training/registry.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/registry.py) allows you to register custom loss functions and collators by task name. You must implement the training logic in a new module under `src/soup_cli/training/` and register it with the `@register_task` decorator, following the pattern used for DPO in [`src/soup_cli/training/dpo.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/training/dpo.py).

### Does Soup CLI support full-model fine-tuning or only LoRA adapters?

Soup supports both modes. By default, most tasks use LoRA adapters for memory efficiency, but the pre-training task and certain classification tasks can perform full-model updates when adapter-only mode is disabled in the configuration. The `--deepspeed` and `--fsdp` flags enable scalable multi-GPU training for both partial and full fine-tuning scenarios.