Supported LLM Fine-Tuning Tasks in Soup CLI: A Complete Guide
The Soup CLI supports 12+ distinct LLM fine-tuning tasks including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), reinforcement learning variants (ORPO, PPO, KTO, GRPO), unlearning (NPO/RMU), text-to-speech, distillation, and MoE routing, all accessible via the unified soup train --task <name> interface.
Soup provides a modular, extensible framework for adapting large language models to specific use cases. Every task is driven by the soup train command and configured through a centralized Pydantic v2 schema, enabling researchers to switch between training paradigms without rewriting data pipelines or training loops.
Core Supervised and Preference-Based Tasks
Supervised Fine-Tuning (SFT)
Supervised Fine-Tuning is the default behavior when running soup train. In src/soup_cli/config/schema.py, the configuration schema defines training.task = "sft", which triggers the standard next-token prediction loss using a causal language modeling data collator. This path is optimized for instruction-following datasets and supports LoRA adapters out of the box.
soup train --config soup.yaml # Implicitly uses task="sft"
Direct Preference Optimization (DPO)
Direct Preference Optimization eliminates the need for a separate reward model by optimizing directly against preference pairs. When invoked with --task dpo, the training loop imports the loss implementation from src/soup_cli/training/dpo.py and swaps the standard cross-entropy loss for the DPO objective.
soup train --config soup.yaml --task dpo --lr 5e-5 --epochs 3
Advanced Preference Optimization Methods
For reinforcement learning style alignment, Soup implements several policy-gradient and KL-regularized objectives selected via --task <name>:
- Ortho-Policy-Optimization (ORPO) – Optimizes the policy while maintaining orthogonality constraints.
- Proximal Policy Optimization (PPO) – Standard RLHF implementation with clipped objectives.
- K-to-One (KTO) – Simplifies preference learning by contrasting against a single negative.
- Simulated Preference Optimization (SIMPO) – Simulates preference pairs on-the-fly.
- Generalized Reinforcement-Learning Preference Optimization (GRPO) – Group-relative policy optimization for multi-turn reasoning.
Each implementation resides in its own module under src/soup_cli/training/ and registers itself with the task registry in src/soup_cli/training/registry.py.
Specialized Fine-Tuning Modes
Pre-training
The pre-training task (--task pretrain) uses the same data pipeline as SFT but disables LoRA adapter-only mode to allow full-model weight updates. This is intended for continued pre-training on large, raw corpora rather than instruction datasets.
Classification, Reranking, and Cross-Encoder Tasks
Soup supports head-only fine-tuning for retrieval-augmented generation workflows through three specific task types:
--task classifier– Standard multi-class or binary classification.--task reranker– Pairwise ranking loss for reordering retrieved documents.--task cross_encoder– Full cross-attention scoring between query and document pairs.
These tasks freeze the transformer backbone and train only a classification head using cross-entropy loss over task-specific labels.
Text-to-Speech (TTS)
The TTS task fine-tunes models for audio token generation. When running --task tts, Soup loads a codec-token data collator from the data pipeline and applies a regression loss on audio embeddings defined in src/soup_cli/training/tts.py.
soup train --config soup.yaml --task tts \
--audio-dir ./audio \
--model whisper-base
Unlearning and Model Editing
Soup provides unlearning capabilities through the unlearn task, which uses negative preference optimization (NPO), SimNPO, or regularized memory updates (RMU) to remove memorized information without full retraining. The logic is implemented in src/soup_cli/training/unlearn.py.
soup train --config soup.yaml --task unlearn \
--unlearn-strategy npo \
--data forget_set.jsonl
Advanced Architectural Patterns
Knowledge Distillation
The distillation task (--task distill) supports both token-level and sequence-level knowledge transfer from a teacher model. Set distill_mode=token for KL-divergence on per-token logits or distill_mode=sequence for sequence-level distribution matching. The training loop routes data through the teacher model for loss computation while updating the student.
Mixture-of-Experts (MoE) and MoLE
For sparse expert architectures, Soup supports:
--task moe_lora_routing– Trains only the router gating parameters while locking expert weights.--task mole– Implements "Mixture-of-LoRA-Experts" by fine-tuning the gating network over a set of LoRA adapters.
Both modes are handled in src/soup_cli/training/moe.py and enable efficient fine-tuning of massive MoE backbones on consumer hardware.
Retriever-Augmented Generation (RA-DIT)
The RA-DIT workflow (soup ra-dit) executes a two-stage training process where a retriever and generator are trained jointly. This command spawns a retriever-training run followed by a generator-training run, automatically wiring retrieval results into the generator's context window according to the protocol defined in docs/commands.md.
Workflow Automation and Specialized Inference
The CLI integrates synthetic data workflows that auto-generate training data before fine-tuning. Commands like soup data best-of-n and soup data evolve internally invoke soup train with the appropriate --task flag after data preparation steps, as documented in docs/commands.md at lines 100–108.
Additionally, Soup ships with an Audio-ASR inference task (soup infer --task asr) for Whisper-style models, implying a matching training task exists for automatic speech recognition fine-tuning.
Architectural Implementation Details
Soup's task system relies on three core architectural decisions:
Unified Config Schema – All tasks share the Pydantic v2 schema in src/soup_cli/config/schema.py. The training.task field serves as the single source of truth, with validation ensuring task-specific hyper-parameters are only accepted when the relevant task is active.
Modular Loss Registry – Each task registers its loss function, collator, and trainer via src/soup_cli/training/registry.py. The main training loop fetches implementations by task_name, enabling plug-and-play addition of new objectives.
Lazy Dependency Loading – Heavy libraries including torch, transformers, and peft are imported inside task functions rather than at module load time, keeping the CLI lightweight when executing non-training commands.
Summary
- Soup CLI supports 12+ LLM fine-tuning tasks ranging from standard SFT to advanced RL methods like GRPO and ORPO.
- All tasks are invoked via
soup train --task <name>with configuration centralized insrc/soup_cli/config/schema.py. - Preference optimization tasks (DPO, KTO, SIMPO) reside in
src/soup_cli/training/dpo.pyand related modules. - Specialized modes include TTS audio fine-tuning, unlearning (NPO/RMU), distillation, and MoE routing.
- Classification tasks (classifier, reranker, cross-encoder) use head-only training for RAG applications.
- The registry pattern in
src/soup_cli/training/registry.pyenables modular loss functions and data collators per task.
Frequently Asked Questions
What LLM fine-tuning tasks does Soup CLI support by default?
Soup CLI supports Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), ORPO, PPO, KTO, SIMPO, GRPO, pre-training, classification, reranking, cross-encoding, TTS, unlearning, distillation, MoE routing, and RA-DIT workflows. Each task is exposed through the soup train --task <name> argument as documented in docs/commands.md.
How does Soup CLI switch between different training objectives?
The CLI uses a unified configuration schema in src/soup_cli/config/schema.py where the training.task field determines which loss function and data collator to load from src/soup_cli/training/registry.py. Changing the task name automatically configures the training loop for that specific objective without requiring changes to the config file structure.
Can I add custom fine-tuning tasks to Soup CLI?
Yes. The modular registry system in src/soup_cli/training/registry.py allows you to register custom loss functions and collators by task name. You must implement the training logic in a new module under src/soup_cli/training/ and register it with the @register_task decorator, following the pattern used for DPO in src/soup_cli/training/dpo.py.
Does Soup CLI support full-model fine-tuning or only LoRA adapters?
Soup supports both modes. By default, most tasks use LoRA adapters for memory efficiency, but the pre-training task and certain classification tasks can perform full-model updates when adapter-only mode is disabled in the configuration. The --deepspeed and --fsdp flags enable scalable multi-GPU training for both partial and full fine-tuning scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →