How to Perform Supervised Fine-Tuning (SFT) with Soup CLI: A Complete Guide
Supervised Fine-Tuning (SFT) with the Soup CLI is performed by running soup train --config soup.yaml after initializing a configuration with soup init --task sft, which invokes the SFTTrainerWrapper class in src/soup_cli/trainer/sft.py to handle streaming data loading, LoRA wrapping, and checkpoint management.
The MakazhanAlpamys/Soup repository provides an end-to-end command-line interface for adapting pretrained language models to downstream tasks using labeled datasets. This guide walks you through the complete workflow for Supervised Fine-Tuning (SFT) with Soup CLI, from configuration to deployment.
What is Supervised Fine-Tuning in Soup?
Supervised Fine-Tuning (SFT) adapts a pretrained model to specific tasks using labeled prompt-response pairs. In the Soup ecosystem, the SFTTrainerWrapper class ([src/soup_cli/trainer/sft.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/sft.py)) orchestrates this process by integrating with the Pydantic configuration schema defined in [src/soup_cli/config/schema.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py).
The trainer inherits from StreamingSetupMixin, which enables dynamic batch sizing, automatic VRAM probing, and memory-efficient streaming of large datasets without full materialization. When you execute soup train, the CLI resolves the task: sft configuration to instantiate this wrapper and begins the training loop.
Prerequisites and Configuration Setup
Initializing the SFT Configuration
Begin by scaffolding a configuration file that defines your SFT task. The soup init command generates a soup.yaml file validated against the single source-of-truth schema:
soup init --template chat \
--model meta-llama/Llama-3.1-8B \
--task sft \
--output soup.yaml
This creates a configuration file containing a task: sft block and placeholders for data paths. All configuration fields—including model identifiers, training hyperparameters, and LoRA settings—are strictly validated by the Pydantic schema before training begins.
Preparing Your Supervised Dataset
Soup CLI expects supervised datasets in JSONL, CSV, or Hugging Face Hub formats where each record contains a prompt and a response field:
{ "prompt": "Translate English to French: Hello", "response": "Bonjour" }
Validate your local dataset before training:
soup data validate my_data.jsonl
Once validated, reference the dataset in soup.yaml:
data:
train: my_data.jsonl
format: chatml
The SFTTrainerWrapper streams this data during training, processing batches through a dataloader that yields (input_ids, labels) pairs with loss computed exclusively on response tokens when training.train_on_responses_only: true is set.
Configuring Model and Training Parameters
Selecting Base Models and LoRA Adapters
Specify your base model using the model field in the configuration. For parameter-efficient fine-tuning, add a lora block:
model: meta-llama/Llama-3.1-8B
lora:
r: 64
alpha: 16
target_modules:
- q_proj
- k_proj
- v_proj
- o_proj
When LoRA configuration is present, the trainer calls get_peft_model inside the SFT wrapper to freeze base model weights and train only the adapter parameters. The trainer supports both the transformers backend (default) and the mlx backend for Apple Silicon, detected at runtime based on your environment.
Hyperparameter Configuration
Edit training parameters under the training section in soup.yaml:
training:
lr: 2e-4
epochs: 3
batch_size: 4
train_on_responses_only: true
These fields map directly to Hugging Face TrainingArguments after schema validation. The SFTTrainerWrapper uses CrossEntropyLoss over response tokens during the warmup and training phases, with automatic gradient accumulation and mixed-precision support handled by the underlying training framework.
Running the SFT Training Loop
Execute the training run with a single command:
soup train --config soup.yaml \
--gpus auto \
--tensorboard
During execution, the CLI performs the following operations according to the MakazhanAlpamys/Soup source code:
- Configuration Parsing: Validates
soup.yamlagainst the Pydantic schema insrc/soup_cli/config/schema.py - Trainer Instantiation: Creates
SFTTrainerWrapperwithStreamingSetupMixincapabilities - Model Loading: Loads the base model via
transformers(ormlxon Apple Silicon) and wraps with PEFT if LoRA is configured - Streaming Data: Initializes a streaming dataloader that probes available VRAM and adjusts batch sizing dynamically
- Training Loop: Performs single-epoch warmup followed by full training, computing loss only on response tokens
- Checkpointing: Writes checkpoints to
output/checkpoint-{step}/directories at configured intervals
The trainer emits Rich-formatted logs to the terminal and optionally streams metrics to TensorBoard for real-time visualization.
Monitoring, Evaluation, and Publishing
Real-time Monitoring with TensorBoard
Enable TensorBoard logging using the --tensorboard flag during training. The SFTTrainerWrapper emits loss curves, learning rate schedules, and throughput metrics to the log directory, viewable via:
tensorboard --logdir output/logs
Every SFT run also generates a repro-receipt.json file containing complete configuration snapshots, dependency versions, and training cost estimates produced by soup cost for full reproducibility.
Inference and Model Serving
After training completes, test the fine-tuned model using local inference:
soup infer --model ./output \
--input <(echo '{"prompt":"What is the capital of France?"}') \
--output result.jsonl
Deploy the model as an OpenAI-compatible API server:
soup serve --model ./output --port 8000
The server exposes endpoints at http://localhost:8000/v1/chat/completions, allowing immediate integration with existing tooling.
Publishing to Hugging Face Hub
Upload your fine-tuned checkpoint with automatic model card generation:
soup push --model ./output \
--repo username/llama-3.1-sft \
--card $(soup card output)
The push command reads checkpoint metadata and the reproducibility receipt to generate a comprehensive model card documenting training parameters, dataset provenance, and evaluation metrics.
Advanced SFT Features
Backend Agnostic Execution: The SFTTrainerWrapper detects the compute backend at runtime, automatically adapting the data loader and optimization routines for NVIDIA GPUs (via transformers) or Apple Silicon (via mlx).
Telemetry and Cost Estimation: Each run writes training cost estimates and resource utilization statistics to soup cost output, enabling budget planning for large-scale fine-tuning operations.
BitNet Support: For specialized quantization requirements, the generic SFTTrainerWrapper architecture is subclassed in src/soup_cli/trainer/bitnet.py to support BitNet training while maintaining the same CLI interface.
Summary
- Initialize configurations using
soup init --task sftto generate validatedsoup.yamlfiles conforming to the Pydantic schema insrc/soup_cli/config/schema.py - Prepare datasets in JSONL format with
promptandresponsefields, validating them viasoup data validatebefore training - Enable efficient tuning by configuring LoRA adapters in the config file, which triggers
get_peft_modelwrapping insideSFTTrainerWrapper - Execute training with
soup train --config soup.yaml, leveraging streaming data loading and automatic VRAM management fromStreamingSetupMixin - Deploy and share using
soup servefor local APIs orsoup pushto publish reproducible checkpoints to Hugging Face Hub with auto-generated model cards
Frequently Asked Questions
What data format does Soup CLI require for SFT?
Soup CLI accepts JSONL, CSV, or Hugging Face Hub datasets where each row contains a prompt field and a response field. The soup data validate command checks format compliance before training begins, and the SFTTrainerWrapper streams this data during the training loop without loading the entire dataset into memory.
How do I enable LoRA for parameter-efficient fine-tuning?
Add a lora: block to your soup.yaml configuration specifying the rank (r), alpha (alpha), and target modules. When present, the SFTTrainerWrapper automatically calls get_peft_model to wrap the base model, enabling adapter training while keeping base parameters frozen.
Can I run SFT on Apple Silicon with Soup CLI?
Yes. Set backend: mlx in your configuration file. The SFTTrainerWrapper detects the backend at runtime and adapts the data loader and training operations to use the MLX framework instead of standard Transformers, enabling native execution on Apple Silicon devices.
Where are training checkpoints and logs stored?
By default, checkpoints are written to the output/ directory (e.g., output/checkpoint-500/), while logs and TensorBoard events are stored in output/logs/. Each run also generates a repro-receipt.json in the output directory containing full configuration snapshots for reproducibility.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →