How to Perform Supervised Fine-Tuning (SFT) with Soup CLI: A Complete Guide

Supervised Fine-Tuning (SFT) with the Soup CLI is performed by running soup train --config soup.yaml after initializing a configuration with soup init --task sft, which invokes the SFTTrainerWrapper class in src/soup_cli/trainer/sft.py to handle streaming data loading, LoRA wrapping, and checkpoint management.

The MakazhanAlpamys/Soup repository provides an end-to-end command-line interface for adapting pretrained language models to downstream tasks using labeled datasets. This guide walks you through the complete workflow for Supervised Fine-Tuning (SFT) with Soup CLI, from configuration to deployment.

What is Supervised Fine-Tuning in Soup?

Supervised Fine-Tuning (SFT) adapts a pretrained model to specific tasks using labeled prompt-response pairs. In the Soup ecosystem, the SFTTrainerWrapper class ([src/soup_cli/trainer/sft.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/sft.py)) orchestrates this process by integrating with the Pydantic configuration schema defined in [src/soup_cli/config/schema.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py).

The trainer inherits from StreamingSetupMixin, which enables dynamic batch sizing, automatic VRAM probing, and memory-efficient streaming of large datasets without full materialization. When you execute soup train, the CLI resolves the task: sft configuration to instantiate this wrapper and begins the training loop.

Prerequisites and Configuration Setup

Initializing the SFT Configuration

Begin by scaffolding a configuration file that defines your SFT task. The soup init command generates a soup.yaml file validated against the single source-of-truth schema:

soup init --template chat \
    --model meta-llama/Llama-3.1-8B \
    --task sft \
    --output soup.yaml

This creates a configuration file containing a task: sft block and placeholders for data paths. All configuration fields—including model identifiers, training hyperparameters, and LoRA settings—are strictly validated by the Pydantic schema before training begins.

Preparing Your Supervised Dataset

Soup CLI expects supervised datasets in JSONL, CSV, or Hugging Face Hub formats where each record contains a prompt and a response field:

{ "prompt": "Translate English to French: Hello", "response": "Bonjour" }

Validate your local dataset before training:

soup data validate my_data.jsonl

Once validated, reference the dataset in soup.yaml:

data:
  train: my_data.jsonl
  format: chatml

The SFTTrainerWrapper streams this data during training, processing batches through a dataloader that yields (input_ids, labels) pairs with loss computed exclusively on response tokens when training.train_on_responses_only: true is set.

Configuring Model and Training Parameters

Selecting Base Models and LoRA Adapters

Specify your base model using the model field in the configuration. For parameter-efficient fine-tuning, add a lora block:

model: meta-llama/Llama-3.1-8B
lora:
  r: 64
  alpha: 16
  target_modules:
    - q_proj
    - k_proj
    - v_proj
    - o_proj

When LoRA configuration is present, the trainer calls get_peft_model inside the SFT wrapper to freeze base model weights and train only the adapter parameters. The trainer supports both the transformers backend (default) and the mlx backend for Apple Silicon, detected at runtime based on your environment.

Hyperparameter Configuration

Edit training parameters under the training section in soup.yaml:

training:
  lr: 2e-4
  epochs: 3
  batch_size: 4
  train_on_responses_only: true

These fields map directly to Hugging Face TrainingArguments after schema validation. The SFTTrainerWrapper uses CrossEntropyLoss over response tokens during the warmup and training phases, with automatic gradient accumulation and mixed-precision support handled by the underlying training framework.

Running the SFT Training Loop

Execute the training run with a single command:

soup train --config soup.yaml \
    --gpus auto \
    --tensorboard

During execution, the CLI performs the following operations according to the MakazhanAlpamys/Soup source code:

  1. Configuration Parsing: Validates soup.yaml against the Pydantic schema in src/soup_cli/config/schema.py
  2. Trainer Instantiation: Creates SFTTrainerWrapper with StreamingSetupMixin capabilities
  3. Model Loading: Loads the base model via transformers (or mlx on Apple Silicon) and wraps with PEFT if LoRA is configured
  4. Streaming Data: Initializes a streaming dataloader that probes available VRAM and adjusts batch sizing dynamically
  5. Training Loop: Performs single-epoch warmup followed by full training, computing loss only on response tokens
  6. Checkpointing: Writes checkpoints to output/checkpoint-{step}/ directories at configured intervals

The trainer emits Rich-formatted logs to the terminal and optionally streams metrics to TensorBoard for real-time visualization.

Monitoring, Evaluation, and Publishing

Real-time Monitoring with TensorBoard

Enable TensorBoard logging using the --tensorboard flag during training. The SFTTrainerWrapper emits loss curves, learning rate schedules, and throughput metrics to the log directory, viewable via:

tensorboard --logdir output/logs

Every SFT run also generates a repro-receipt.json file containing complete configuration snapshots, dependency versions, and training cost estimates produced by soup cost for full reproducibility.

Inference and Model Serving

After training completes, test the fine-tuned model using local inference:

soup infer --model ./output \
    --input <(echo '{"prompt":"What is the capital of France?"}') \
    --output result.jsonl

Deploy the model as an OpenAI-compatible API server:

soup serve --model ./output --port 8000

The server exposes endpoints at http://localhost:8000/v1/chat/completions, allowing immediate integration with existing tooling.

Publishing to Hugging Face Hub

Upload your fine-tuned checkpoint with automatic model card generation:

soup push --model ./output \
    --repo username/llama-3.1-sft \
    --card $(soup card output)

The push command reads checkpoint metadata and the reproducibility receipt to generate a comprehensive model card documenting training parameters, dataset provenance, and evaluation metrics.

Advanced SFT Features

Backend Agnostic Execution: The SFTTrainerWrapper detects the compute backend at runtime, automatically adapting the data loader and optimization routines for NVIDIA GPUs (via transformers) or Apple Silicon (via mlx).

Telemetry and Cost Estimation: Each run writes training cost estimates and resource utilization statistics to soup cost output, enabling budget planning for large-scale fine-tuning operations.

BitNet Support: For specialized quantization requirements, the generic SFTTrainerWrapper architecture is subclassed in src/soup_cli/trainer/bitnet.py to support BitNet training while maintaining the same CLI interface.

Summary

  • Initialize configurations using soup init --task sft to generate validated soup.yaml files conforming to the Pydantic schema in src/soup_cli/config/schema.py
  • Prepare datasets in JSONL format with prompt and response fields, validating them via soup data validate before training
  • Enable efficient tuning by configuring LoRA adapters in the config file, which triggers get_peft_model wrapping inside SFTTrainerWrapper
  • Execute training with soup train --config soup.yaml, leveraging streaming data loading and automatic VRAM management from StreamingSetupMixin
  • Deploy and share using soup serve for local APIs or soup push to publish reproducible checkpoints to Hugging Face Hub with auto-generated model cards

Frequently Asked Questions

What data format does Soup CLI require for SFT?

Soup CLI accepts JSONL, CSV, or Hugging Face Hub datasets where each row contains a prompt field and a response field. The soup data validate command checks format compliance before training begins, and the SFTTrainerWrapper streams this data during the training loop without loading the entire dataset into memory.

How do I enable LoRA for parameter-efficient fine-tuning?

Add a lora: block to your soup.yaml configuration specifying the rank (r), alpha (alpha), and target modules. When present, the SFTTrainerWrapper automatically calls get_peft_model to wrap the base model, enabling adapter training while keeping base parameters frozen.

Can I run SFT on Apple Silicon with Soup CLI?

Yes. Set backend: mlx in your configuration file. The SFTTrainerWrapper detects the backend at runtime and adapts the data loader and training operations to use the MLX framework instead of standard Transformers, enabling native execution on Apple Silicon devices.

Where are training checkpoints and logs stored?

By default, checkpoints are written to the output/ directory (e.g., output/checkpoint-500/), while logs and TensorBoard events are stored in output/logs/. Each run also generates a repro-receipt.json in the output directory containing full configuration snapshots for reproducibility.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →