# How to Perform Supervised Fine-Tuning (SFT) with Soup CLI: A Complete Guide

> Master Supervised Fine-Tuning SFT with the Soup CLI. Follow our guide to easily train models using streaming data, LoRA, and checkpointing.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Supervised Fine-Tuning (SFT) with the Soup CLI is performed by running `soup train --config soup.yaml` after initializing a configuration with `soup init --task sft`, which invokes the `SFTTrainerWrapper` class in [`src/soup_cli/trainer/sft.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/sft.py) to handle streaming data loading, LoRA wrapping, and checkpoint management.**

The MakazhanAlpamys/Soup repository provides an end-to-end command-line interface for adapting pretrained language models to downstream tasks using labeled datasets. This guide walks you through the complete workflow for Supervised Fine-Tuning (SFT) with Soup CLI, from configuration to deployment.


## What is Supervised Fine-Tuning in Soup?

**Supervised Fine-Tuning (SFT)** adapts a pretrained model to specific tasks using labeled prompt-response pairs. In the Soup ecosystem, the `SFTTrainerWrapper` class ([[`src/soup_cli/trainer/sft.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/sft.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/sft.py)) orchestrates this process by integrating with the Pydantic configuration schema defined in [[`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py).

The trainer inherits from `StreamingSetupMixin`, which enables dynamic batch sizing, automatic VRAM probing, and memory-efficient streaming of large datasets without full materialization. When you execute `soup train`, the CLI resolves the `task: sft` configuration to instantiate this wrapper and begins the training loop.


## Prerequisites and Configuration Setup

### Initializing the SFT Configuration

Begin by scaffolding a configuration file that defines your SFT task. The `soup init` command generates a [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) file validated against the single source-of-truth schema:

```bash
soup init --template chat \
    --model meta-llama/Llama-3.1-8B \
    --task sft \
    --output soup.yaml

```

This creates a configuration file containing a `task: sft` block and placeholders for data paths. All configuration fields—including model identifiers, training hyperparameters, and LoRA settings—are strictly validated by the Pydantic schema before training begins.


### Preparing Your Supervised Dataset

Soup CLI expects supervised datasets in JSONL, CSV, or Hugging Face Hub formats where each record contains a `prompt` and a `response` field:

```json
{ "prompt": "Translate English to French: Hello", "response": "Bonjour" }

```

Validate your local dataset before training:

```bash
soup data validate my_data.jsonl

```

Once validated, reference the dataset in [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml):

```yaml
data:
  train: my_data.jsonl
  format: chatml

```

The `SFTTrainerWrapper` streams this data during training, processing batches through a dataloader that yields `(input_ids, labels)` pairs with loss computed exclusively on response tokens when `training.train_on_responses_only: true` is set.


## Configuring Model and Training Parameters

### Selecting Base Models and LoRA Adapters

Specify your base model using the `model` field in the configuration. For parameter-efficient fine-tuning, add a `lora` block:

```yaml
model: meta-llama/Llama-3.1-8B
lora:
  r: 64
  alpha: 16
  target_modules:
    - q_proj
    - k_proj
    - v_proj
    - o_proj

```

When LoRA configuration is present, the trainer calls `get_peft_model` inside the SFT wrapper to freeze base model weights and train only the adapter parameters. The trainer supports both the `transformers` backend (default) and the `mlx` backend for Apple Silicon, detected at runtime based on your environment.


### Hyperparameter Configuration

Edit training parameters under the `training` section in [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml):

```yaml
training:
  lr: 2e-4
  epochs: 3
  batch_size: 4
  train_on_responses_only: true

```

These fields map directly to Hugging Face `TrainingArguments` after schema validation. The `SFTTrainerWrapper` uses `CrossEntropyLoss` over response tokens during the warmup and training phases, with automatic gradient accumulation and mixed-precision support handled by the underlying training framework.


## Running the SFT Training Loop

Execute the training run with a single command:

```bash
soup train --config soup.yaml \
    --gpus auto \
    --tensorboard

```

During execution, the CLI performs the following operations according to the MakazhanAlpamys/Soup source code:

1. **Configuration Parsing**: Validates [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) against the Pydantic schema in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py)
2. **Trainer Instantiation**: Creates `SFTTrainerWrapper` with `StreamingSetupMixin` capabilities
3. **Model Loading**: Loads the base model via `transformers` (or `mlx` on Apple Silicon) and wraps with PEFT if LoRA is configured
4. **Streaming Data**: Initializes a streaming dataloader that probes available VRAM and adjusts batch sizing dynamically
5. **Training Loop**: Performs single-epoch warmup followed by full training, computing loss only on response tokens
6. **Checkpointing**: Writes checkpoints to `output/checkpoint-{step}/` directories at configured intervals

The trainer emits Rich-formatted logs to the terminal and optionally streams metrics to TensorBoard for real-time visualization.


## Monitoring, Evaluation, and Publishing

### Real-time Monitoring with TensorBoard

Enable TensorBoard logging using the `--tensorboard` flag during training. The `SFTTrainerWrapper` emits loss curves, learning rate schedules, and throughput metrics to the log directory, viewable via:

```bash
tensorboard --logdir output/logs

```

Every SFT run also generates a [`repro-receipt.json`](https://github.com/MakazhanAlpamys/Soup/blob/main/repro-receipt.json) file containing complete configuration snapshots, dependency versions, and training cost estimates produced by `soup cost` for full reproducibility.


### Inference and Model Serving

After training completes, test the fine-tuned model using local inference:

```bash
soup infer --model ./output \
    --input <(echo '{"prompt":"What is the capital of France?"}') \
    --output result.jsonl

```

Deploy the model as an OpenAI-compatible API server:

```bash
soup serve --model ./output --port 8000

```

The server exposes endpoints at `http://localhost:8000/v1/chat/completions`, allowing immediate integration with existing tooling.


### Publishing to Hugging Face Hub

Upload your fine-tuned checkpoint with automatic model card generation:

```bash
soup push --model ./output \
    --repo username/llama-3.1-sft \
    --card $(soup card output)

```

The `push` command reads checkpoint metadata and the reproducibility receipt to generate a comprehensive model card documenting training parameters, dataset provenance, and evaluation metrics.


## Advanced SFT Features

**Backend Agnostic Execution**: The `SFTTrainerWrapper` detects the compute backend at runtime, automatically adapting the data loader and optimization routines for NVIDIA GPUs (via `transformers`) or Apple Silicon (via `mlx`).

**Telemetry and Cost Estimation**: Each run writes training cost estimates and resource utilization statistics to `soup cost` output, enabling budget planning for large-scale fine-tuning operations.

**BitNet Support**: For specialized quantization requirements, the generic `SFTTrainerWrapper` architecture is subclassed in [`src/soup_cli/trainer/bitnet.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/bitnet.py) to support BitNet training while maintaining the same CLI interface.


## Summary

- **Initialize configurations** using `soup init --task sft` to generate validated [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) files conforming to the Pydantic schema in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py)
- **Prepare datasets** in JSONL format with `prompt` and `response` fields, validating them via `soup data validate` before training
- **Enable efficient tuning** by configuring LoRA adapters in the config file, which triggers `get_peft_model` wrapping inside `SFTTrainerWrapper`
- **Execute training** with `soup train --config soup.yaml`, leveraging streaming data loading and automatic VRAM management from `StreamingSetupMixin`
- **Deploy and share** using `soup serve` for local APIs or `soup push` to publish reproducible checkpoints to Hugging Face Hub with auto-generated model cards


## Frequently Asked Questions

### What data format does Soup CLI require for SFT?

Soup CLI accepts JSONL, CSV, or Hugging Face Hub datasets where each row contains a `prompt` field and a `response` field. The `soup data validate` command checks format compliance before training begins, and the `SFTTrainerWrapper` streams this data during the training loop without loading the entire dataset into memory.

### How do I enable LoRA for parameter-efficient fine-tuning?

Add a `lora:` block to your [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) configuration specifying the rank (`r`), alpha (`alpha`), and target modules. When present, the `SFTTrainerWrapper` automatically calls `get_peft_model` to wrap the base model, enabling adapter training while keeping base parameters frozen.

### Can I run SFT on Apple Silicon with Soup CLI?

Yes. Set `backend: mlx` in your configuration file. The `SFTTrainerWrapper` detects the backend at runtime and adapts the data loader and training operations to use the MLX framework instead of standard Transformers, enabling native execution on Apple Silicon devices.

### Where are training checkpoints and logs stored?

By default, checkpoints are written to the `output/` directory (e.g., `output/checkpoint-500/`), while logs and TensorBoard events are stored in `output/logs/`. Each run also generates a [`repro-receipt.json`](https://github.com/MakazhanAlpamys/Soup/blob/main/repro-receipt.json) in the output directory containing full configuration snapshots for reproducibility.