# How to Configure LLM Fine-Tuning Using YAML Files in Soup: A Complete Guide

> Learn to configure LLM fine-tuning with Soup's YAML file. This guide explains parsing YAML into Pydantic models for robust, validated workflows. Get started now!

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Soup drives its entire fine-tuning workflow from a single [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) file that is parsed into Pydantic models for strict up-front validation.**

The open-source Soup framework (available at `MakazhanAlpamys/Soup`) eliminates configuration boilerplate by centralizing every training hyperparameter, data source, and LoRA setting in one declarative YAML file. When you run commands like `soup train -c soup.yaml`, the CLI validates your configuration against a strict schema before executing any GPU-intensive code, catching typos and invalid combinations immediately.

## The soup.yaml Schema Architecture

Soup's configuration system treats [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) as the single source of truth. Every YAML section maps to a Pydantic `BaseModel` subclass, creating a type-safe contract between your configuration file and the training runtime.

### Core Configuration Models

The top-level `SoupConfig` object composes several specialized models:

- **DataConfig** (`L124–L158`): Defines training data sources, formats (Alpaca, ShareGPT), streaming options, and validation splits
- **TrainingConfig** (`L224–L279`): Controls epochs, learning rates, batch sizes, optimizer selection, and quantization settings
- **LoraConfig** (`L53–L110`): Manages adapter rank, alpha values, dropout, and pattern-based target module selection

When the CLI executes `soup train`, it calls `load_config_from_string` in [`src/soup_cli/config/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/loader.py) (`L86`), which parses the YAML string and instantiates these models. This triggers automatic field validation, ensuring that parameters like `batch_size: auto` or `quantization: 4bit` conform to expected types and ranges.

## Validation and Safety Guards

Soup implements multiple safety layers to prevent misconfigurations from reaching the training loop:

**Path Containment** (`L21–L28`): The validator `is_under_cwd` ensures all file paths referenced in the YAML are relative and contained within the current working directory, blocking absolute paths that could access sensitive system locations.

**Mutual Exclusivity Checks** (`L57–L72`): The schema enforces that incompatible PEFT methods cannot be combined, preventing runtime conflicts between different adapter types.

**ReDoS Protection** (`L39–L44`): Regex patterns used in `unfrozen_parameters` and `rank_pattern` are screened against ReDoS (Regular Expression Denial of Service) attacks before compilation.

## Minimal Configuration Example

Create a [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) file with only the sections you need—omitted fields automatically use safe defaults:

```yaml

# soup.yaml

data:
  train: data/my_dataset.jsonl
  format: alpaca
  val_split: 0.1

training:
  epochs: 5
  batch_size: auto
  lr: 2e-5
  lora:
    r: 32
    alpha: 16
    dropout: 0.05

quantization: 4bit
optimizer: adamw_torch

```

In this configuration:
- `data.train` accepts either local paths or Hugging Face dataset names (validated by `DataConfig`)
- `batch_size: auto` enables Soup's GPU memory probing to find the optimal batch size
- `quantization: 4bit` activates QLoRA when combined with LoRA settings

## Advanced Configuration Patterns

For production fine-tuning, you can enable streaming datasets, pattern-based LoRA ranks, and per-module learning rates:

```yaml
data:
  train:
    - data/part1.jsonl
    - data/part2.jsonl
  format: sharegpt
  streaming: true
  buffer_size: 50000
  interleave: concat
  image_dir: images/

training:
  epochs: 3
  batch_size: 8
  lora:
    r: 0                # full fine-tuning (no adapter)

    target_modules: ["q_proj", "v_proj"]
    rank_pattern:
      "expert.*.w1": 16
      "expert.*.w2": 24
  quantization: none
  optimizer: adamw_torch
  lr_groups:
    - pattern: ".*lm_head.*"
      lr: 5e-5
    - pattern: ".*"
      lr: 1e-5

```

Key advanced features demonstrated:
- **Streaming** (`L124–L158`): Loads data on-the-fly with configurable `buffer_size`, essential for datasets larger than RAM
- **Full Fine-Tuning**: Setting `lora.r: 0` disables adapters entirely according to `LoraConfig` validation rules (`L54–L60`)
- **Pattern-Based Ranks** (`L12–L20`): The `rank_pattern` dictionary applies different LoRA ranks to specific layer patterns (e.g., MoE experts)
- **Learning Rate Groups** (`L75–L81`): `lr_groups` assigns specific learning rates to parameter groups matching regex patterns

## Programmatic Validation

Before launching expensive training runs, validate your YAML programmatically using the same loader the CLI employs:

```python
from soup_cli.config.loader import load_config_from_string

yaml_text = open("soup.yaml").read()
config = load_config_from_string(yaml_text)   # raises ValueError if invalid

# Access validated fields safely

print(config.training.batch_size)
print(config.data.format)

```

This approach raises `ValueError` immediately if your YAML contains schema violations, path traversals, or mutually exclusive settings, allowing you to fix configuration errors before allocating GPU resources.

## Summary

- Soup uses a single [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml) file as the source of truth for all fine-tuning parameters, defined by Pydantic models in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py)
- The `load_config_from_string` function in [`src/soup_cli/config/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/loader.py) validates configurations upfront using strict type checking and custom validators
- Safety guards prevent path traversal attacks (`is_under_cwd`), ReDoS vulnerabilities, and incompatible PEFT method combinations
- You can configure everything from basic QLoRA training to advanced pattern-based full fine-tuning using the same YAML structure
- Programmatic validation allows you to test configurations before executing `soup train` or `soup ship` for distributed runs

## Frequently Asked Questions

### What happens if I make a typo in my soup.yaml file?

Soup's Pydantic schema validation will catch the error immediately when you run `soup train` or call `load_config_from_string`. The validation occurs before any model weights are loaded or GPU memory is allocated, providing clear error messages about which field contains the invalid value according to the definitions in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py).

### Can I use absolute paths for my training data?

No, the `is_under_cwd` validator in the configuration loader blocks absolute paths and paths that traverse outside the current working directory. All data paths must be relative to your project root to prevent accidental access to system files. This security check is implemented in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) (`L21–L28`).

### How do I switch between LoRA and full fine-tuning?

Set [`training.lora.r`](https://github.com/MakazhanAlpamys/Soup/blob/main/training.lora.r) to a positive integer (typically 8-64) for LoRA/QLoRA training, or set it to `0` to disable adapters entirely and perform full fine-tuning. When `r: 0`, Soup ignores LoRA-specific parameters and trains all unfrozen parameters directly, as validated in `LoraConfig` (`L54–L60`). Ensure you adjust `quantization` accordingly—`4bit` requires LoRA, while full fine-tuning typically uses `quantization: none`.

### Does Soup support multiple datasets in one configuration?

Yes, the `data.train` field accepts both single strings and lists of paths. When providing multiple files, you can specify `interleave: concat` or other strategies to combine them. Each path is validated for existence and containment within the working directory before training begins.