Soup CLI LoRA Configurations: Complete Guide to Fine-Tuning with Low-Rank Adaptation
The Soup CLI exposes comprehensive LoRA configuration options through the LoraConfig Pydantic model in src/soup_cli/config/schema.py, supporting standard LoRA alongside specialized variants like DoRA, VeRA, and OLoRA with multiple initialization strategies.
The Soup CLI (MakazhanAlpamys/Soup) provides a unified interface for training Large Language Models using Low-Rank Adaptation (LoRA). At its core, the LoraConfig class serves as the single source of truth for all adapter configurations, validated through Pydantic model validators that enforce mutual exclusivity and compatibility constraints before training begins.
Core LoRA Parameters
The foundation of any LoRA training run relies on four primary parameters defined in src/soup_cli/config/schema.py.
Rank and Scaling Factor
The r parameter controls the low-rank dimension of the adapter. Setting r = 0 disables LoRA entirely, enabling full fine-tuning of the base model without adapter layers. The alpha parameter scales the LoRA weight updates according to the formula α = r * alpha_factor, determining the magnitude of parameter updates during training.
Dropout Regularization
The dropout parameter specifies the probability of dropping LoRA updates during training, defaulting to 0.05. When using target_parameters to directly modify specific nn.Parameter tensors, you must set dropout = 0 to ensure deterministic updates to expert weights or other targeted parameters.
Target Module Selection
The target_modules parameter accepts either a list of module names or the string "auto", which delegates module inference to PEFT. For advanced use cases, target_parameters accepts a direct list of 2D or 3D parameter tensors (such as expert weights in Mixture-of-Experts models), though this requires dropout = 0 and is only compatible with plain LoRA.
Advanced LoRA Variants
The Soup CLI supports four distinct LoRA extensions beyond the standard implementation, each controlled via boolean flags that are mutually exclusive.
DoRA (Weight-Decomposed LoRA)
Setting use_dora = true enables Weight-Decomposed Low-Rank Adaptation, which learns an additional decomposition of the pre-trained weights for improved training stability. DoRA is incompatible with use_vera and use_olora, and cannot be combined with init_strategy values of "pissa" or "loftq".
VeRA (Vector-based Random Matrix Adaptation)
The use_vera = true flag activates VeRA, which uses a shared random initialization across all adapted layers while learning scaling vectors. This approach shares a single rank across all target modules and is incompatible with rank_pattern, alpha_pattern, and certain initialization strategies.
OLoRA (Orthogonal LoRA)
Enabling use_olora = true triggers orthogonal initialization via QR decomposition, which improves convergence properties. The validator at src/soup_cli/config/schema.py automatically sets init_strategy = 'olora' when this flag is enabled, maintaining backward compatibility with configurations from versions prior to v0.33.0. OLoRA is mutually exclusive with both DoRA and VeRA.
r-sLoRA (Rank-Stabilized LoRA)
The use_rslora = true parameter activates rank-stabilized scaling, a variant designed to improve training stability when using high ranks. Unlike DoRA, VeRA, and OLoRA, r-sLoRA is compatible with plain LoRA and does not impose mutual exclusivity constraints.
Initialization Strategies
The init_strategy parameter determines how LoRA weights are initialized, with four supported options validated by the schema.
"random": Standard Kaiming initialization (default)."pissa": Principal Singular Spectrum Adaptation, using SVD-based initialization for faster early convergence. Incompatible withuse_doraanduse_vera."olora": Orthogonal initialization via QR decomposition (equivalent to enablinguse_olora)."loftq": LoftQ quantization-aware initialization, which jointly creates low-rank A/B matrices and a low-bit base weight. Requires specifyingloftq_iter(1-10 iterations) andloftq_bits(2, 4, or 8 bits). Incompatible withuse_doraanduse_vera.
Fine-Grained Pattern Control
For models with heterogeneous sub-structures such as Mixture-of-Experts (MoE), the CLI supports per-module configuration through regular expression patterns.
Rank and Alpha Patterns
The rank_pattern dictionary maps regex patterns to specific ranks (e.g., {"q_proj": 8, "experts.*.w1": 4}), allowing different adapter capacities for different model components. Similarly, alpha_pattern maps patterns to specific alpha values. These patterns are validated to ensure they do not exceed 256 keys with values capped at 1024 to maintain predictable memory usage. Pattern-based configuration is incompatible with VeRA.
Validation and Constraints
The LoraConfig model employs Pydantic validators operating in both mode="before" and mode="after" to enforce configuration integrity. These validators ensure that:
use_dora,use_vera, anduse_oloraare never enabled simultaneously.target_parametersusage requiresdropout = 0.init_strategyvalues of"pissa"or"loftq"are not combined with DoRA or VeRA.loftq_iterandloftq_bitsare only specified when using LoftQ initialization.
Validated configurations are passed to src/soup_cli/utils/peft_builder.py, where the build_lora_config_kwargs function translates Soup CLI parameters into PEFT-compatible arguments for peft.get_peft_model.
Command-Line Examples
Configure LoRA directly via CLI flags using the --lora prefix:
# Standard LoRA with rank 64
soup train --lora.r 64 --lora.alpha 128
# Full fine-tuning (disable LoRA)
soup train --lora.r 0
# DoRA with specific rank
soup train --lora.use_dora true --lora.r 32
# VeRA adaptation
soup train --lora.use_vera true --lora.r 16
# OLoRA orthogonal initialization
soup train --lora.use_olora true --lora.r 8
# PiSSA initialization
soup train --lora.init_strategy pissa --lora.r 8
# LoftQ with 4-bit quantization
soup train --lora.init_strategy loftq --lora.loftq_bits 4 --lora.loftq_iter 3
# Heterogeneous MoE configuration
soup train \
--lora.rank_pattern "q_proj=8,experts.*.w1=4" \
--lora.alpha_pattern "q_proj=16,experts.*.w1=8"
# Direct expert parameter targeting
soup train --lora.target_parameters "auto" --lora.dropout 0
Summary
- The
LoraConfigclass insrc/soup_cli/config/schema.pydefines all LoRA parameters as a single source of truth. - Setting
r = 0disables LoRA entirely, enabling full fine-tuning. - DoRA, VeRA, and OLoRA are mutually exclusive variants that extend standard LoRA capabilities.
- Initialization strategies include random, PiSSA, OLoRA, and LoftQ quantization-aware initialization.
- Pattern-based configuration via
rank_patternandalpha_patternenables heterogeneous adapter capacities for complex model architectures. - Pydantic validators enforce compatibility constraints before training begins, preventing silent configuration errors.
Frequently Asked Questions
What happens when I set r = 0 in Soup CLI LoRA configuration?
Setting r = 0 disables the LoRA adapter entirely, causing the CLI to perform full fine-tuning of the base model without injecting any low-rank adaptation layers. This effectively bypasses the PEFT adapter setup and trains all original model parameters directly.
Can I use DoRA and VeRA together in the same Soup CLI training run?
No, DoRA (use_dora), VeRA (use_vera), and OLoRA (use_olora) are mutually exclusive options enforced by the LoraConfig validators in src/soup_cli/config/schema.py. Attempting to enable more than one simultaneously will trigger a validation error before training begins.
How do I configure different LoRA ranks for different layers in my model?
Use the rank_pattern parameter with regular expression keys mapping to integer ranks (e.g., {"q_proj": 8, "v_proj": 4}). This is particularly useful for Mixture-of-Experts models where expert layers may require different adapter capacities than attention layers. Note that pattern-based configuration is incompatible with VeRA.
What is the difference between target_modules and target_parameters in Soup CLI?
target_modules accepts string names of modules (or "auto") to which LoRA adapters are applied via PEFT's standard mechanism. target_parameters accepts direct references to 2D or 3D nn.Parameter tensors (such as expert weights), requiring dropout = 0 and only working with plain LoRA (not DoRA, VeRA, or OLoRA).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →