LoRA vs DoRA vs VeRA vs OLoRA: How to Choose the Right PEFT Method
LoRA, DoRA, VeRA, and OLoRA are all parameter-efficient fine-tuning (PEFT) methods that add lightweight trainable adapters to frozen base models, but they differ in how adapter weights are structured, initialized, and scaled.
All four techniques in the MakazhanAlpamys/Soup repository let you fine-tune large language models without updating billions of parameters. The choice between them depends on your memory constraints, training rank, and whether you need runtime control over adapter strength.
What Is LoRA and How Does It Work?
LoRA (Low-Rank Adaptation) is the baseline PEFT method. Instead of directly updating weight matrices, LoRA adds two small matrices A and B such that ΔW = B·A. The original weights remain frozen, and only these low-rank matrices are trained.
The core configuration lives in LoraConfig.r, which controls the rank of the decomposition:
# src/soup_cli/config/schema.py – baseline LoRA rank definition
class LoraConfig(BaseModel):
r: int = Field(
default=8,
ge=1,
description="LoRA attention dimension (the 'rank'). Higher values = more "
"expressive adapters but more parameters. Common values: 8, 16, 32, 64, 128.",
)
LoRA works for any model architecture and is the safest default when you're unsure which method to use.
DoRA: Weight-Decomposed LoRA for Dynamic Scaling
DoRA (Weight-Decomposed Low-Rank Adaptation) decomposes the adapter into separate direction and magnitude components, similar to singular-value decomposition.
Enable DoRA with the use_dora flag:
# src/soup_cli/config/schema.py – DoRA toggle
use_dora: bool = Field(
default=False,
description="Enable DoRA (Weight-Decomposed LoRA) from 'LyCORIS'.",
)
DoRA shines when you need to adjust adapter strength without reloading weights. The direction and scale components can be modified independently at inference time, making it ideal for multi-task scenarios or prompt-controlled adaptation.
VeRA: Extreme Memory Efficiency with Random Vectors
VeRA (Vector-based Random Matrix Adaptation) replaces LoRA's two-matrix structure with a single shared random vector per rank. This dramatically reduces memory overhead.
Enable VeRA with:
# src/soup_cli/config/schema.py – VeRA toggle
use_vera: bool = Field(
default=False,
description="Enable VeRA (Vector-based Random Matrix Adaptation). Note: VeRA "
"uses different rank semantics and ignores 'init_lora_weights'. "
"See: https://arxiv.org/abs/2310.12321",
)
Choose VeRA when working with very large models (70B+ parameters) or tight GPU memory constraints. The memory savings can reach 10× compared to standard LoRA at equivalent expressiveness.
OLoRA: Orthogonal Initialization for High-Rank Training
OLoRA (Orthogonal LoRA) initializes the adapter matrices using QR decomposition, producing an orthonormal basis for the "A" matrix.
Enable OLoRA with:
# src/soup_cli/config/schema.py – OLoRA toggle
use_olora: bool = Field(
default=False,
description="Enable OLoRA (Orthogonal LoRA) which initializes LoRA matrices "
"using QR decomposition for better high-rank convergence. "
"See: https://arxiv.org/abs/2406.01775v3",
)
OLoRA is specifically designed for high-rank adapters (typically r > 64). The orthogonal initialization provides more stable gradients and faster convergence when you need larger adapter capacity.
Mutual Exclusivity: Only One Method at a Time
The three advanced methods are mutually exclusive. The schema enforces this through the _validate_peft_exclusivity validator:
# src/soup_cli/config/schema.py – exclusivity validator
@model_validator(mode="after")
def _validate_peft_exclusivity(self):
enabled = []
if self.use_dora:
enabled.append("use_dora")
if self.use_vera:
enabled.append("use_vera")
if self.use_olora:
enabled.append("use_olora")
if len(enabled) > 1:
raise ValueError(
f"PEFT methods are mutually exclusive, got {len(enabled)} enabled: "
f"{', '.join(enabled)}. Pick at most one of use_dora, use_vera, use_olora."
)
return self
Attempting to enable multiple flags simultaneously raises a clear ValueError with instructions on which fields conflict.
Building PEFT Configs: The Runtime Translation
The build_peft_config function in the PEFT builder translates your schema selection into the appropriate peft library class:
# src/soup_cli/utils/peft_builder.py – method selection logic
def build_peft_config(lora_cfg: LoraConfig) -> dict:
if lora_cfg.use_vera:
return {
"peft_cls": "VeraConfig",
"init_kwargs": {"r": lora_cfg.r, ...}
}
# LoRA variants (including DoRA and OLoRA) use LoraConfig with flags
init_kwargs = {
"r": lora_cfg.r,
"lora_alpha": lora_cfg.alpha,
"use_dora": lora_cfg.use_dora,
}
if lora_cfg.use_olora:
init_kwargs["init_lora_weights"] = "olora"
return {
"peft_cls": "LoraConfig",
"init_kwargs": init_kwargs,
}
VeRA requires VeraConfig, while DoRA and OLoRA modify LoraConfig through specific parameters.
Practical Configuration Examples
Plain LoRA (Default Choice)
lora_cfg = LoraConfig(
r=64,
alpha=16,
dropout=0.05,
target_modules="auto",
)
# Results in: {"peft_cls": "LoraConfig", "init_kwargs": {...}}
DoRA with Dynamic Scaling
lora_cfg = LoraConfig(
r=32,
alpha=16,
use_dora=True, # Enable weight decomposition
)
# Results in: {"peft_cls": "LoraConfig", "init_kwargs": {"use_dora": True, ...}}
VeRA for Memory-Constrained Training
lora_cfg = LoraConfig(
r=64,
use_vera=True, # Switches to VeraConfig
# Note: init_lora_weights is ignored for VeRA
)
# Results in: {"peft_cls": "VeraConfig", "init_kwargs": {...}}
OLoRA for High-Rank Stability
lora_cfg = LoraConfig(
r=128, # High rank benefits most
alpha=32,
use_olora=True, # QR-based orthogonal initialization
)
# Results in: {"peft_cls": "LoraConfig", "init_kwargs": {"init_lora_weights": "olora", ...}}
All configurations can also be expressed in YAML:
lora:
r: 64
alpha: 16
dropout: 0.05
target_modules: auto
# Enable exactly ONE of the following:
# use_dora: true # For dynamic scaling
# use_vera: true # For memory efficiency
# use_olora: true # For high-rank stability
Decision Guide: Which Method to Choose?
| Your Situation | Recommended Method | Rationale |
|---|---|---|
| Unsure or new to PEFT | LoRA | Most tested, broadest compatibility, safest default |
| Need runtime adapter strength control | DoRA | Direction/scale decomposition enables on-the-fly adjustments |
| Training 70B+ models, limited VRAM | VeRA | ~10× memory reduction through shared random vectors |
| Using ranks above 64 | OLoRA | Orthogonal initialization stabilizes high-rank training |
| Multiple concerns apply | LoRA | Start here, then experiment; exclusivity validator prevents errors |
Summary
-
LoRA is the foundational PEFT method using low-rank matrix decomposition—start here for any fine-tuning task.
-
DoRA adds direction/scale decomposition for dynamic adapter control, configured via
use_dorain [schema.py](/src/soup_cli/config/schema.py#L68-L71). -
VeRA minimizes memory with shared random vectors, switching to
VeraConfigwhenuse_vera=Trueper [peft_builder.py](/src/soup_cli/utils/peft_builder.py#L30-L41). -
OLoRA stabilizes high-rank training via QR orthogonal initialization, activated with
use_olora. -
The three advanced methods are mutually exclusive—the
_validate_peft_exclusivityvalidator raises errors for invalid combinations.
Frequently Asked Questions
Can I combine DoRA with OLoRA?
No. The validator in [schema.py](/src/soup_cli/config/schema.py#L38-L53) explicitly forbids enabling multiple flags. DoRA and OLoRA both modify LoraConfig behavior, and their interaction is undefined. Choose the method that better matches your needs: DoRA for runtime scaling, or OLoRA for high-rank initialization.
Why does VeRA use a different PEFT class?
VeRA fundamentally changes the adapter structure from matrix pairs to shared random vectors. The peft library implements this as VeraConfig rather than a LoraConfig flag. The Soup builder handles this translation automatically when use_vera=True, returning {"peft_cls": "VeraConfig"} instead of the default LoraConfig wrapper.
What happens if I set rank patterns with VeRA?
VeRA ignores rank_pattern and related per-module rank specifications because its vector-sharing mechanism uses different semantics than LoRA's module-specific matrices. The tests/test_rank_pattern.py suite verifies that such configurations are rejected or warned when VeRA is active.
Is OLoRA worth it for small ranks (r ≤ 32)?
Unlikely. The OLoRA paper and the use_olora description target high-rank scenarios where standard initialization struggles. For typical ranks of 8–32, standard LoRA or DoRA provide better parameter efficiency without the orthogonalization overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →