How to Configure `rank_pattern` and `alpha_pattern` for MoE Models in Soup
rank_pattern and alpha_pattern are configuration fields in LoraConfig that let you override the default LoRA rank and alpha per target-module name pattern, enabling efficient fine-tuning of Mixture-of-Experts (MoE) models by applying lower ranks to heavy expert layers while keeping higher ranks for attention layers.
The Soup framework provides granular control over LoRA hyperparameters through pattern-based overrides. For MoE architectures—such as Mixtral, Qwen-MoE, and DeepSeek-V3—this capability is essential because expert feed-forward networks contain far more parameters than attention layers, making uniform LoRA ranks inefficient.
What rank_pattern and alpha_pattern Control
LoRA (Low-Rank Adaptation) injects trainable low-rank matrices into frozen base weights. By default, every target module uses the same rank (r) and scaling factor (alpha). The rank_pattern and alpha_pattern dictionaries break this uniformity:
- rank_pattern: Maps module name patterns to custom ranks (e.g.,
{"experts.*.w1": 8}) - alpha_pattern: Maps module name patterns to custom scaling factors (e.g.,
{"experts.*.w1": 4})
The final LoRA update follows ΔW = (α / r) · A · B, so both parameters jointly control update magnitude.
In src/soup_cli/config/schema.py, these fields are defined as part of the LoraConfig dataclass with strict validation rules. The runtime wiring that forwards these dictionaries to PEFT occurs in src/soup_cli/utils/peft_builder.py at lines 54-55, where they populate init_kwargs["rank_pattern"] and init_kwargs["alpha_pattern"].
Validation Rules Enforced by LoraConfig
The schema in src/soup_cli/config/schema.py (lines 195-226 and 230-240) enforces these constraints:
| Rule | Constraint | Violation Example |
|---|---|---|
| Type | Must be dict[str, int] |
Passing a list raises rank_pattern/alpha_pattern must be a dict[str, int] |
| Key count | Maximum 256 keys (_MAX_LORA_RANK_PATTERN_KEYS) |
Exceeding triggers rank_pattern/alpha_pattern caps at 256 keys |
| Key format | Non-empty strings, no null bytes (\x00) |
Empty key raises rank_pattern/alpha_pattern keys must be non-empty strings |
| Value range | Positive integers ≤ 1024 (_MAX_LORA_RANK_PATTERN_VALUE) |
Zero or 1025 raises rank_pattern/alpha_pattern values must be in (0, 1024] |
| VeRA incompatibility | Cannot coexist with use_vera=True |
Conflict raises rank_pattern is incompatible with use_vera=True |
VeRA (Vector-based Random Matrix Adaptation) uses a single shared rank across all layers, making per-module patterns logically incompatible. Ensure use_vera: false (or omit the field) when configuring patterns.
YAML Configuration for MoE Models
Add rank_pattern and alpha_pattern under training.lora in your soup.yaml:
training:
lora:
r: 64 # default global rank for attention layers
alpha: 16 # default global alpha
target_modules:
- "q_proj"
- "k_proj"
- "v_proj"
- "o_proj"
- "experts.*.w1"
- "experts.*.w2"
- "experts.*.w3"
rank_pattern:
"experts.*.w1": 8 # lower rank for all expert W1 projections
"experts.*.w2": 8 # lower rank for all expert W2 projections
"experts.*.w3": 8 # lower rank for all expert W3 projections
alpha_pattern:
"experts.*.w1": 4 # reduced scaling for expert W1
"experts.*.w2": 4 # reduced scaling for expert W2
"experts.*.w3": 4 # reduced scaling for expert W3
Patterns use Python's re.search for matching against full module names from model.named_modules(). The first matching pattern wins when multiple patterns could apply.
Python API for Advanced Use Cases
Instantiate LoraConfig directly for programmatic control:
from soup_cli.config.schema import LoraConfig
lora_cfg = LoraConfig(
r=64,
alpha=16,
target_modules=["q_proj", "k_proj", "v_proj", "experts.*.w1", "experts.*.w2"],
rank_pattern={
"experts.*.w1": 8,
"experts.*.w2": 8,
},
alpha_pattern={
"experts.*.w1": 4,
"experts.*.w2": 4,
},
use_dora=False, # compatible with patterns; applies to base LoRA rank
use_vera=False, # REQUIRED: must be False when using patterns
)
Pass this configuration to soup_cli.utils.peft_builder.build_peft_adapter() or your trainer initialization. The builder extracts patterns and injects them into PEFT's initialization kwargs at src/soup_cli/utils/peft_builder.py.
Common MoE Pattern Reference
| Model Family | Recommended Pattern | Target Module |
|---|---|---|
| Mixtral 8x7B / 8x22B | "experts.*.w1" , "experts.*.w2" , "experts.*.w3" |
SWiGLU FFN experts (gate, up, down projections) |
| Qwen1.5-MoE / Qwen2-MoE | "mlp.experts.*.gate_proj" , "mlp.experts.*.up_proj" , "mlp.experts.*.down_proj" |
Full expert MLP submodules |
| DeepSeek-V3 | "experts.*.ffn" or "mlp.experts.*" |
Consolidated expert FFN blocks |
| DBRX | "ffn.experts.*.w1" , "ffn.experts.*.w2" |
Granular expert weights |
| OLMoE | "transformer.blocks.*.mlp.experts.*" |
Block-scoped expert patterns |
To verify patterns match your specific checkpoint, inspect module names:
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("your-moe-model")
for name, _ in list(model.named_modules())[:30]:
print(name)
Look for strings containing experts, mlp, or ffn to identify the correct pattern syntax.
Troubleshooting Common Issues
Pattern not applied: Confirm use_vera=False and check that patterns actually match module names. Misplaced wildcards (experts.*w1 vs experts.*.w1) are frequent errors.
Validation errors on key count: Consolidate patterns using broader wildcards. Instead of listing 64 individual experts, use experts.*.w1.
Unexpected memory usage: Lower ranks reduce parameter count but not necessarily activation memory. Combine rank_pattern with gradient checkpointing for MoE models.
DoRA interaction: use_dora=True is compatible with rank_pattern. The rank override applies to the base LoRA matrices before DoRA's weight-decomposed transformation.
Summary
rank_patternandalpha_patterninLoraConfigenable per-module LoRA hyperparameters via regex-like string matching.- MoE models benefit significantly from reduced expert ranks (8-16) while keeping attention ranks high (64-256).
- Validation limits: 256 keys max, values 1-1024, incompatible with
use_vera=True. - Configuration path:
src/soup_cli/config/schema.pydefines validation;src/soup_cli/utils/peft_builder.pyforwards to PEFT. - Best practice: Verify module names with
named_modules(), start with conservative rank reductions, and monitor training dynamics through the alpha scaling factor.
Frequently Asked Questions
What happens if a module matches multiple patterns in rank_pattern?
PEFT applies the first matching pattern in dictionary iteration order. Python 3.7+ preserves insertion order, so list your most specific patterns first. For predictable behavior, ensure pattern sets are mutually exclusive or ordered by specificity.
Can I use rank_pattern with DoRA (Weight-Decomposed Low-Rank Adaptation)?
Yes. According to the Soup source in src/soup_cli/utils/peft_builder.py, DoRA is compatible with rank_pattern and alpha_pattern. The rank override applies to the base LoRA matrices (A and B), while DoRA's magnitude vector (m) operates on the decomposed result. Set use_dora=True alongside your patterns.
Why does setting rank_pattern trigger a validation error about VeRA?
VeRA uses a single shared random matrix across all layers, eliminating the concept of per-module ranks. The validator in src/soup_cli/config/schema.py explicitly rejects this combination with rank_pattern is incompatible with use_vera=True. Choose either VeRA (simple, memory-efficient) or patterned LoRA (flexible, MoE-optimized)—not both.
How do I find the exact module names for my MoE model?
Run this diagnostic to dump all module names:
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("model-name", device_map="auto")
for name, module in model.named_modules():
if "expert" in name.lower() or "mlp" in name.lower():
print(f"{name}: {type(module).__name__}")
Match the printed names against your intended patterns. Test patterns with re.search(pattern, name) to confirm matches before adding to your configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →