# How to Configure `rank_pattern` and `alpha_pattern` for MoE Models in Soup

> Learn to configure rank_pattern and alpha_pattern for MoE models in Soup. Optimize fine-tuning by adjusting LoRA ranks for expert and attention layers.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**`rank_pattern` and `alpha_pattern` are configuration fields in `LoraConfig` that let you override the default LoRA rank and alpha per target-module name pattern, enabling efficient fine-tuning of Mixture-of-Experts (MoE) models by applying lower ranks to heavy expert layers while keeping higher ranks for attention layers.**

The Soup framework provides granular control over LoRA hyperparameters through pattern-based overrides. For MoE architectures—such as Mixtral, Qwen-MoE, and DeepSeek-V3—this capability is essential because expert feed-forward networks contain far more parameters than attention layers, making uniform LoRA ranks inefficient.

## What `rank_pattern` and `alpha_pattern` Control

**LoRA** (Low-Rank Adaptation) injects trainable low-rank matrices into frozen base weights. By default, every target module uses the same rank (`r`) and scaling factor (`alpha`). The **rank_pattern** and **alpha_pattern** dictionaries break this uniformity:

- **rank_pattern**: Maps module name patterns to custom ranks (e.g., `{"experts.*.w1": 8}`)
- **alpha_pattern**: Maps module name patterns to custom scaling factors (e.g., `{"experts.*.w1": 4}`)

The final LoRA update follows `ΔW = (α / r) · A · B`, so both parameters jointly control update magnitude.

In [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py), these fields are defined as part of the `LoraConfig` dataclass with strict validation rules. The runtime wiring that forwards these dictionaries to PEFT occurs in [`src/soup_cli/utils/peft_builder.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/peft_builder.py) at lines 54-55, where they populate `init_kwargs["rank_pattern"]` and `init_kwargs["alpha_pattern"]`.

## Validation Rules Enforced by `LoraConfig`

The schema in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) (lines 195-226 and 230-240) enforces these constraints:

| Rule | Constraint | Violation Example |
|------|-----------|-------------------|
| **Type** | Must be `dict[str, int]` | Passing a list raises `rank_pattern/alpha_pattern must be a dict[str, int]` |
| **Key count** | Maximum 256 keys (`_MAX_LORA_RANK_PATTERN_KEYS`) | Exceeding triggers `rank_pattern/alpha_pattern caps at 256 keys` |
| **Key format** | Non-empty strings, no null bytes (`\x00`) | Empty key raises `rank_pattern/alpha_pattern keys must be non-empty strings` |
| **Value range** | Positive integers ≤ 1024 (`_MAX_LORA_RANK_PATTERN_VALUE`) | Zero or 1025 raises `rank_pattern/alpha_pattern values must be in (0, 1024]` |
| **VeRA incompatibility** | Cannot coexist with `use_vera=True` | Conflict raises `rank_pattern is incompatible with use_vera=True` |

**VeRA** (Vector-based Random Matrix Adaptation) uses a single shared rank across all layers, making per-module patterns logically incompatible. Ensure `use_vera: false` (or omit the field) when configuring patterns.

## YAML Configuration for MoE Models

Add `rank_pattern` and `alpha_pattern` under `training.lora` in your [`soup.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/soup.yaml):

```yaml
training:
  lora:
    r: 64                     # default global rank for attention layers

    alpha: 16                 # default global alpha

    target_modules:
      - "q_proj"
      - "k_proj"
      - "v_proj"
      - "o_proj"
      - "experts.*.w1"
      - "experts.*.w2"
      - "experts.*.w3"
    rank_pattern:
      "experts.*.w1": 8       # lower rank for all expert W1 projections

      "experts.*.w2": 8       # lower rank for all expert W2 projections

      "experts.*.w3": 8       # lower rank for all expert W3 projections

    alpha_pattern:
      "experts.*.w1": 4       # reduced scaling for expert W1

      "experts.*.w2": 4       # reduced scaling for expert W2

      "experts.*.w3": 4       # reduced scaling for expert W3

```

Patterns use Python's `re.search` for matching against full module names from `model.named_modules()`. The first matching pattern wins when multiple patterns could apply.

## Python API for Advanced Use Cases

Instantiate `LoraConfig` directly for programmatic control:

```python
from soup_cli.config.schema import LoraConfig

lora_cfg = LoraConfig(
    r=64,
    alpha=16,
    target_modules=["q_proj", "k_proj", "v_proj", "experts.*.w1", "experts.*.w2"],
    rank_pattern={
        "experts.*.w1": 8,
        "experts.*.w2": 8,
    },
    alpha_pattern={
        "experts.*.w1": 4,
        "experts.*.w2": 4,
    },
    use_dora=False,  # compatible with patterns; applies to base LoRA rank

    use_vera=False,  # REQUIRED: must be False when using patterns

)

```

Pass this configuration to `soup_cli.utils.peft_builder.build_peft_adapter()` or your trainer initialization. The builder extracts patterns and injects them into PEFT's initialization kwargs at [`src/soup_cli/utils/peft_builder.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/peft_builder.py).

## Common MoE Pattern Reference

| Model Family | Recommended Pattern | Target Module |
|-------------|---------------------|---------------|
| **Mixtral 8x7B / 8x22B** | `"experts.*.w1"` , `"experts.*.w2"` , `"experts.*.w3"` | SWiGLU FFN experts (gate, up, down projections) |
| **Qwen1.5-MoE / Qwen2-MoE** | `"mlp.experts.*.gate_proj"` , `"mlp.experts.*.up_proj"` , `"mlp.experts.*.down_proj"` | Full expert MLP submodules |
| **DeepSeek-V3** | `"experts.*.ffn"` or `"mlp.experts.*"` | Consolidated expert FFN blocks |
| **DBRX** | `"ffn.experts.*.w1"` , `"ffn.experts.*.w2"` | Granular expert weights |
| **OLMoE** | `"transformer.blocks.*.mlp.experts.*"` | Block-scoped expert patterns |

To verify patterns match your specific checkpoint, inspect module names:

```python
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("your-moe-model")
for name, _ in list(model.named_modules())[:30]:
    print(name)

```

Look for strings containing `experts`, `mlp`, or `ffn` to identify the correct pattern syntax.

## Troubleshooting Common Issues

**Pattern not applied**: Confirm `use_vera=False` and check that patterns actually match module names. Misplaced wildcards (`experts.*w1` vs `experts.*.w1`) are frequent errors.

**Validation errors on key count**: Consolidate patterns using broader wildcards. Instead of listing 64 individual experts, use `experts.*.w1`.

**Unexpected memory usage**: Lower ranks reduce parameter count but not necessarily activation memory. Combine `rank_pattern` with gradient checkpointing for MoE models.

**DoRA interaction**: `use_dora=True` is compatible with `rank_pattern`. The rank override applies to the base LoRA matrices before DoRA's weight-decomposed transformation.

## Summary

- **`rank_pattern` and `alpha_pattern`** in `LoraConfig` enable per-module LoRA hyperparameters via regex-like string matching.
- **MoE models benefit significantly** from reduced expert ranks (8-16) while keeping attention ranks high (64-256).
- **Validation limits**: 256 keys max, values 1-1024, incompatible with `use_vera=True`.
- **Configuration path**: [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) defines validation; [`src/soup_cli/utils/peft_builder.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/peft_builder.py) forwards to PEFT.
- **Best practice**: Verify module names with `named_modules()`, start with conservative rank reductions, and monitor training dynamics through the alpha scaling factor.

## Frequently Asked Questions

### What happens if a module matches multiple patterns in `rank_pattern`?

PEFT applies the **first matching pattern** in dictionary iteration order. Python 3.7+ preserves insertion order, so list your most specific patterns first. For predictable behavior, ensure pattern sets are mutually exclusive or ordered by specificity.

### Can I use `rank_pattern` with DoRA (Weight-Decomposed Low-Rank Adaptation)?

Yes. According to the Soup source in [`src/soup_cli/utils/peft_builder.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/peft_builder.py), DoRA is compatible with `rank_pattern` and `alpha_pattern`. The rank override applies to the base LoRA matrices (`A` and `B`), while DoRA's magnitude vector (`m`) operates on the decomposed result. Set `use_dora=True` alongside your patterns.

### Why does setting `rank_pattern` trigger a validation error about VeRA?

**VeRA** uses a single shared random matrix across all layers, eliminating the concept of per-module ranks. The validator in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) explicitly rejects this combination with `rank_pattern is incompatible with use_vera=True`. Choose either VeRA (simple, memory-efficient) or patterned LoRA (flexible, MoE-optimized)—not both.

### How do I find the exact module names for my MoE model?

Run this diagnostic to dump all module names:

```python
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("model-name", device_map="auto")
for name, module in model.named_modules():
    if "expert" in name.lower() or "mlp" in name.lower():
        print(f"{name}: {type(module).__name__}")

```

Match the printed names against your intended patterns. Test patterns with `re.search(pattern, name)` to confirm matches before adding to your configuration.