# Embedding Models Compatible with Unsloth: Supported Architectures and Fine-Tuning Guide

> Explore Unsloth compatible embedding models like BGE M3 and Gemma. Learn how to fine-tune these models efficiently with Unsloth's optimized approach for better performance and memory usage.

- Repository: [Unsloth AI/unsloth](https://github.com/unslothai/unsloth)
- Tags: how-to-guide
- Published: 2026-03-20

---

**Unsloth natively supports sentence-embedding models including all-MiniLM-L6-v2, BGE-M3, EmbeddingGemma, GTE-ModernBERT, and Qwen3-Embedding, enabling memory-efficient fine-tuning through a dedicated optimizer path that treats the embedding matrix with a separate learning rate.**

The unslothai/unsloth repository extends its memory-efficient training capabilities beyond large language models to dedicated embedding architectures. Understanding which embedding models are compatible with Unsloth and how they are utilized allows practitioners to fine-tune dense retrieval models with the same 2-5x memory savings that Unsloth provides for LLMs.

## Supported Embedding Models in Unsloth

Unsloth maintains a curated registry of embedding-compatible architectures in [`studio/backend/utils/models/model_config.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/utils/models/model_config.py) (lines 38-60). The following sentence-transformer models are officially supported with dedicated YAML configurations:

| Canonical Config | Hugging Face Identifier | Alias Resolution |
|-----------------|------------------------|------------------|
| [`unsloth_all-MiniLM-L6-v2.yaml`](https://github.com/unslothai/unsloth/blob/main/unsloth_all-MiniLM-L6-v2.yaml) | `unsloth/all-MiniLM-L6-v2` | `sentence-transformers/all-MiniLM-L6-v2` |
| [`unsloth_bge-m3.yaml`](https://github.com/unslothai/unsloth/blob/main/unsloth_bge-m3.yaml) | `unsloth/bge-m3` | `BAAI/bge-m3` |
| [`unsloth_embeddinggemma-300m.yaml`](https://github.com/unslothai/unsloth/blob/main/unsloth_embeddinggemma-300m.yaml) | `unsloth/embeddinggemma-300m` | `google/embeddinggemma-300m` |
| [`unsloth_gte-modernbert-base.yaml`](https://github.com/unslothai/unsloth/blob/main/unsloth_gte-modernbert-base.yaml) | `unsloth/gte-modernbert-base` | `Alibaba-NLP/gte-modernbert-base` |
| [`unsloth_Qwen3-Embedding-0.6B.yaml`](https://github.com/unslothai/unsloth/blob/main/unsloth_Qwen3-Embedding-0.6B.yaml) | `unsloth/Qwen3-Embedding-0.6B` | `Qwen/Qwen3-Embedding-0.6B` (also 4B) |

These mappings allow Unsloth to resolve community model names to canonical Hugging Face repositories while applying embedding-specific optimizations.

## How Unsloth Utilizes Embedding Models

Unsloth implements a specialized training pipeline for embedding models that differs from standard LLM fine-tuning. The architecture handles embedding-only models through three core mechanisms.

### Disabling SDPA for Embedding-Only Architectures

The model loader in [`unsloth/models/loader.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader.py) (lines 120-126) maintains a `DISABLE_SDPA_MODEL_NAMES` list that excludes sentence-embedding models from scaled-dot-product attention optimizations. Since embedding models typically bypass the causal attention mechanisms used in generative models, this flag prevents incompatible acceleration paths from loading.

### Dual Optimizer Groups for the Embedding Matrix

The `UnslothTrainer` class implements a split-parameter optimizer strategy in [`unsloth/trainer.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/trainer.py) (lines 43-77). When `embedding_learning_rate` is specified—either through the CLI flag or `UnslothTrainingArguments` (lines 33-36)—the `_create_unsloth_optimizer` function constructs two distinct parameter groups:

- **Non-embedding parameters**: Updated with the standard learning rate
- **Embedding weights** (`*.modules_to_save.default.weight`): Updated with the custom `embedding_learning_rate`

This separation enables stable fine-tuning of the dense embedding vectors without destabilizing the underlying transformer representations. The wrapper in `create_optimizer` (lines 82-99) finalizes this configuration before training begins.

### FastSentenceTransformer Wrapper

The `FastSentenceTransformer` class defined in [`unsloth/models/sentence_transformer.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/sentence_transformer.py) (line 522) inherits from the base `FastModel` and patches the transformer environment to correctly expose the embedding matrix (see the implementation note at line 190). This wrapper ensures that embedding models integrate with Unsloth's memory-efficient backends while maintaining compatibility with the `sentence-transformers` ecosystem for saving and export.

## Fine-Tuning Embedding Models with Unsloth

Unsloth exposes embedding model training through both its CLI and Python API, with automatic handling of the specialized optimizer paths.

### Command-Line Training

Pass the `--embedding-learning-rate` flag to trigger the dual-optimizer behavior directly from the terminal:

```bash
unsloth train \
  --model unsloth/embeddinggemma-300m \
  --dataset my_embeddings.jsonl \
  --output_dir ./finetuned_embeddinggemma \
  --embedding-learning-rate 1e-4 \
  --epochs 3 \
  --batch_size 64

```

The CLI handler in [`unsloth_cli/commands/train.py`](https://github.com/unslothai/unsloth/blob/main/unsloth_cli/commands/train.py) forwards this flag to the trainer configuration, ensuring the embedding matrix receives the specified learning rate while other parameters use the default schedule.

### Python API Training

For programmatic control, use `FastSentenceTransformer` with `UnslothTrainingArguments`:

```python
from unsloth import FastSentenceTransformer, UnslothTrainer, UnslothTrainingArguments

# Load the base embedding model

model = FastSentenceTransformer.from_pretrained(
    "unsloth/embeddinggemma-300m",
    max_seq_length=256,
)

# Prepare dataset of (text, vector) pairs

train_dataset = [
    {"text": "I love pizza", "label": [0.12, -0.34, 0.56]},
    {"text": "Python is great", "label": [0.05, 0.78, -0.21]},
]

args = UnslothTrainingArguments(
    output_dir="./finetuned_embeddinggemma",
    per_device_train_batch_size=32,
    num_train_epochs=3,
    embedding_learning_rate=5e-5,  # Custom LR for embedding layer

)

trainer = UnslothTrainer(
    model=model,
    args=args,
    train_dataset=train_dataset,
)

trainer.train()
model.save_pretrained("./finetuned_embeddinggemma")

```

The `embedding_learning_rate` parameter in `UnslothTrainingArguments` triggers the dual-group optimizer creation automatically.

### Loading Fine-Tuned Embeddings

After training, load the model for inference using the standard sentence-transformers interface:

```python
from unsloth import FastSentenceTransformer

model = FastSentenceTransformer.from_pretrained("./finetuned_embeddinggemma")
embeddings = model.encode(["First sentence", "Second sentence"])
print(embeddings.shape)  # Output: (2, 384) for MiniLM-L6-v2

```

The `encode` method forwards to the model's `get_input_embeddings`, maintaining compatibility with existing retrieval pipelines.

## Summary

- Unsloth officially supports five embedding architectures: **all-MiniLM-L6-v2**, **BGE-M3**, **EmbeddingGemma-300M**, **GTE-ModernBERT**, and **Qwen3-Embedding**, mapped in [`studio/backend/utils/models/model_config.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/utils/models/model_config.py).
- The loader disables **SDPA** for embedding models via `DISABLE_SDPA_MODEL_NAMES` in [`unsloth/models/loader.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader.py) to prevent attention mechanism conflicts.
- Training utilizes a **dual-optimizer strategy** where `UnslothTrainer` applies a separate `embedding_learning_rate` to the embedding matrix while keeping standard LR for transformer weights.
- **FastSentenceTransformer** provides the integration layer between Unsloth's efficient kernels and the sentence-transformers ecosystem.
- Both CLI (`--embedding-learning-rate`) and Python API (`UnslothTrainingArguments`) expose embedding-specific training controls.

## Frequently Asked Questions

### Which embedding models can I fine-tune with Unsloth?

Unsloth supports **sentence-transformers/all-MiniLM-L6-v2**, **BAAI/bge-m3**, **google/embeddinggemma-300m**, **Alibaba-NLP/gte-modernbert-base**, and **Qwen/Qwen3-Embedding-0.6B** (along with the 4B variant). These are registered in [`studio/backend/utils/models/model_config.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/utils/models/model_config.py) with canonical YAML configurations that resolve to the underlying Hugging Face repositories.

### Why does Unsloth disable SDPA for embedding models?

Unsloth disables **scaled-dot-product attention (SDPA)** for embedding-only models through the `DISABLE_SDPA_MODEL_NAMES` list in [`unsloth/models/loader.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader.py) (lines 120-126). Embedding architectures typically do not utilize causal self-attention mechanisms, making SDPA optimizations incompatible or unnecessary for these model types.

### How do I set a different learning rate for the embedding layer?

Pass the `--embedding-learning-rate` flag via CLI or specify `embedding_learning_rate` in `UnslothTrainingArguments`. The `UnslothTrainer` automatically splits parameters into two optimizer groups in `_create_unsloth_optimizer` ([`unsloth/trainer.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/trainer.py) lines 43-77), applying your custom rate only to the embedding matrix weights (`*.modules_to_save.default.weight`) while using the standard learning rate for all other parameters.

### Can I export Unsloth fine-tuned embedding models to GGUF?

Yes. The `FastSentenceTransformer` class in [`unsloth/models/sentence_transformer.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/sentence_transformer.py) handles GGUF conversion and export through its `save_pretrained` method. After training with `UnslothTrainer`, calling `model.save_pretrained()` generates compatible artifacts for the sentence-transformers ecosystem or quantized GGUF formats.