Embedding Models Compatible with Unsloth: Supported Architectures and Fine-Tuning Guide

Unsloth natively supports sentence-embedding models including all-MiniLM-L6-v2, BGE-M3, EmbeddingGemma, GTE-ModernBERT, and Qwen3-Embedding, enabling memory-efficient fine-tuning through a dedicated optimizer path that treats the embedding matrix with a separate learning rate.

The unslothai/unsloth repository extends its memory-efficient training capabilities beyond large language models to dedicated embedding architectures. Understanding which embedding models are compatible with Unsloth and how they are utilized allows practitioners to fine-tune dense retrieval models with the same 2-5x memory savings that Unsloth provides for LLMs.

Supported Embedding Models in Unsloth

Unsloth maintains a curated registry of embedding-compatible architectures in studio/backend/utils/models/model_config.py (lines 38-60). The following sentence-transformer models are officially supported with dedicated YAML configurations:

Canonical Config Hugging Face Identifier Alias Resolution
unsloth_all-MiniLM-L6-v2.yaml unsloth/all-MiniLM-L6-v2 sentence-transformers/all-MiniLM-L6-v2
unsloth_bge-m3.yaml unsloth/bge-m3 BAAI/bge-m3
unsloth_embeddinggemma-300m.yaml unsloth/embeddinggemma-300m google/embeddinggemma-300m
unsloth_gte-modernbert-base.yaml unsloth/gte-modernbert-base Alibaba-NLP/gte-modernbert-base
unsloth_Qwen3-Embedding-0.6B.yaml unsloth/Qwen3-Embedding-0.6B Qwen/Qwen3-Embedding-0.6B (also 4B)

These mappings allow Unsloth to resolve community model names to canonical Hugging Face repositories while applying embedding-specific optimizations.

How Unsloth Utilizes Embedding Models

Unsloth implements a specialized training pipeline for embedding models that differs from standard LLM fine-tuning. The architecture handles embedding-only models through three core mechanisms.

Disabling SDPA for Embedding-Only Architectures

The model loader in unsloth/models/loader.py (lines 120-126) maintains a DISABLE_SDPA_MODEL_NAMES list that excludes sentence-embedding models from scaled-dot-product attention optimizations. Since embedding models typically bypass the causal attention mechanisms used in generative models, this flag prevents incompatible acceleration paths from loading.

Dual Optimizer Groups for the Embedding Matrix

The UnslothTrainer class implements a split-parameter optimizer strategy in unsloth/trainer.py (lines 43-77). When embedding_learning_rate is specified—either through the CLI flag or UnslothTrainingArguments (lines 33-36)—the _create_unsloth_optimizer function constructs two distinct parameter groups:

  • Non-embedding parameters: Updated with the standard learning rate
  • Embedding weights (*.modules_to_save.default.weight): Updated with the custom embedding_learning_rate

This separation enables stable fine-tuning of the dense embedding vectors without destabilizing the underlying transformer representations. The wrapper in create_optimizer (lines 82-99) finalizes this configuration before training begins.

FastSentenceTransformer Wrapper

The FastSentenceTransformer class defined in unsloth/models/sentence_transformer.py (line 522) inherits from the base FastModel and patches the transformer environment to correctly expose the embedding matrix (see the implementation note at line 190). This wrapper ensures that embedding models integrate with Unsloth's memory-efficient backends while maintaining compatibility with the sentence-transformers ecosystem for saving and export.

Fine-Tuning Embedding Models with Unsloth

Unsloth exposes embedding model training through both its CLI and Python API, with automatic handling of the specialized optimizer paths.

Command-Line Training

Pass the --embedding-learning-rate flag to trigger the dual-optimizer behavior directly from the terminal:

unsloth train \
  --model unsloth/embeddinggemma-300m \
  --dataset my_embeddings.jsonl \
  --output_dir ./finetuned_embeddinggemma \
  --embedding-learning-rate 1e-4 \
  --epochs 3 \
  --batch_size 64

The CLI handler in unsloth_cli/commands/train.py forwards this flag to the trainer configuration, ensuring the embedding matrix receives the specified learning rate while other parameters use the default schedule.

Python API Training

For programmatic control, use FastSentenceTransformer with UnslothTrainingArguments:

from unsloth import FastSentenceTransformer, UnslothTrainer, UnslothTrainingArguments

# Load the base embedding model

model = FastSentenceTransformer.from_pretrained(
    "unsloth/embeddinggemma-300m",
    max_seq_length=256,
)

# Prepare dataset of (text, vector) pairs

train_dataset = [
    {"text": "I love pizza", "label": [0.12, -0.34, 0.56]},
    {"text": "Python is great", "label": [0.05, 0.78, -0.21]},
]

args = UnslothTrainingArguments(
    output_dir="./finetuned_embeddinggemma",
    per_device_train_batch_size=32,
    num_train_epochs=3,
    embedding_learning_rate=5e-5,  # Custom LR for embedding layer

)

trainer = UnslothTrainer(
    model=model,
    args=args,
    train_dataset=train_dataset,
)

trainer.train()
model.save_pretrained("./finetuned_embeddinggemma")

The embedding_learning_rate parameter in UnslothTrainingArguments triggers the dual-group optimizer creation automatically.

Loading Fine-Tuned Embeddings

After training, load the model for inference using the standard sentence-transformers interface:

from unsloth import FastSentenceTransformer

model = FastSentenceTransformer.from_pretrained("./finetuned_embeddinggemma")
embeddings = model.encode(["First sentence", "Second sentence"])
print(embeddings.shape)  # Output: (2, 384) for MiniLM-L6-v2

The encode method forwards to the model's get_input_embeddings, maintaining compatibility with existing retrieval pipelines.

Summary

  • Unsloth officially supports five embedding architectures: all-MiniLM-L6-v2, BGE-M3, EmbeddingGemma-300M, GTE-ModernBERT, and Qwen3-Embedding, mapped in studio/backend/utils/models/model_config.py.
  • The loader disables SDPA for embedding models via DISABLE_SDPA_MODEL_NAMES in unsloth/models/loader.py to prevent attention mechanism conflicts.
  • Training utilizes a dual-optimizer strategy where UnslothTrainer applies a separate embedding_learning_rate to the embedding matrix while keeping standard LR for transformer weights.
  • FastSentenceTransformer provides the integration layer between Unsloth's efficient kernels and the sentence-transformers ecosystem.
  • Both CLI (--embedding-learning-rate) and Python API (UnslothTrainingArguments) expose embedding-specific training controls.

Frequently Asked Questions

Which embedding models can I fine-tune with Unsloth?

Unsloth supports sentence-transformers/all-MiniLM-L6-v2, BAAI/bge-m3, google/embeddinggemma-300m, Alibaba-NLP/gte-modernbert-base, and Qwen/Qwen3-Embedding-0.6B (along with the 4B variant). These are registered in studio/backend/utils/models/model_config.py with canonical YAML configurations that resolve to the underlying Hugging Face repositories.

Why does Unsloth disable SDPA for embedding models?

Unsloth disables scaled-dot-product attention (SDPA) for embedding-only models through the DISABLE_SDPA_MODEL_NAMES list in unsloth/models/loader.py (lines 120-126). Embedding architectures typically do not utilize causal self-attention mechanisms, making SDPA optimizations incompatible or unnecessary for these model types.

How do I set a different learning rate for the embedding layer?

Pass the --embedding-learning-rate flag via CLI or specify embedding_learning_rate in UnslothTrainingArguments. The UnslothTrainer automatically splits parameters into two optimizer groups in _create_unsloth_optimizer (unsloth/trainer.py lines 43-77), applying your custom rate only to the embedding matrix weights (*.modules_to_save.default.weight) while using the standard learning rate for all other parameters.

Can I export Unsloth fine-tuned embedding models to GGUF?

Yes. The FastSentenceTransformer class in unsloth/models/sentence_transformer.py handles GGUF conversion and export through its save_pretrained method. After training with UnslothTrainer, calling model.save_pretrained() generates compatible artifacts for the sentence-transformers ecosystem or quantized GGUF formats.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →