How to Apply Model Distillation with SetFit and Model2Vec in Sieves
Use the task.distill() API in Sieves to fine-tune a lightweight SetFit or Model2Vec student model on predictions from a large zero-shot teacher, enabling fast, cost-effective inference without external LLM calls.
Sieves provides a unified framework for model distillation that bridges high-accuracy zero-shot classification with efficient, task-specific student models. Whether you need the balanced accuracy of SetFit or the extreme speed of Model2Vec, the distill() method in the Classification task orchestrates the entire teacher-student pipeline.
Understanding the Distillation Pipeline in Sieves
Model distillation in Sieves follows a three-stage workflow implemented in sieves/tasks/predictive/classification/core.py. First, a teacher model—typically a large language model like GPT-4o-mini—generates predictions on your documents. Second, these predictions are exported to a Hugging Face datasets.Dataset via task.to_hf_dataset(). Finally, a student model is fine-tuned on this dataset using either SetFit or Model2Vec, depending on your latency and accuracy requirements.
The framework switch occurs at lines 310-340 of the core classification file, where Sieves branches between DistillationFramework.setfit and DistillationFramework.model2vec.
Prerequisites and Installation
Distillation requires optional dependencies that are not included in the base Sieves package. Install them using the distill extra:
pip install "sieves[distill]"
This installs SetFit, Model2Vec, and Sentence Transformers. If these libraries are missing, sieves/tasks/distillation/distillation_import.py raises a clear import error at runtime (lines 11-27).
Step-by-Step Distillation Workflow
Step 1: Generate Teacher Predictions with a Zero-Shot Model
Begin by running a zero-shot classification task to generate soft labels for your dataset. The teacher can be any supported LLM, such as OpenAI's GPT-4o-mini.
from sieves import Doc, Pipeline
from sieves.tasks.predictive.classification.core import Classification
# Prepare documents
docs = [
Doc(text="AI breakthrough announced in neural architecture search"),
Doc(text="Election results show tight race in swing states"),
Doc(text="Championship game ends in overtime thriller")
]
# Initialize teacher model
teacher = Classification(
labels=["tech", "politics", "sports"],
model="openai:gpt-4o-mini",
mode="single",
)
# Generate predictions (stored in doc.results["classification"])
teacher(docs)
Step 2: Distill Knowledge into a SetFit Student Model
SetFit is ideal when you need strong accuracy with modest training data. It uses Sentence Transformers to generate embeddings, then trains a classifier head. This approach typically outperforms standard fine-tuning on small datasets.
from sieves.tasks.distillation.types import DistillationFramework
from pathlib import Path
# Distill using SetFit
teacher.distill(
base_model_id="sentence-transformers/paraphrase-MiniLM-L3-v2",
framework=DistillationFramework.setfit,
data=docs,
output_path=Path("distilled_setfit"),
val_frac=0.3,
init_kwargs={"multi_target_strategy": "multi-output"},
train_kwargs={"num_train_epochs": 5},
)
Implementation details: The SetFit branch (lines 44-66 in core.py) creates a SetFitModel, configures a Trainer with your specified arguments, and persists the model to distilled_setfit/. Evaluation metrics are saved to distilled_setfit/metrics.json.
Step 3: Alternative - Distill with Model2Vec for Maximum Speed
Model2Vec provides ultra-lightweight linear classifiers based on static embeddings. Choose this framework when inference speed or memory usage is critical—Model2Vec models are typically orders of magnitude faster than transformer-based approaches, though they may sacrifice some accuracy on complex tasks.
# Distill using Model2Vec
teacher.distill(
base_model_id="facebook/fasttext-wiki-news-subwords-300",
framework=DistillationFramework.model2vec,
data=docs,
output_path=Path("distilled_model2vec"),
val_frac=0.2,
train_kwargs={"max_iter": 200},
)
Implementation details: The Model2Vec branch (lines 70-88 in core.py) loads a StaticModelForClassification, fits it on the training split, converts it to a Pipeline, and saves the artifacts to distilled_model2vec/. Metrics are stored in distilled_model2vec/metrics.json.
Loading and Deploying Your Distilled Model
Once training completes, load the distilled student model directly into a Classification task. Sieves recognizes the framework prefix (setfit: or model2vec:) and resolves the local path automatically.
# Load SetFit student
student = Classification(
labels=["tech", "politics", "sports"],
model="setfit:distilled_setfit",
)
# Run inference - no external API calls required
student(docs)
print(docs[0].results["classification"])
For Model2Vec, use the model2vec: prefix:
student = Classification(
labels=["tech", "politics", "sports"],
model="model2vec:distilled_model2vec",
)
Implementation Details and Source Code
The distillation logic is centralized in sieves/tasks/predictive/classification/core.py. The distill method validates your dataset, splits it into train/validation sets via _split_dataset, and branches based on the DistillationFramework enum at lines 310-340.
-
SetFit implementation (lines 44-66): Instantiates
SetFitModelfrom the Hugging Face checkpoint, configures aTrainerwithTrainingArguments, executes fine-tuning, and persists the model directory andmetrics.json. -
Model2Vec implementation (lines 70-88): Loads
StaticModelForClassification, fits the model on the training data, converts it to aPipelineobject, and saves the model artifacts alongside evaluation metrics.
Optional dependencies are managed through sieves/tasks/distillation/distillation_import.py (lines 11-27), which raises informative errors if setfit, model2vec, or sentence_transformers are not installed.
Summary
- Sieves unifies distillation through the
task.distill()method in theClassificationtask, enabling you to convert expensive LLM predictions into efficient student models. - Choose SetFit when you need high accuracy with limited training data; it leverages Sentence Transformers and trains a robust classifier head.
- Choose Model2Vec when inference speed and memory efficiency are paramount; it uses static embeddings and linear classifiers for near-instant prediction.
- Implementation resides in
sieves/tasks/predictive/classification/core.py, with framework-specific logic at lines 44-88 and the main orchestration at lines 310-340. - Deployment is seamless: Load distilled models using the
setfit:ormodel2vec:prefixes in theClassificationtask constructor.
Frequently Asked Questions
What is the difference between SetFit and Model2Vec distillation in Sieves?
SetFit uses Sentence Transformer embeddings to train a classifier head, offering strong accuracy even with small datasets but requiring more inference compute than Model2Vec. Model2Vec relies on static embeddings (like FastText) and linear classifiers, making it significantly faster and lighter at the cost of some accuracy on complex semantic tasks. Choose SetFit for accuracy, Model2Vec for speed.
How do I install the required dependencies for distillation?
Run pip install "sieves[distill]" to install the optional extras. This command installs setfit, model2vec, and sentence_transformers. If you attempt to use distill() without these packages, sieves/tasks/distillation/distillation_import.py raises a clear import error indicating which library is missing.
Can I use a custom Hugging Face model for the student?
Yes. Pass any valid Hugging Face model ID to the base_model_id parameter in distill(). For SetFit, use any Sentence Transformers checkpoint (e.g., "sentence-transformers/all-MiniLM-L6-v2"). For Model2Vec, use any compatible static embedding model (e.g., "minishlab/potion-base-8M"). The framework handles model initialization automatically.
Where are the trained student models and metrics saved?
The output_path parameter in distill() specifies the directory for artifacts. For both frameworks, the fine-tuned model is saved to that directory, and evaluation metrics are written to metrics.json inside the same directory. When loading the model later, Sieves reads these artifacts to reconstruct the Classification task.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →