How to Optimize Prompts Using DSPy's MIPROv2 in Sieves: A Complete Guide

Sieves provides a high-level Optimization task that wraps DSPy's MIPROv2 optimizer to automatically improve prompt templates and few-shot examples through Bayesian search, requiring no direct DSPy code from users.

The mantisai/sieves repository simplifies prompt engineering by integrating DSPy's state-of-the-art optimizers directly into its task framework. When you need to optimize prompts using DSPy's MIPROv2 in Sieves, the library abstracts the complex Bayesian optimization process into a single Optimization task that handles instruction tuning and few-shot selection automatically.

Understanding MIPROv2 Integration in Sieves

Core Implementation Details

The integration lives in sieves/tasks/optimization/core.py, where the Optimization class encapsulates DSPy's MIPROv2 optimizer. According to the class docstring, it "Uses MIPROv2 to optimize instructions and few‑shot examples"【/sieves/tasks/optimization/core.py†L17-L18】.

The optimizer instantiation occurs at lines 63-66:

teleprompter = dspy.MIPROv2(
    metric=evaluate, 
    **(self._init_kwargs or {}), 
    verbose=False
)

This creates a MIPROv2 instance that uses your specified metric to evaluate prompt candidates during the optimization loop.

How Bayesian Optimization Works

MIPROv2 operates by Bayesian-searching over the space of prompt templates and few-shot example selections. The optimizer:

  1. Generates candidate prompt configurations using a validation split of your labeled data
  2. Compiles each candidate using teleprompter.compile()
  3. Scores each configuration using the supplied evaluate metric (such as accuracy for classification tasks)
  4. Returns the highest-scoring configuration as an optimized task

This process is documented in docs/guides/optimization.md, which explains the workflow, cost considerations, and tuning parameters like num_trials and num_candidates【/docs/guides/optimization.md†L12-L20】【/docs/guides/optimization.md†L97-L104】.

Setting Up Your Optimization Task

To optimize prompts using DSPy's MIPROv2 in Sieves, you wrap your existing predictive task with the Optimization class. Here is the complete workflow:


# 1️⃣ Import the required Sieves components

from sieves import Doc, Pipeline
from sieves.tasks.predictive.classification.core import Classification
from sieves.tasks.optimization.core import Optimization
import dspy

# 2️⃣ Define a simple classification task (uses Outlines by default)

classifier = Classification(
    model=dspy.LM(model="gpt-4o-mini"),   # any dspy LM compatible model

    metric="accuracy",                     # metric used for optimization

)

# 3️⃣ Create a few‑shot training set (list of Docs with gold labels)

train_examples = [
    Doc(text="I love this product!", gold={"label": "positive"}),
    Doc(text="Terrible experience.", gold={"label": "negative"}),
    # …add more labeled examples

]

# 4️⃣ Wrap the task with the Optimization helper

optim = Optimization(
    task=classifier,
    training_data=train_examples,
    num_trials=30,          # how many Bayesian trials to run

    num_candidates=5,       # how many prompt/example combos per trial

    metric="accuracy",      # matches the task’s metric

)

# 5️⃣ Run the optimizer – this will invoke DSPy MIPROv2 internally

optimized_task = optim()

# 6️⃣ Use the optimized task in a pipeline

pipe = Pipeline([optimized_task])
docs = [Doc(text="The service was okay.")]
result = pipe(docs)
print(result[0].results[optimized_task.id])   # → optimized prediction

This example demonstrates how Sieves abstracts the underlying DSPy implementation. The Optimization task handles the instantiation of dspy.MIPROv2, the Bayesian search loop, and the compilation of the best-performing configuration.

Configuring MIPROv2 Parameters for Better Results

Tuning num_trials and num_candidates

The effectiveness of your optimization depends on two critical parameters exposed through the Optimization task:

  • num_trials: Controls how many Bayesian optimization iterations MIPROv2 performs. Higher values explore more of the prompt space but increase API costs.
  • num_candidates: Determines how many prompt/example combinations are generated per trial. This affects the diversity of candidates evaluated during each iteration.

According to the documentation in docs/guides/optimization.md, these parameters directly impact both the quality of the optimized prompts and the computational cost of the process【/docs/guides/optimization.md†L97-L104】.

Selecting the Right Metric

MIPROv2 requires a metric function to evaluate prompt candidates. In Sieves, this is handled automatically when you specify a metric during task creation (such as "accuracy" for classification). The metric is passed to the MIPROv2 constructor as seen in sieves/tasks/optimization/core.py:

teleprompter = dspy.MIPROv2(metric=evaluate, ...)

The evaluate function wraps your specified metric and compares model outputs against the gold labels in your training Doc objects. Ensure your metric aligns with your task type—classification tasks typically use accuracy or F1 score, while extraction tasks might use token-level precision.

Summary

  • Sieves abstracts DSPy's MIPROv2 through the Optimization task in sieves/tasks/optimization/core.py, eliminating the need for direct DSPy code.
  • Bayesian optimization powers the prompt improvement process, searching over instruction templates and few-shot example combinations.
  • Configuration happens through parameters like num_trials and num_candidates, which trade off optimization quality against API costs.
  • Integration requires only wrapping your existing predictive task with Optimization and providing labeled training data as Doc objects.

Frequently Asked Questions

What is the difference between MIPROv2 and other DSPy optimizers?

MIPROv2 specifically focuses on optimizing instructions and few-shot examples through Bayesian search, whereas other DSPy optimizers like BootstrapFewShot or COPRO use different strategies such as bootstrapping demonstrations or coordinate ascent. According to the Sieves source code in sieves/tasks/optimization/core.py, the Optimization task specifically wraps MIPROv2 because it provides superior performance for instruction tuning while maintaining compatibility with Sieves' task architecture.

How do I choose the right metric for MIPROv2 optimization?

Select a metric that aligns with your specific task objective. For classification tasks, use "accuracy" or "f1" depending on your class balance. For extraction or NER tasks, consider token-level metrics like precision or recall. The metric is passed to the underlying dspy.MIPROv2 constructor and used to score each candidate configuration during the Bayesian search process. Ensure your training Doc objects include the corresponding gold labels so the metric can compare predictions against ground truth.

Can I use MIPROv2 with custom predictive tasks in Sieves?

Yes, the Optimization task is designed to work with any Sieves PredictiveTask that exposes a compatible interface. As long as your custom task provides the necessary methods for compilation and defines a metric for evaluation, you can wrap it with Optimization exactly as you would with built-in tasks like Classification. The optimizer will treat your custom task as a black box during the Bayesian search, testing different prompt configurations to maximize your specified metric.

What are the cost considerations when running MIPROv2 optimization?

MIPROv2 optimization consumes API tokens proportional to the product of num_trials, num_candidates, and the size of your validation dataset. Each trial evaluates multiple candidate prompts against your labeled examples, meaning costs scale with the thoroughness of the search. The documentation in docs/guides/optimization.md recommends starting with smaller values (e.g., num_trials=10, num_candidates=3) to estimate costs before running full optimization. Consider using cheaper models like gpt-4o-mini during optimization, then switching to more powerful models for production inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →