How to Add Few-Shot Examples to Tasks in Sieves: A Complete Guide
Pass a list of task-specific FewshotExample objects to the fewshot_examples parameter when initializing any PredictiveTask subclass in Sieves to inject contextual examples into the model prompt.
Sieves is a modular NLP framework by Mantis AI that standardizes how predictive tasks interact with language models. Adding few-shot examples to tasks in Sieves allows you to condition model outputs by providing concrete input-output pairs directly in the prompt, improving accuracy without fine-tuning.
Understanding Few-Shot Support in Sieves
The PredictiveTask Base Class
All few-shot functionality in Sieves originates in the PredictiveTask abstract base class located in sieves/tasks/predictive/core.py. When you initialize any predictive task—whether classification, sentiment analysis, or a custom implementation—the constructor accepts a fewshot_examples parameter at lines 45-55:
def __init__(
self,
...,
fewshot_examples: Sequence[FewshotExample] | None = None,
...
):
self._fewshot_examples = fewshot_examples
self._validate_fewshot_examples()
The base class stores these examples internally and immediately triggers validation through the _validate_fewshot_examples method.
Validation and Schema Enforcement
Each concrete task must define what constitutes a valid few-shot example. This is enforced through two mechanisms:
-
The
fewshot_example_typeproperty: Every task subclass must implement this property to return the specific Pydantic model class representing its few-shot schema. For example,SentimentAnalysisinsieves/tasks/predictive/sentiment_analysis/core.py(lines 78-86) returns its specificFewshotExampleschema. -
The
_validate_fewshot_exampleshook: Tasks override this method to enforce business logic. InClassification(sieves/tasks/predictive/classification/core.py, lines 53-80), this method ensures that every few-shot example contains labels that exist in the task's declared label set, raising aValueErrorimmediately if you provide an invalid label.
During inference, the bridge layer in sieves/tasks/predictive/bridges.py (lines 92-99) formats these validated examples as XML (or model-specific format) and injects them into the prompt before sending to the model wrapper.
How to Add Few-Shot Examples to Classification Tasks
For single-label classification, import FewshotExampleSingleLabel from sieves/tasks/predictive/schemas/classification.py and pass a list to the fewshot_examples parameter:
from sieves.tasks.predictive.classification.core import Classification
from sieves.tasks.predictive.schemas.classification import FewshotExampleSingleLabel
# Build few-shot examples matching the task's label set
few_shot = [
FewshotExampleSingleLabel(
text="The movie was thrilling and full of action.",
label="positive",
score=0.96,
),
FewshotExampleSingleLabel(
text="I disliked the plot and the characters were flat.",
label="negative",
score=0.92,
),
]
# Instantiate task with few-shot examples
cls_task = Classification(
model="gpt-4o-mini",
labels=("positive", "negative"),
fewshot_examples=few_shot,
)
# Run inference - few-shot examples are injected into the prompt
docs = [Doc(text="An amazing visual experience.")]
result_docs = cls_task(docs)
The Classification class validates these examples against its label set in _validate_fewshot_examples (lines 53-80 of sieves/tasks/predictive/classification/core.py), ensuring you cannot accidentally pass a label like "neutral" if it wasn't declared in the task's labels parameter.
How to Add Few-Shot Examples to Sentiment Analysis
For aspect-based sentiment analysis, use the FewshotExample schema from sieves/tasks/predictive/schemas/sentiment_analysis.py:
from sieves.tasks.predictive.sentiment_analysis.core import SentimentAnalysis
from sieves.tasks.predictive.schemas.sentiment_analysis import FewshotExample
few_shot = [
FewshotExample(
text="The battery life is impressive, but the screen is dull.",
overall=0.7,
aspect_sentiment={"battery": 0.9, "screen": 0.3},
),
FewshotExample(
text="Fast processor, terrible camera.",
overall=0.6,
aspect_sentiment={"processor": 0.95, "camera": 0.2},
),
]
sent_task = SentimentAnalysis(
model="gpt-4o-mini",
aspects=("battery", "screen", "processor", "camera"),
fewshot_examples=few_shot,
)
docs = [Doc(text="The camera is sharp but the battery drains quickly.")]
out = sent_task(docs)
The SentimentAnalysis task declares its few-shot schema via the fewshot_example_type property (lines 78-86 of sieves/tasks/predictive/sentiment_analysis/core.py), which returns the FewshotExample class defined in the corresponding schema file.
Creating Custom Predictive Tasks with Few-Shot Support
To add few-shot capability to a custom predictive task, inherit from PredictiveTask and implement two hooks:
from sieves.tasks.predictive.core import PredictiveTask
from sieves.tasks.predictive.schemas.core import FewshotExample as BaseFS
import pydantic
class MyFewshotExample(BaseFS):
"""Custom few-shot schema."""
custom_output: str
class MyTask(PredictiveTask[MyPrompt, MyResult, MyBridge]):
@property
def fewshot_example_type(self):
"""Return the Pydantic model for validation."""
return MyFewshotExample
def _validate_fewshot_examples(self):
"""Custom validation logic."""
for ex in self._fewshot_examples or []:
if not hasattr(ex, "custom_output"):
raise ValueError("Missing `custom_output` in few-shot example.")
# Add task-specific validation here
Now instantiate with examples:
task = MyTask(
model="gpt-4o-mini",
fewshot_examples=[
MyFewshotExample(text="Input text here", custom_output="Expected output")
]
)
The base class handles storage and bridge integration, while your implementation controls what constitutes a valid example.
Summary
- Few-shot examples in Sieves are passed via the
fewshot_examplesparameter available on allPredictiveTasksubclasses. - Validation is task-specific: Each task defines its schema via
fewshot_example_typeand validates examples in_validate_fewshot_examples(e.g.,Classificationchecks against label sets insieves/tasks/predictive/classification/core.pylines 53-80). - Schema selection matters: Import the correct
FewshotExamplesubclass for your task (e.g.,FewshotExampleSingleLabelfor single-label classification fromsieves/tasks/predictive/schemas/classification.py). - Custom tasks require two hooks: Implement
fewshot_example_type(property) and_validate_fewshot_examples(method) when extendingPredictiveTask. - Prompt injection is automatic: The bridge layer (
sieves/tasks/predictive/bridges.pylines 92-99) formats examples as XML and injects them into the model prompt duringtask(docs)execution.
Frequently Asked Questions
What is the maximum number of few-shot examples I can add?
There is no hardcoded limit in the Sieves framework itself. The practical limit depends on your model's context window and the length of your input documents. Each few-shot example consumes tokens in the prompt, so monitor your context window usage when adding more than 10-20 examples. The PredictiveTask base class stores examples in a simple list (self._fewshot_examples), so memory overhead is minimal until the bridge formats them for the prompt.
How does Sieves validate few-shot examples?
Validation occurs immediately during task construction through the _validate_fewshot_examples hook. Each concrete task implements its own validation logic. For example, the Classification task in sieves/tasks/predictive/classification/core.py (lines 53-80) verifies that every example's label exists in the task's declared labels parameter. The SentimentAnalysis task checks that aspect names match the declared aspects. If validation fails, a ValueError is raised before any inference occurs, preventing runtime prompt errors.
Can I use different few-shot examples for different model providers?
Yes, because few-shot examples are task-level configuration, not model-level. You instantiate the task with specific examples, then pass that task to any supported model bridge. However, if you need provider-specific formatting, you would create separate task instances with different example sets, or implement a custom bridge in sieves/tasks/predictive/bridges.py that handles provider-specific XML formatting. The examples themselves (Pydantic models) remain provider-agnostic until the bridge formats them for the specific model's prompt structure.
What happens if my few-shot example has an invalid label?
The task raises a ValueError immediately upon instantiation. For instance, if you create a Classification task with labels=("positive", "negative") but include a few-shot example with label="neutral", the _validate_fewshot_examples method in sieves/tasks/predictive/classification/core.py detects the mismatch and raises an error with a descriptive message. This fail-fast approach prevents silent failures during inference and ensures prompt consistency. You must fix the label in your example data or expand the task's allowed labels to include the new category.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →