How to Create Custom Tasks and Bridges in Sieves Pipelines: A Complete Guide
To create custom tasks and bridges in Sieves pipelines, subclass Task for deterministic operations or PredictiveTask with a custom Bridge for LLM-powered tasks, then implement the required abstract methods like _call, integrate, and prompt_signature.
The mantisai/sieves library provides a modular framework for composing NLP pipelines through reusable components. When you need to create custom tasks and bridges in Sieves pipelines to handle domain-specific logic or integrate custom language models, the framework offers lightweight, fully typed base classes that handle batching, caching, and serialization automatically.
Understanding the Core Architecture of Sieves Tasks
Before implementing custom components, you need to understand the three fundamental abstractions in sieves/tasks/core.py, sieves/tasks/predictive/core.py, and sieves/tasks/predictive/bridges.py.
The Task Base Class (sieves/tasks/core.py)
The Task class in sieves/tasks/core.py is the foundation for all pipeline operations. It handles batching, conditional execution, and result plumbing through these key mechanisms:
__init__(task_id, include_meta, batch_size, condition)– Configures the task ID, metadata handling, batch size, and an optional per-document filter function.__call__(docs)– Iterates over documents, respects thecondition, batches passing documents, and forwards them to_call. Documents failing the condition receiveNoneas their result._call(self, docs)– Abstract method that subclasses must implement; receives only the documents that passed the condition.serialize / deserialize– Enable pipeline persistence and caching.
The PredictiveTask for LLM Operations (sieves/tasks/predictive/core.py)
PredictiveTask extends Task for operations that call language models through a Bridge. It is generic over three types:
PredictiveTask[PromptSignature, Result, BridgeClass]
PromptSignature– The Pydantic model describing the model's output structure.Result– The type stored indoc.results.BridgeClass– The concreteBridgeimplementation for a specific model wrapper (Outlines, DSPy, LangChain).
Key responsibilities include:
- Delegating prompt building, extraction, integration, and consolidation to its bridge.
- Supplying helpers for few-shot examples and dataset export via
to_hf_dataset. - Declaring supported model types through the
supportsproperty.
The Bridge Abstraction (sieves/tasks/predictive/bridges.py)
Bridge in sieves/tasks/predictive/bridges.py is the model-agnostic contract that connects predictive tasks to concrete model wrappers. Constructor arguments include the task ID, optional custom prompt, model settings, Pydantic signature, model type, and few-shot examples.
Core abstract properties to implement:
_default_prompt_instructions– Fallback instructions when the user does not provide custom ones.inference_mode– TheModelWrapperInferenceMode(e.g.,json,choice,regex).integrate(self, results, docs)– Writes the model's output back into eachDoc(typically intodoc.results).consolidate(self, results, docs_offsets)– Combines chunk-level results into document-level results (required for chunk-aware pipelines).
Helper properties like _prompt_instructions, _prompt_example_xml, and prompt_template automatically construct ready-to-use prompt strings from defaults, few-shot examples, and conclusions.
How to Create a Custom Task in Sieves
For deterministic operations that don't require language models, subclass Task and implement the _call method. This example from sieves/tests/docs/test_custom_tasks.py demonstrates a character counting task:
from sieves.tasks.core import Task
from sieves.data import Doc
from typing import Iterable
class CharCountTask(Task):
"""Counts characters in doc.text."""
def _call(self, docs: Iterable[Doc]) -> Iterable[Doc]:
for doc in docs:
doc.results[self.id] = len(doc.text or "")
yield doc
The task automatically inherits batching and conditional execution from Task.__call__. You can instantiate it with standard arguments:
task = CharCountTask(task_id="CharCount", include_meta=False, batch_size=-1)
docs = [Doc(text="Hello"), Doc(text="World!")]
list(task(docs)) # Docs with results["CharCount"] = 5, 6
Use the condition argument to skip documents based on custom logic:
task = CharCountTask(condition=lambda d: len(d.text) > 0)
How to Create a Custom Predictive Task and Bridge
For LLM-powered operations, you must implement both a Bridge (handling model-specific prompt building and result integration) and a PredictiveTask (wiring the bridge into the pipeline).
Step 1: Define the Output Schema with Pydantic
Create a Pydantic model that describes the structured output you expect from the language model:
import pydantic
class SentimentEstimate(pydantic.BaseModel):
reasoning: str
score: pydantic.confloat(ge=0, le=1)
Step 2: Implement the Custom Bridge (sieves/tasks/predictive/bridges.py)
Subclass Bridge with your schema types and implement the required abstract properties:
from sieves.tasks.predictive.bridges import Bridge
from sieves.model_wrappers import ModelType, outlines_
from functools import cached_property
class OutlinesSentimentAnalysis(
Bridge[SentimentEstimate, SentimentEstimate, outlines_.InferenceMode]
):
@property
def model_type(self) -> ModelType:
return ModelType.outlines
@property
def _default_prompt_instructions(self) -> str:
return (
"Estimate the sentiment in this text as a float between 0 and 1. "
"Provide your reasoning before the score."
)
@property
def _prompt_conclusion(self) -> str | None:
return "========\nText: {{ text }}\nOutput:"
@property
def inference_mode(self) -> outlines_.InferenceMode:
return self._model_settings.inference_mode or outlines_.InferenceMode.json
@cached_property
def prompt_signature(self) -> type[pydantic.BaseModel]:
return SentimentEstimate
def integrate(self, results, docs):
for doc, result in zip(docs, results):
doc.results[self._task_id] = result.score
return docs
def consolidate(self, results, docs_offsets):
consolidated = []
for start, end in docs_offsets:
chunk_scores = [r.score for r in results[start:end] if r]
chunk_reasonings = [r.reasoning for r in results[start:end] if r]
consolidated.append(
SentimentEstimate(
score=sum(chunk_scores) / len(chunk_scores),
reasoning=" ".join(chunk_reasonings),
)
)
return consolidated
The integrate method writes results back to documents, while consolidate handles chunk-aware pipelines by aggregating multiple chunk results into single document-level results.
Step 3: Create the PredictiveTask Subclass (sieves/tasks/predictive/core.py)
Wire your bridge into a task by subclassing PredictiveTask:
from sieves.tasks.predictive.core import PredictiveTask
from sieves.model_wrappers import ModelType
from typing import Sequence, Any
import datasets
class SentimentAnalysis(
PredictiveTask[SentimentEstimate, SentimentEstimate, OutlinesSentimentAnalysis]
):
@property
def metric(self) -> str:
return "MSE"
@property
def prompt_signature(self) -> type[pydantic.BaseModel]:
return SentimentEstimate
def _init_bridge(self, model_type: ModelType) -> OutlinesSentimentAnalysis:
if model_type == ModelType.outlines:
return OutlinesSentimentAnalysis(
task_id=self._task_id,
prompt_instructions=self._custom_prompt_instructions,
overwrite=False,
model_settings=self._model_settings,
prompt_signature=self.prompt_signature,
model_type=model_type,
fewshot_examples=self._fewshot_examples,
)
raise KeyError(f"Unsupported model type {model_type}")
@property
def supports(self) -> set[ModelType]:
return {ModelType.outlines}
def to_hf_dataset(self, docs: Iterable[Doc]) -> datasets.Dataset:
info = datasets.DatasetInfo(
description=f"Sentiment dataset generated by Sieves v{Config.get_version()}",
features=datasets.Features({"text": datasets.Value("string"), "score": datasets.Value("float32")}),
)
def gen():
for doc in docs:
yield {"text": doc.text, "score": doc.results[self._task_id]}
return datasets.Dataset.from_generator(gen, features=info.features, info=info)
The _init_bridge method acts as a factory that instantiates your bridge with the appropriate configuration. The supports property declares which model wrappers this task can use.
Step 4: Use in a Pipeline
Instantiate your custom predictive task and add it to a Pipeline:
from sieves.pipeline import Pipeline
from sieves.data import Doc
from sieves.model_wrappers import ModelWrapperInferenceMode, outlines_
pipeline = Pipeline(
tasks=[
SentimentAnalysis(
task_id="Sentiment",
include_meta=False,
batch_size=-1,
model_type=ModelType.outlines,
model_settings=outlines_.ModelSettings(
inference_mode=ModelWrapperInferenceMode.json
),
)
],
use_cache=False,
)
docs = [Doc(text="I love this product!"), Doc(text="Terrible experience.")]
results = list(pipeline(docs))
print(results[0].results["Sentiment"]) # → score between 0-1
Because SentimentAnalysis inherits from PredictiveTask, all heavy lifting—including prompt creation, model execution, and chunk consolidation—is handled automatically by the bridge implementation.
Summary
-
Custom deterministic tasks require subclassing
Taskfromsieves/tasks/core.pyand implementing the_callmethod to processIterable[Doc]and yield processed documents. -
Custom predictive tasks require subclassing
PredictiveTaskfromsieves/tasks/predictive/core.pyand implementingprompt_signature,metric,_init_bridge, andsupportsto wire in your custom bridge. -
Custom bridges require subclassing
Bridgefromsieves/tasks/predictive/bridges.pyand implementing_default_prompt_instructions,inference_mode,integrate, and optionallyconsolidateto handle model-specific prompt building and result integration. -
All components support the
+operator for pipeline chaining and automatically inherit batching, conditional execution, and serialization from their base classes.
Frequently Asked Questions
What is the difference between a Task and a PredictiveTask in Sieves?
A Task is the base class for any deterministic operation that transforms documents, requiring only the implementation of the _call method to process Iterable[Doc]. A PredictiveTask extends this for language model operations, adding generic type parameters for prompt signatures, results, and bridges, plus abstract methods like metric and _init_bridge to handle model-specific inference and evaluation.
How do I handle chunked documents in a custom Bridge?
Implement the consolidate method in your Bridge subclass to aggregate results from multiple chunks into a single document-level result. This method receives results (all chunk outputs) and docs_offsets (tuples indicating which chunks belong to which document), allowing you to compute averages, concatenate reasoning, or apply any other aggregation logic before returning a consolidated result per document.
Can I use multiple model types with the same custom PredictiveTask?
Yes, by implementing _init_bridge as a factory that returns different bridge instances based on the model_type parameter. Your supports property should return a set containing all compatible ModelType values (e.g., {ModelType.outlines, ModelType.dspy}), and _init_bridge should instantiate the appropriate bridge class for each model type, raising KeyError for unsupported types.
Where should I store custom task implementations in my project?
You can define custom tasks and bridges in any Python module within your project, as Sieves uses standard Python import paths. For organization, place deterministic tasks in a tasks/ directory and predictive tasks with their bridges in tasks/predictive/, mirroring the structure in sieves/tests/docs/test_custom_tasks.py where the library demonstrates end-to-end custom task implementations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →