How to Create Custom Tasks and Bridges in Sieves Pipelines: A Complete Guide

To create custom tasks and bridges in Sieves pipelines, subclass Task for deterministic operations or PredictiveTask with a custom Bridge for LLM-powered tasks, then implement the required abstract methods like _call, integrate, and prompt_signature.

The mantisai/sieves library provides a modular framework for composing NLP pipelines through reusable components. When you need to create custom tasks and bridges in Sieves pipelines to handle domain-specific logic or integrate custom language models, the framework offers lightweight, fully typed base classes that handle batching, caching, and serialization automatically.

Understanding the Core Architecture of Sieves Tasks

Before implementing custom components, you need to understand the three fundamental abstractions in sieves/tasks/core.py, sieves/tasks/predictive/core.py, and sieves/tasks/predictive/bridges.py.

The Task Base Class (sieves/tasks/core.py)

The Task class in sieves/tasks/core.py is the foundation for all pipeline operations. It handles batching, conditional execution, and result plumbing through these key mechanisms:

  • __init__(task_id, include_meta, batch_size, condition) – Configures the task ID, metadata handling, batch size, and an optional per-document filter function.
  • __call__(docs) – Iterates over documents, respects the condition, batches passing documents, and forwards them to _call. Documents failing the condition receive None as their result.
  • _call(self, docs) – Abstract method that subclasses must implement; receives only the documents that passed the condition.
  • serialize / deserialize – Enable pipeline persistence and caching.

The PredictiveTask for LLM Operations (sieves/tasks/predictive/core.py)

PredictiveTask extends Task for operations that call language models through a Bridge. It is generic over three types:

PredictiveTask[PromptSignature, Result, BridgeClass]
  • PromptSignature – The Pydantic model describing the model's output structure.
  • Result – The type stored in doc.results.
  • BridgeClass – The concrete Bridge implementation for a specific model wrapper (Outlines, DSPy, LangChain).

Key responsibilities include:

  • Delegating prompt building, extraction, integration, and consolidation to its bridge.
  • Supplying helpers for few-shot examples and dataset export via to_hf_dataset.
  • Declaring supported model types through the supports property.

The Bridge Abstraction (sieves/tasks/predictive/bridges.py)

Bridge in sieves/tasks/predictive/bridges.py is the model-agnostic contract that connects predictive tasks to concrete model wrappers. Constructor arguments include the task ID, optional custom prompt, model settings, Pydantic signature, model type, and few-shot examples.

Core abstract properties to implement:

  • _default_prompt_instructions – Fallback instructions when the user does not provide custom ones.
  • inference_mode – The ModelWrapperInferenceMode (e.g., json, choice, regex).
  • integrate(self, results, docs) – Writes the model's output back into each Doc (typically into doc.results).
  • consolidate(self, results, docs_offsets) – Combines chunk-level results into document-level results (required for chunk-aware pipelines).

Helper properties like _prompt_instructions, _prompt_example_xml, and prompt_template automatically construct ready-to-use prompt strings from defaults, few-shot examples, and conclusions.

How to Create a Custom Task in Sieves

For deterministic operations that don't require language models, subclass Task and implement the _call method. This example from sieves/tests/docs/test_custom_tasks.py demonstrates a character counting task:

from sieves.tasks.core import Task
from sieves.data import Doc
from typing import Iterable

class CharCountTask(Task):
    """Counts characters in doc.text."""
    def _call(self, docs: Iterable[Doc]) -> Iterable[Doc]:
        for doc in docs:
            doc.results[self.id] = len(doc.text or "")
            yield doc

The task automatically inherits batching and conditional execution from Task.__call__. You can instantiate it with standard arguments:

task = CharCountTask(task_id="CharCount", include_meta=False, batch_size=-1)
docs = [Doc(text="Hello"), Doc(text="World!")]
list(task(docs))  # Docs with results["CharCount"] = 5, 6

Use the condition argument to skip documents based on custom logic:

task = CharCountTask(condition=lambda d: len(d.text) > 0)

How to Create a Custom Predictive Task and Bridge

For LLM-powered operations, you must implement both a Bridge (handling model-specific prompt building and result integration) and a PredictiveTask (wiring the bridge into the pipeline).

Step 1: Define the Output Schema with Pydantic

Create a Pydantic model that describes the structured output you expect from the language model:

import pydantic

class SentimentEstimate(pydantic.BaseModel):
    reasoning: str
    score: pydantic.confloat(ge=0, le=1)

Step 2: Implement the Custom Bridge (sieves/tasks/predictive/bridges.py)

Subclass Bridge with your schema types and implement the required abstract properties:

from sieves.tasks.predictive.bridges import Bridge
from sieves.model_wrappers import ModelType, outlines_
from functools import cached_property

class OutlinesSentimentAnalysis(
    Bridge[SentimentEstimate, SentimentEstimate, outlines_.InferenceMode]
):
    @property
    def model_type(self) -> ModelType:
        return ModelType.outlines

    @property
    def _default_prompt_instructions(self) -> str:
        return (
            "Estimate the sentiment in this text as a float between 0 and 1. "
            "Provide your reasoning before the score."
        )

    @property
    def _prompt_conclusion(self) -> str | None:
        return "========\nText: {{ text }}\nOutput:"

    @property
    def inference_mode(self) -> outlines_.InferenceMode:
        return self._model_settings.inference_mode or outlines_.InferenceMode.json

    @cached_property
    def prompt_signature(self) -> type[pydantic.BaseModel]:
        return SentimentEstimate

    def integrate(self, results, docs):
        for doc, result in zip(docs, results):
            doc.results[self._task_id] = result.score
        return docs

    def consolidate(self, results, docs_offsets):
        consolidated = []
        for start, end in docs_offsets:
            chunk_scores = [r.score for r in results[start:end] if r]
            chunk_reasonings = [r.reasoning for r in results[start:end] if r]
            consolidated.append(
                SentimentEstimate(
                    score=sum(chunk_scores) / len(chunk_scores),
                    reasoning=" ".join(chunk_reasonings),
                )
            )
        return consolidated

The integrate method writes results back to documents, while consolidate handles chunk-aware pipelines by aggregating multiple chunk results into single document-level results.

Step 3: Create the PredictiveTask Subclass (sieves/tasks/predictive/core.py)

Wire your bridge into a task by subclassing PredictiveTask:

from sieves.tasks.predictive.core import PredictiveTask
from sieves.model_wrappers import ModelType
from typing import Sequence, Any
import datasets

class SentimentAnalysis(
    PredictiveTask[SentimentEstimate, SentimentEstimate, OutlinesSentimentAnalysis]
):
    @property
    def metric(self) -> str:
        return "MSE"

    @property
    def prompt_signature(self) -> type[pydantic.BaseModel]:
        return SentimentEstimate

    def _init_bridge(self, model_type: ModelType) -> OutlinesSentimentAnalysis:
        if model_type == ModelType.outlines:
            return OutlinesSentimentAnalysis(
                task_id=self._task_id,
                prompt_instructions=self._custom_prompt_instructions,
                overwrite=False,
                model_settings=self._model_settings,
                prompt_signature=self.prompt_signature,
                model_type=model_type,
                fewshot_examples=self._fewshot_examples,
            )
        raise KeyError(f"Unsupported model type {model_type}")

    @property
    def supports(self) -> set[ModelType]:
        return {ModelType.outlines}

    def to_hf_dataset(self, docs: Iterable[Doc]) -> datasets.Dataset:
        info = datasets.DatasetInfo(
            description=f"Sentiment dataset generated by Sieves v{Config.get_version()}",
            features=datasets.Features({"text": datasets.Value("string"), "score": datasets.Value("float32")}),
        )
        def gen():
            for doc in docs:
                yield {"text": doc.text, "score": doc.results[self._task_id]}
        return datasets.Dataset.from_generator(gen, features=info.features, info=info)

The _init_bridge method acts as a factory that instantiates your bridge with the appropriate configuration. The supports property declares which model wrappers this task can use.

Step 4: Use in a Pipeline

Instantiate your custom predictive task and add it to a Pipeline:

from sieves.pipeline import Pipeline
from sieves.data import Doc
from sieves.model_wrappers import ModelWrapperInferenceMode, outlines_

pipeline = Pipeline(
    tasks=[
        SentimentAnalysis(
            task_id="Sentiment",
            include_meta=False,
            batch_size=-1,
            model_type=ModelType.outlines,
            model_settings=outlines_.ModelSettings(
                inference_mode=ModelWrapperInferenceMode.json
            ),
        )
    ],
    use_cache=False,
)

docs = [Doc(text="I love this product!"), Doc(text="Terrible experience.")]
results = list(pipeline(docs))
print(results[0].results["Sentiment"])  # → score between 0-1

Because SentimentAnalysis inherits from PredictiveTask, all heavy lifting—including prompt creation, model execution, and chunk consolidation—is handled automatically by the bridge implementation.

Summary

  • Custom deterministic tasks require subclassing Task from sieves/tasks/core.py and implementing the _call method to process Iterable[Doc] and yield processed documents.

  • Custom predictive tasks require subclassing PredictiveTask from sieves/tasks/predictive/core.py and implementing prompt_signature, metric, _init_bridge, and supports to wire in your custom bridge.

  • Custom bridges require subclassing Bridge from sieves/tasks/predictive/bridges.py and implementing _default_prompt_instructions, inference_mode, integrate, and optionally consolidate to handle model-specific prompt building and result integration.

  • All components support the + operator for pipeline chaining and automatically inherit batching, conditional execution, and serialization from their base classes.

Frequently Asked Questions

What is the difference between a Task and a PredictiveTask in Sieves?

A Task is the base class for any deterministic operation that transforms documents, requiring only the implementation of the _call method to process Iterable[Doc]. A PredictiveTask extends this for language model operations, adding generic type parameters for prompt signatures, results, and bridges, plus abstract methods like metric and _init_bridge to handle model-specific inference and evaluation.

How do I handle chunked documents in a custom Bridge?

Implement the consolidate method in your Bridge subclass to aggregate results from multiple chunks into a single document-level result. This method receives results (all chunk outputs) and docs_offsets (tuples indicating which chunks belong to which document), allowing you to compute averages, concatenate reasoning, or apply any other aggregation logic before returning a consolidated result per document.

Can I use multiple model types with the same custom PredictiveTask?

Yes, by implementing _init_bridge as a factory that returns different bridge instances based on the model_type parameter. Your supports property should return a set containing all compatible ModelType values (e.g., {ModelType.outlines, ModelType.dspy}), and _init_bridge should instantiate the appropriate bridge class for each model type, raising KeyError for unsupported types.

Where should I store custom task implementations in my project?

You can define custom tasks and bridges in any Python module within your project, as Sieves uses standard Python import paths. For organization, place deterministic tasks in a tasks/ directory and predictive tasks with their bridges in tasks/predictive/, mirroring the structure in sieves/tests/docs/test_custom_tasks.py where the library demonstrates end-to-end custom task implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →