# How to Create Custom Tasks and Bridges in Sieves Pipelines: A Complete Guide

> Learn to create custom Sieves tasks and bridges for deterministic or LLM-powered pipelines. This guide covers subclassing Task and PredictiveTask, implementing abstract methods like _call, integrate, and prompt_signature.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**To create custom tasks and bridges in Sieves pipelines, subclass `Task` for deterministic operations or `PredictiveTask` with a custom `Bridge` for LLM-powered tasks, then implement the required abstract methods like `_call`, `integrate`, and `prompt_signature`.**

The `mantisai/sieves` library provides a modular framework for composing NLP pipelines through reusable components. When you need to create custom tasks and bridges in Sieves pipelines to handle domain-specific logic or integrate custom language models, the framework offers lightweight, fully typed base classes that handle batching, caching, and serialization automatically.

## Understanding the Core Architecture of Sieves Tasks

Before implementing custom components, you need to understand the three fundamental abstractions in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py), [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py), and [`sieves/tasks/predictive/bridges.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/bridges.py).

### The Task Base Class (sieves/tasks/core.py)

The `Task` class in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) is the foundation for all pipeline operations. It handles **batching**, **conditional execution**, and **result plumbing** through these key mechanisms:

- `__init__(task_id, include_meta, batch_size, condition)` – Configures the task ID, metadata handling, batch size, and an optional per-document filter function.
- `__call__(docs)` – Iterates over documents, respects the `condition`, batches passing documents, and forwards them to `_call`. Documents failing the condition receive `None` as their result.
- `_call(self, docs)` – **Abstract method** that subclasses must implement; receives only the documents that passed the condition.
- `serialize / deserialize` – Enable pipeline persistence and caching.

### The PredictiveTask for LLM Operations (sieves/tasks/predictive/core.py)

`PredictiveTask` extends `Task` for operations that call language models through a **Bridge**. It is generic over three types:

```python
PredictiveTask[PromptSignature, Result, BridgeClass]

```

- `PromptSignature` – The Pydantic model describing the model's output structure.
- `Result` – The type stored in `doc.results`.
- `BridgeClass` – The concrete `Bridge` implementation for a specific model wrapper (Outlines, DSPy, LangChain).

Key responsibilities include:
- Delegating prompt building, extraction, integration, and consolidation to its bridge.
- Supplying helpers for few-shot examples and dataset export via `to_hf_dataset`.
- Declaring supported model types through the `supports` property.

### The Bridge Abstraction (sieves/tasks/predictive/bridges.py)

`Bridge` in [`sieves/tasks/predictive/bridges.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/bridges.py) is the **model-agnostic contract** that connects predictive tasks to concrete model wrappers. Constructor arguments include the task ID, optional custom prompt, model settings, Pydantic signature, model type, and few-shot examples.

Core abstract properties to implement:

- `_default_prompt_instructions` – Fallback instructions when the user does not provide custom ones.
- `inference_mode` – The `ModelWrapperInferenceMode` (e.g., `json`, `choice`, `regex`).
- `integrate(self, results, docs)` – Writes the model's output back into each `Doc` (typically into `doc.results`).
- `consolidate(self, results, docs_offsets)` – Combines chunk-level results into document-level results (required for chunk-aware pipelines).

Helper properties like `_prompt_instructions`, `_prompt_example_xml`, and `prompt_template` automatically construct ready-to-use prompt strings from defaults, few-shot examples, and conclusions.

## How to Create a Custom Task in Sieves

For deterministic operations that don't require language models, subclass `Task` and implement the `_call` method. This example from [`sieves/tests/docs/test_custom_tasks.py`](https://github.com/mantisai/sieves/blob/main/sieves/tests/docs/test_custom_tasks.py) demonstrates a character counting task:

```python
from sieves.tasks.core import Task
from sieves.data import Doc
from typing import Iterable

class CharCountTask(Task):
    """Counts characters in doc.text."""
    def _call(self, docs: Iterable[Doc]) -> Iterable[Doc]:
        for doc in docs:
            doc.results[self.id] = len(doc.text or "")
            yield doc

```

The task automatically inherits batching and conditional execution from `Task.__call__`. You can instantiate it with standard arguments:

```python
task = CharCountTask(task_id="CharCount", include_meta=False, batch_size=-1)
docs = [Doc(text="Hello"), Doc(text="World!")]
list(task(docs))  # Docs with results["CharCount"] = 5, 6

```

Use the `condition` argument to skip documents based on custom logic:

```python
task = CharCountTask(condition=lambda d: len(d.text) > 0)

```

## How to Create a Custom Predictive Task and Bridge

For LLM-powered operations, you must implement both a `Bridge` (handling model-specific prompt building and result integration) and a `PredictiveTask` (wiring the bridge into the pipeline).

### Step 1: Define the Output Schema with Pydantic

Create a Pydantic model that describes the structured output you expect from the language model:

```python
import pydantic

class SentimentEstimate(pydantic.BaseModel):
    reasoning: str
    score: pydantic.confloat(ge=0, le=1)

```

### Step 2: Implement the Custom Bridge (sieves/tasks/predictive/bridges.py)

Subclass `Bridge` with your schema types and implement the required abstract properties:

```python
from sieves.tasks.predictive.bridges import Bridge
from sieves.model_wrappers import ModelType, outlines_
from functools import cached_property

class OutlinesSentimentAnalysis(
    Bridge[SentimentEstimate, SentimentEstimate, outlines_.InferenceMode]
):
    @property
    def model_type(self) -> ModelType:
        return ModelType.outlines

    @property
    def _default_prompt_instructions(self) -> str:
        return (
            "Estimate the sentiment in this text as a float between 0 and 1. "
            "Provide your reasoning before the score."
        )

    @property
    def _prompt_conclusion(self) -> str | None:
        return "========\nText: {{ text }}\nOutput:"

    @property
    def inference_mode(self) -> outlines_.InferenceMode:
        return self._model_settings.inference_mode or outlines_.InferenceMode.json

    @cached_property
    def prompt_signature(self) -> type[pydantic.BaseModel]:
        return SentimentEstimate

    def integrate(self, results, docs):
        for doc, result in zip(docs, results):
            doc.results[self._task_id] = result.score
        return docs

    def consolidate(self, results, docs_offsets):
        consolidated = []
        for start, end in docs_offsets:
            chunk_scores = [r.score for r in results[start:end] if r]
            chunk_reasonings = [r.reasoning for r in results[start:end] if r]
            consolidated.append(
                SentimentEstimate(
                    score=sum(chunk_scores) / len(chunk_scores),
                    reasoning=" ".join(chunk_reasonings),
                )
            )
        return consolidated

```

The `integrate` method writes results back to documents, while `consolidate` handles chunk-aware pipelines by aggregating multiple chunk results into single document-level results.

### Step 3: Create the PredictiveTask Subclass (sieves/tasks/predictive/core.py)

Wire your bridge into a task by subclassing `PredictiveTask`:

```python
from sieves.tasks.predictive.core import PredictiveTask
from sieves.model_wrappers import ModelType
from typing import Sequence, Any
import datasets

class SentimentAnalysis(
    PredictiveTask[SentimentEstimate, SentimentEstimate, OutlinesSentimentAnalysis]
):
    @property
    def metric(self) -> str:
        return "MSE"

    @property
    def prompt_signature(self) -> type[pydantic.BaseModel]:
        return SentimentEstimate

    def _init_bridge(self, model_type: ModelType) -> OutlinesSentimentAnalysis:
        if model_type == ModelType.outlines:
            return OutlinesSentimentAnalysis(
                task_id=self._task_id,
                prompt_instructions=self._custom_prompt_instructions,
                overwrite=False,
                model_settings=self._model_settings,
                prompt_signature=self.prompt_signature,
                model_type=model_type,
                fewshot_examples=self._fewshot_examples,
            )
        raise KeyError(f"Unsupported model type {model_type}")

    @property
    def supports(self) -> set[ModelType]:
        return {ModelType.outlines}

    def to_hf_dataset(self, docs: Iterable[Doc]) -> datasets.Dataset:
        info = datasets.DatasetInfo(
            description=f"Sentiment dataset generated by Sieves v{Config.get_version()}",
            features=datasets.Features({"text": datasets.Value("string"), "score": datasets.Value("float32")}),
        )
        def gen():
            for doc in docs:
                yield {"text": doc.text, "score": doc.results[self._task_id]}
        return datasets.Dataset.from_generator(gen, features=info.features, info=info)

```

The `_init_bridge` method acts as a factory that instantiates your bridge with the appropriate configuration. The `supports` property declares which model wrappers this task can use.

### Step 4: Use in a Pipeline

Instantiate your custom predictive task and add it to a `Pipeline`:

```python
from sieves.pipeline import Pipeline
from sieves.data import Doc
from sieves.model_wrappers import ModelWrapperInferenceMode, outlines_

pipeline = Pipeline(
    tasks=[
        SentimentAnalysis(
            task_id="Sentiment",
            include_meta=False,
            batch_size=-1,
            model_type=ModelType.outlines,
            model_settings=outlines_.ModelSettings(
                inference_mode=ModelWrapperInferenceMode.json
            ),
        )
    ],
    use_cache=False,
)

docs = [Doc(text="I love this product!"), Doc(text="Terrible experience.")]
results = list(pipeline(docs))
print(results[0].results["Sentiment"])  # → score between 0-1

```

Because `SentimentAnalysis` inherits from `PredictiveTask`, all heavy lifting—including prompt creation, model execution, and chunk consolidation—is handled automatically by the bridge implementation.

## Summary

- **Custom deterministic tasks** require subclassing `Task` from [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) and implementing the `_call` method to process `Iterable[Doc]` and yield processed documents.

- **Custom predictive tasks** require subclassing `PredictiveTask` from [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) and implementing `prompt_signature`, `metric`, `_init_bridge`, and `supports` to wire in your custom bridge.

- **Custom bridges** require subclassing `Bridge` from [`sieves/tasks/predictive/bridges.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/bridges.py) and implementing `_default_prompt_instructions`, `inference_mode`, `integrate`, and optionally `consolidate` to handle model-specific prompt building and result integration.

- All components support the `+` operator for pipeline chaining and automatically inherit batching, conditional execution, and serialization from their base classes.

## Frequently Asked Questions

### What is the difference between a Task and a PredictiveTask in Sieves?

A `Task` is the base class for any deterministic operation that transforms documents, requiring only the implementation of the `_call` method to process `Iterable[Doc]`. A `PredictiveTask` extends this for language model operations, adding generic type parameters for prompt signatures, results, and bridges, plus abstract methods like `metric` and `_init_bridge` to handle model-specific inference and evaluation.

### How do I handle chunked documents in a custom Bridge?

Implement the `consolidate` method in your `Bridge` subclass to aggregate results from multiple chunks into a single document-level result. This method receives `results` (all chunk outputs) and `docs_offsets` (tuples indicating which chunks belong to which document), allowing you to compute averages, concatenate reasoning, or apply any other aggregation logic before returning a consolidated result per document.

### Can I use multiple model types with the same custom PredictiveTask?

Yes, by implementing `_init_bridge` as a factory that returns different bridge instances based on the `model_type` parameter. Your `supports` property should return a set containing all compatible `ModelType` values (e.g., `{ModelType.outlines, ModelType.dspy}`), and `_init_bridge` should instantiate the appropriate bridge class for each model type, raising `KeyError` for unsupported types.

### Where should I store custom task implementations in my project?

You can define custom tasks and bridges in any Python module within your project, as Sieves uses standard Python import paths. For organization, place deterministic tasks in a `tasks/` directory and predictive tasks with their bridges in `tasks/predictive/`, mirroring the structure in [`sieves/tests/docs/test_custom_tasks.py`](https://github.com/mantisai/sieves/blob/main/sieves/tests/docs/test_custom_tasks.py) where the library demonstrates end-to-end custom task implementations.