How to Compose Pipelines Using the + Operator in Sieves: A Complete Guide
Use the + operator to functionally combine Pipeline objects or append Task instances, creating a new pipeline that sequentially executes all tasks while preserving the left operand's cache settings.
The Sieves library (mantisai/sieves) provides an intuitive way to build modular NLP workflows. When you need to chain multiple processing steps, the library overloads the binary + operator to make pipeline composition readable and functional. This approach allows you to assemble complex document processing workflows from smaller, reusable components.
Understanding Pipeline Composition in Sieves
In Sieves, a pipeline is the orchestrator that sequentially runs a list of Task objects on one or more Doc instances. The + operator provides a Pythonic way to concatenate these pipelines without mutating the original objects.
What Happens When You Use the + Operator
When you write combined = pipeline_a + pipeline_b or combined = pipeline_a + some_task, the Pipeline.__add__ method defined in [sieves/pipeline/core.py](https://github.com/mantisai/sieves/blob/main/sieves/pipeline/core.py) (lines 239-262) is invoked.
The method executes the following logic:
- Type check: Determines whether the right-hand side (
other) is aPipelineor a singleTask - Task concatenation:
- If
otheris aPipeline, builds[*self._tasks, *other._tasks] - If
otheris aTask, builds[*self._tasks, other]
- If
- Cache handling: The new pipeline inherits the left pipeline's
use_cacheflag, ensuring consistent caching semantics - Validation: Returns a brand-new
Pipelineinstance; original pipelines remain unchanged (functional, not in-place)
If an unsupported type is supplied, the method raises TypeError(f"Cannot chain Pipeline with {type(other).__name__}").
The In-Place Alternative with +=
For scenarios where you want to modify an existing pipeline rather than create a new one, Sieves provides the += operator via Pipeline.__iadd__ (lines 262-282 in sieves/pipeline/core.py).
This method mutates the left-hand pipeline by appending new task(s) and then re-validating task IDs to ensure uniqueness across the combined list.
Technical Implementation Details
Task Concatenation Logic
The concatenation mechanism uses Python's unpacking syntax to merge task lists efficiently:
# When combining two pipelines
new_tasks = [*self._tasks, *other._tasks]
# When appending a single task
new_tasks = [*self._tasks, other]
This approach preserves the execution order: tasks from the left operand execute first, followed by tasks from the right operand.
Cache Propagation Behavior
A critical architectural decision in Sieves is that cache settings propagate from the left operand. When you compose pipelines using +, the resulting pipeline uses the use_cache value from the left-hand side:
# combined will have use_cache=True (from pipe_cached)
combined = pipe_cached + pipe_uncached
# combined will have use_cache=False (from pipe_uncached)
combined = pipe_uncached + pipe_cached
This prevents accidental disabling of caching when chaining pipelines with different cache configurations.
Task ID Validation
After any composition operation (both + and +=), Sieves calls _validate_tasks() to guarantee that each task ID remains unique across the combined list. This prevents naming collisions that could cause errors during pipeline execution.
Practical Code Examples
Combining Two Pipelines
Create separate pipelines for different processing stages and combine them into a unified workflow:
from sieves import Pipeline, Doc
from sieves.tasks import ClassificationTask, NERTask
# Create two simple pipelines
pipe_text = Pipeline([ClassificationTask(model="gpt-4o-mini", id="cls")])
pipe_entities = Pipeline([NERTask(model="gpt-4o-mini", id="ner")])
# Combine them with the + operator (functional composition)
combined_pipe = pipe_text + pipe_entities
# The combined pipeline runs classification first, then NER
docs = [Doc(text="Alice lives in Paris.")]
for processed in combined_pipe(docs):
print(processed.results) # {'cls': ..., 'ner': ...}
Adding a Single Task to a Pipeline
You can also append individual tasks directly without wrapping them in a separate pipeline:
# Mixing a single task with a pipeline
single_task = ClassificationTask(model="gpt-4o-mini", id="sentiment")
pipeline = single_task + pipe_entities # returns a new Pipeline
Note that the + operator is commutative in practice—you can add a task to a pipeline or a pipeline to a task, as both implement the necessary logic to return a new Pipeline instance.
In-Place Composition with +=
When you want to extend an existing pipeline without creating intermediate objects, use the in-place addition operator:
# In-place concatenation with +=
pipe_text += NERTask(model="gpt-4o-mini", id="ner")
# Now pipe_text itself contains both tasks
This modifies pipe_text directly, appending the new task and re-validating task IDs to ensure no conflicts exist.
Summary
- The
+operator creates newPipelineinstances without mutating the original pipelines, enabling safe reuse of components across different workflows. - Cache settings propagate from the left operand, ensuring consistent caching behavior regardless of what is appended.
- Task IDs are validated after composition to prevent naming collisions that could cause runtime errors.
- Both
Pipeline-to-PipelineandPipeline-to-Taskcomposition are supported, with the+operator handling type checking and appropriate concatenation logic. - For in-place modification, use
+=which mutates the left-hand pipeline and re-validates task IDs.
Frequently Asked Questions
Can I compose more than two pipelines at once?
Yes, you can chain multiple compositions left-to-right. The expression pipeline_a + pipeline_b + pipeline_c works because each + operation returns a new Pipeline instance that can be immediately combined with the next operand. According to the implementation in sieves/pipeline/core.py, the intermediate pipelines maintain proper task ordering and cache settings from the leftmost operand.
What happens if task IDs conflict during composition?
The pipeline runs _validate_tasks() after any composition operation (both + and +=). If duplicate task IDs are detected across the combined task list, the validation raises an error preventing the creation of an invalid pipeline. This ensures that each task remains uniquely identifiable for result tracking and debugging purposes.
Does using + modify the original pipelines?
No, the + operator is functional and immutable. When you write combined = pipe_a + pipe_b, the Pipeline.__add__ method (lines 239-262 in sieves/pipeline/core.py) constructs a brand-new Pipeline instance with concatenated tasks. Both pipe_a and pipe_b remain unchanged and can be reused in other workflows. For in-place modification, use the += operator instead.
Where is the + operator implemented in the Sieves source code?
The binary + operator is implemented in the Pipeline.__add__ method located in [sieves/pipeline/core.py](https://github.com/mantisai/sieves/blob/main/sieves/pipeline/core.py) at lines 239-262. The in-place += operator is implemented in __iadd__ at lines 262-282 in the same file. Both methods handle type checking, task concatenation, cache propagation, and task ID validation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →