How to Serialize and Deserialize Sieves Pipelines with YAML Configuration
Use pipeline.dump() to serialize a Sieves pipeline to YAML and Pipeline.load() with task_kwargs to reconstruct it later, enabling version-safe persistence across sessions and machines.
The mantisai/sieves library provides a robust mechanism for saving and reloading machine learning pipelines through YAML-based serialization. This capability allows you to serialize and deserialize Sieves pipelines with YAML configuration, preserving task configurations while allowing runtime injection of non-serializable objects like model instances.
Understanding the Serialization Architecture
The persistence system relies on three core components working together to ensure reliable round-trip serialization.
The Config Class
Located in sieves/serialization.py, the Config class is a Pydantic model that stores the pipeline's structural metadata. It records the pipeline class name, version information, the use_cache flag, and serialized representations of all tasks. The Config.dump() method handles the actual YAML file writing using PyYAML, while Config.load() validates version compatibility through Config.validate_init_params before reconstruction.
The Attribute Wrapper
The Attribute class (also in sieves/serialization.py) wraps individual values within tasks. When a value cannot be serialized (such as a live model object), the system marks it as a placeholder and stores a string reference. This design keeps YAML files clean and human-readable while allowing you to inject real objects during deserialization via the task_kwargs parameter.
Pipeline Orchestration
The Pipeline class in sieves/pipeline/core.py orchestrates the entire process. It implements serialize() to generate the Config object, dump() to write YAML files, and the class method load() to read and reconstruct pipelines from disk.
Serializing a Pipeline to YAML
Serialization converts a live pipeline into a portable YAML configuration file.
Using the serialize() Method
The Pipeline.serialize() method (lines 75-86 in sieves/pipeline/core.py) creates a Config object containing:
- The pipeline class identifier
- The
use_cacheboolean flag - Serialized task configurations via
task.serialize()for each task in the pipeline
Writing to Disk with dump()
To save the configuration, call Pipeline.dump(path) (lines 53-60 in sieves/pipeline/core.py). This method internally calls self.serialize().dump(path), which uses PyYAML to write the JSON-compatible data to a .yml file.
from pathlib import Path
from sieves import Pipeline
from sieves.tasks import Classification
# Build a pipeline
clf_task = Classification(
name="sentiment",
model="gpt-4o-mini",
labels=["pos", "neg"]
)
pipe = Pipeline(tasks=[clf_task], use_cache=True)
# Serialize to YAML
config_path = Path("my_pipeline.yml")
pipe.dump(config_path)
Deserializing a Pipeline from YAML
Loading reconstructs the pipeline from the YAML file while allowing injection of runtime-specific objects.
Loading with Pipeline.load()
The class method Pipeline.load(path, task_kwargs) (lines 65-73 in sieves/pipeline/core.py) performs three operations:
- Reads the YAML file using
Config.load(path)(lines 78-86 insieves/serialization.py) - Validates version and class name compatibility via
Config.validate_init_params(lines 53-62 insieves/serialization.py) - Reconstructs tasks by calling
Task.deserializewith the providedtask_kwargs
Providing task_kwargs for Placeholders
The task_kwargs parameter is a list of dictionaries, with one dictionary per task in the pipeline. Each dictionary supplies values for placeholder attributes that could not be serialized to YAML (such as live model instances).
from sieves import Pipeline
from pathlib import Path
config_path = Path("my_pipeline.yml")
# Provide runtime objects for placeholders
task_kwargs = [
{"model": "gpt-4o-mini"} # Matches the first task's placeholder
]
# Load the pipeline
restored_pipe = Pipeline.load(config_path, task_kwargs)
The Deserialization Process
Inside Pipeline.deserialize (lines 88-104 in sieves/pipeline/core.py), the method iterates through the serialized task configurations, instantiates each task class, and calls task_cls.deserialize(task_config, **task_kwargs) to reconstruct the task with its runtime dependencies.
Complete Working Example
This example demonstrates the full round-trip from pipeline creation through serialization to deserialization and execution.
from pathlib import Path
from sieves import Pipeline, Doc
from sieves.tasks import Classification
# Step 1: Build the pipeline
clf_task = Classification(
name="sentiment",
model="gpt-4o-mini",
labels=["positive", "negative"],
)
pipe = Pipeline(tasks=[clf_task], use_cache=True)
# Step 2: Serialize to YAML
config_path = Path("sentiment_pipeline.yml")
pipe.dump(config_path)
# Step 3: Later, or on a different machine, load the pipeline
# Provide any placeholder values that couldn't be serialized
task_kwargs = [
{"model": "gpt-4o-mini"}
]
restored_pipe = Pipeline.load(config_path, task_kwargs)
# Step 4: Execute the restored pipeline
docs = [
Doc(text="I love this product!"),
Doc(text="This was a terrible experience.")
]
results = list(restored_pipe(docs))
print(results[0].results["sentiment"]) # Output: positive
print(results[1].results["sentiment"]) # Output: negative
The YAML file generated by pipe.dump() contains the pipeline structure, task configurations, and placeholder references, but excludes the actual model object, which is injected during loading via task_kwargs.
Summary
- Serialization: Call
pipeline.dump(path)to write a YAML configuration file containing the pipeline structure and task settings. - Deserialization: Use
Pipeline.load(path, task_kwargs)to reconstruct the pipeline, providing runtime objects for any placeholders. - Core Components: The
Configclass insieves/serialization.pyhandles YAML I/O, whilePipelineinsieves/pipeline/core.pyorchestrates the process. - Version Safety: The system validates library versions and class names during loading to prevent compatibility issues.
- Placeholder System: Non-serializable objects are stored as string references and injected via the
task_kwargsparameter during deserialization.
Frequently Asked Questions
What file format does Sieves use for pipeline serialization?
Sieves uses YAML files for pipeline serialization. The Pipeline.dump() method writes to .yml files using PyYAML, producing human-readable configuration files that store pipeline structure, task parameters, and placeholder references for non-serializable objects.
How do I handle model objects that cannot be serialized to YAML?
Use the task_kwargs parameter in Pipeline.load(). Pass a list of dictionaries where each dictionary corresponds to a task in your pipeline and contains the runtime objects (like model instances) that replace placeholder values stored in the YAML file. This design keeps configuration files clean while allowing injection of complex objects during reconstruction.
Does Sieves validate compatibility when loading a serialized pipeline?
Yes. The Config.validate_init_params method in sieves/serialization.py checks that the library version in the YAML file matches the current installed version and verifies that the class name matches the expected pipeline class. These validations prevent errors from version mismatches or incompatible configuration formats.
Can I serialize a pipeline that uses caching?
Yes. The use_cache flag is preserved during serialization. When you call pipeline.dump(), the Config object stores the use_cache boolean value, and Pipeline.load() restores it when reconstructing the pipeline. This ensures that caching behavior persists across serialization round-trips.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →