How Pipeline Caching Works in Sieves: A Complete Technical Guide
Pipeline caching in Sieves stores hashed document results to skip redundant processing, saving API calls and ensuring deterministic outputs across multiple runs.
Pipeline caching is a core performance feature in Sieves (mantisai/sieves) that eliminates redundant computation by tracking which documents have already been processed. When enabled, the pipeline maintains a per-instance hash map of document states, allowing subsequent runs to retrieve cached results instead of re-executing expensive tasks like LLM inference.
Understanding Pipeline Caching Architecture
The caching system in Sieves operates at the Pipeline class level, maintaining isolated state for each pipeline instance rather than using a global cache. This design ensures that cache boundaries respect pipeline composition and configuration boundaries.
The Cache Storage Mechanism
When you instantiate a Pipeline with use_cache=True (the default), Sieves initializes an internal hash map in the constructor at sieves/pipeline/core.py lines 26-31. The cache flag is stored in the private attribute _use_cache (line 36) and exposed through the read-only property use_cache (lines 58-63).
The actual deduplication logic resides in the private method _get_unseen_unique_docs at line 106, which scans incoming document iterables and filters out any Doc objects whose hash already exists in the cache.
Enabling and Disabling Pipeline Caching
You control caching behavior during Pipeline instantiation through the use_cache parameter.
Default Caching Behavior
By default, all pipelines cache results:
from sieves import Pipeline
from sieves.tasks.predictive import Classification
# Cache enabled by default
pipeline = Pipeline(
tasks=[Classification(labels=["spam", "ham"], model="gpt-4o-mini")]
)
Explicitly Disabling the Cache
For debugging or when you require fresh predictions after model updates, disable caching:
# Force fresh processing on every run
pipeline_no_cache = Pipeline(
tasks=[Classification(labels=["spam", "ham"], model="gpt-4o-mini")],
use_cache=False
)
How Pipeline Caching Works During Execution
The caching mechanism operates through a three-stage process implemented in sieves/pipeline/core.py around line 135.
Stage 1: Duplicate Detection
When you call a pipeline with a list of documents, the _get_unseen_unique_docs method (line 106) computes hashes for each Doc and compares them against the internal cache. Documents with matching hashes are flagged as duplicates and removed from the processing queue.
Stage 2: Skipping Cached Documents
Only the filtered, unseen documents proceed to the task chain. This conditional routing at line 135 ensures that cached documents bypass expensive operations like LLM API calls entirely, significantly reducing latency and cost on subsequent runs.
Stage 3: Cache Updates
After task execution completes, the pipeline stores the resulting Doc objects—including their new results entries—back into the hash map. This update ensures that future runs can retrieve the complete processed state without re-execution.
Pipeline Caching in Composition
When composing pipelines using the + operator, cache behavior inherits from the left-hand side pipeline.
Cache Inheritance Rules
The composition logic in sieves/tasks/core.py (line 127) and the Pipeline.__add__ implementation at sieves/pipeline/core.py lines 255-258 ensure that the resulting combined pipeline adopts the use_cache setting from the left operand.
# Left pipeline has caching disabled
left = Pipeline(
[Classification(labels=["spam"], model="gpt-4o-mini")],
use_cache=False
)
# Right pipeline has caching enabled
right = Pipeline(
[Classification(labels=["ham"], model="gpt-4o-mini")],
use_cache=True
)
# Combined inherits use_cache=False from left
combined = left + right
print(combined.use_cache) # Output: False
Why Pipeline Caching Matters
Implementing proper caching strategies in Sieves pipelines delivers three critical benefits for production workloads.
- Performance Optimization: Eliminates redundant API calls to expensive LLM services, reducing latency and operational costs when processing duplicate documents or re-running analytics.
- Deterministic Outputs: Guarantees identical results across multiple pipeline executions for the same input, isolating your application from non-deterministic model behaviors or temperature variations.
- Operational Control: Provides explicit mechanisms to bypass caching via
use_cache=Falsefor debugging scenarios, model A/B testing, or when fresh predictions are required after updating underlying models.
Summary
Pipeline caching in Sieves operates through a per-instance hash map that tracks processed documents to eliminate redundant computation. The system defaults to use_cache=True, filtering duplicate documents through the _get_unseen_unique_docs method before task execution and updating the cache with results afterward. When composing pipelines, the left-hand side's cache setting determines behavior for the combined workflow. Disabling caching via use_cache=False forces fresh processing for debugging or model updates.
Frequently Asked Questions
How do I completely disable pipeline caching in Sieves?
Set use_cache=False when instantiating your Pipeline. This forces the pipeline to process every document on every run, bypassing the _get_unseen_unique_docs deduplication logic entirely. Use this mode when debugging task behavior or when you need fresh predictions after updating your underlying models.
Is the pipeline cache shared across different Pipeline instances?
No, the cache is per-pipeline instance only. Each Pipeline object maintains its own private hash map in the _use_cache attribute. Creating two separate Pipeline instances means they have independent caches—even if configured identically, documents processed by one instance will appear as "unseen" to the other.
How does caching work when I compose multiple pipelines together?
When using the + operator to compose pipelines (e.g., combined = pipeline1 + pipeline2), the resulting pipeline inherits the use_cache setting from the left-hand side operand (pipeline1). This inheritance is handled in Pipeline.__add__ at lines 255-258 of sieves/pipeline/core.py, ensuring consistent caching behavior throughout composed workflows.
Does Sieves cache the actual model outputs or just document processing status?
Sieves caches the complete processed document state, including all results entries generated by tasks. After a task finishes, the updated Doc object—with its computed labels, scores, or annotations—is stored in the hash map. This means subsequent cache hits retrieve the full result set, not just a flag indicating the document was seen.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →