How to Evaluate Pipeline Performance Against Ground Truth Data in Sieves
To evaluate pipeline performance in Sieves, populate the Doc.gold attribute with ground-truth labels, run the pipeline to fill Doc.results, then call pipeline.evaluate(docs) to generate a PipelineEvaluationReport comparing predictions against references.
The open-source Sieves library (mantisai/sieves) provides a unified evaluation framework that compares model predictions against annotated ground truth at both the task and pipeline levels. By storing references in Doc.gold and predictions in Doc.results, you can benchmark classification, NER, and information extraction tasks using either deterministic metrics or LLM-based judges.
Understanding the Evaluation Architecture
Document Structure: Gold vs. Results
Every document in Sieves is represented by a Doc object that carries dual dictionaries for evaluation:
Doc.results– Stores structured outputs generated by the pipeline, keyed by task IDs (e.g.,"cls","ner","ie").Doc.gold– Holds the ground-truth labels or annotations supplied by the user, keyed by the same task IDs.
When you run a pipeline, tasks populate doc.results[task_id] with their predictions. During evaluation, the system compares these entries against doc.gold[task_id] to compute performance metrics.
Task-Level Evaluation
Each task class implements an evaluate method defined in sieves/tasks/core.py that performs the core comparison logic. The base Task class provides the signature:
def evaluate(self, docs: Iterable[Doc], judge: dspy.LM | None = None) -> TaskEvaluationReport
According to the source code at line 99, this method iterates through documents, compares doc.results[task_id] with doc.gold[task_id], and produces a TaskEvaluationReport containing task-specific metrics. Predictive tasks (classification, NER, information extraction) inherit from PredictiveTask in sieves/tasks/predictive/core.py and may enrich reports with additional fields such as confidence-score analysis or per-label breakdowns at line 287.
Pipeline-Level Aggregation
The Pipeline class orchestrates evaluation across all tasks. When you call pipeline.evaluate(docs, judge), it executes the task-level evaluate method for every task in the sequence, then collates individual TaskEvaluationReport objects into a single PipelineEvaluationReport. This aggregation logic is implemented in sieves/pipeline/core.py at line 284, providing a holistic view of performance across the entire workflow.
Step-by-Step Evaluation Workflow
- Prepare ground-truth data – Populate
doc.gold[task_id]for each document before running the pipeline, ensuring keys match your task IDs. - Run inference – Execute
pipeline(docs)to filldoc.resultswith model predictions. - Call evaluate – Invoke
report = pipeline.evaluate(docs, judge=my_judge)to generate the evaluation report. Omit thejudgeparameter for deterministic comparison, or provide adspy.LMinstance for LLM-based scoring. - Inspect results – Access
report.task_reports[task_id]to view specific metrics like accuracy, F1, or exact-match scores.
Code Examples
Classification with Deterministic Scoring
For classification tasks, use deterministic exact-match evaluation by omitting the judge parameter:
from sieves import Doc, Pipeline
from sieves.tasks.predictive.classification import ClassificationTask
# 1. Build gold-label documents
docs = [
Doc(text="I love this product!", gold={"cls": "positive"}),
Doc(text="Terrible experience.", gold={"cls": "negative"}),
]
# 2. Create a classification task
cls_task = ClassificationTask(
task_id="cls",
label_set=["positive", "negative"],
)
# 3. Assemble and run the pipeline
pipe = Pipeline([cls_task])
pipe(docs)
# 4. Evaluate against ground truth (deterministic)
report = pipe.evaluate(docs)
print(report.summary())
print(report.task_reports["cls"]) # Detailed classification metrics
Information Extraction with an LLM Judge
For open-ended extraction tasks, supply a judge LLM to perform model-based rubric evaluation:
from sieves import Doc, Pipeline
from sieves.tasks.predictive.information_extraction import InformationExtractionTask
import dspy
# Documents with gold-standard entity annotations
docs = [
Doc(
text="John bought 3 apples on 2024-01-01.",
gold={"ie": [{"entity": "John", "type": "PERSON"},
{"entity": "3", "type": "QUANTITY"},
{"entity": "2024-01-01", "type": "DATE"}]},
),
]
# Configure extraction task
ie_task = InformationExtractionTask(
task_id="ie",
schema={"entity": str, "type": str},
mode="single",
)
pipe = Pipeline([ie_task])
pipe(docs)
# Instantiate judge LLM via DSPy
judge = dspy.OpenAI(model="gpt-4o-mini")
# Evaluate with model-based rubric
report = pipe.evaluate(docs, judge=judge)
print(report.task_reports["ie"].f1) # F1 score computed by the judge
Programmatic Metric Access
Access specific metrics programmatically by iterating through the pipeline evaluation report:
# report is a PipelineEvaluationReport
for task_id, task_report in report.task_reports.items():
print(f"Task {task_id}:")
for metric, value in task_report.metrics.items():
print(f" {metric}: {value:.3f}")
Key Implementation Details
The evaluation API supports two comparison modes controlled by the optional judge parameter. When judge=None (the default), Sieves uses built-in deterministic logic such as exact string matching or set-based overlap. When a dspy.LM instance is provided, the system delegates comparison to the language model using task-specific rubrics, enabling evaluation of generative or open-ended tasks where exact matching is insufficient.
All tasks inherit the uniform evaluate interface from the base Task class, ensuring consistent benchmarking across different pipeline configurations. The PipelineEvaluationReport returned by pipeline.evaluate() aggregates these individual reports while maintaining access to per-task breakdowns via the task_reports dictionary attribute.
Summary
- Populate
Doc.goldwith task-specific keys before processing to establish ground truth references. - Run
pipeline.evaluate(docs)without a judge for deterministic scoring, or pass adspy.LMinstance for LLM-based evaluation. - Access detailed metrics through
PipelineEvaluationReport.task_reports, which contains individualTaskEvaluationReportobjects for each pipeline step. - Evaluation logic is centralized in
sieves/tasks/core.pyand aggregated insieves/pipeline/core.py, providing a consistent API across classification, NER, and extraction tasks.
Frequently Asked Questions
What is the difference between Doc.results and Doc.gold?
Doc.results stores the structured predictions generated by pipeline tasks, while Doc.gold contains the reference annotations you provide for evaluation. Both are dictionaries keyed by task IDs (e.g., "cls", "ner"), allowing the evaluate method to compare predicted values against ground truth using deterministic logic or an LLM judge.
Can I use a custom LLM as a judge for evaluation?
Yes. The evaluate method accepts an optional judge parameter of type dspy.LM. Pass any DSPy language model instance (such as dspy.OpenAI or dspy.Ollama) to enable model-based rubric evaluation. When judge=None, Sieves falls back to deterministic comparison methods like exact match or set overlap.
How do I access specific metrics like F1 or accuracy?
Individual metrics are available through the TaskEvaluationReport objects stored in PipelineEvaluationReport.task_reports. For predictive tasks, you can access attributes directly (e.g., report.task_reports["ie"].f1) or iterate through the metrics dictionary to retrieve accuracy, precision, recall, or task-specific scores computed during evaluation.
Does evaluation work for all task types in Sieves?
Yes. All tasks inherit from the base Task class defined in sieves/tasks/core.py, which implements the core evaluate method. Predictive tasks (classification, NER, information extraction) extend this through PredictiveTask in sieves/tasks/predictive/core.py to provide specialized metrics, but the uniform API ensures every task in a pipeline can be evaluated against ground truth data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →