How to Evaluate Pipeline Performance Against Ground Truth Data in Sieves

To evaluate pipeline performance in Sieves, populate the Doc.gold attribute with ground-truth labels, run the pipeline to fill Doc.results, then call pipeline.evaluate(docs) to generate a PipelineEvaluationReport comparing predictions against references.

The open-source Sieves library (mantisai/sieves) provides a unified evaluation framework that compares model predictions against annotated ground truth at both the task and pipeline levels. By storing references in Doc.gold and predictions in Doc.results, you can benchmark classification, NER, and information extraction tasks using either deterministic metrics or LLM-based judges.

Understanding the Evaluation Architecture

Document Structure: Gold vs. Results

Every document in Sieves is represented by a Doc object that carries dual dictionaries for evaluation:

  • Doc.results – Stores structured outputs generated by the pipeline, keyed by task IDs (e.g., "cls", "ner", "ie").
  • Doc.gold – Holds the ground-truth labels or annotations supplied by the user, keyed by the same task IDs.

When you run a pipeline, tasks populate doc.results[task_id] with their predictions. During evaluation, the system compares these entries against doc.gold[task_id] to compute performance metrics.

Task-Level Evaluation

Each task class implements an evaluate method defined in sieves/tasks/core.py that performs the core comparison logic. The base Task class provides the signature:

def evaluate(self, docs: Iterable[Doc], judge: dspy.LM | None = None) -> TaskEvaluationReport

According to the source code at line 99, this method iterates through documents, compares doc.results[task_id] with doc.gold[task_id], and produces a TaskEvaluationReport containing task-specific metrics. Predictive tasks (classification, NER, information extraction) inherit from PredictiveTask in sieves/tasks/predictive/core.py and may enrich reports with additional fields such as confidence-score analysis or per-label breakdowns at line 287.

Pipeline-Level Aggregation

The Pipeline class orchestrates evaluation across all tasks. When you call pipeline.evaluate(docs, judge), it executes the task-level evaluate method for every task in the sequence, then collates individual TaskEvaluationReport objects into a single PipelineEvaluationReport. This aggregation logic is implemented in sieves/pipeline/core.py at line 284, providing a holistic view of performance across the entire workflow.

Step-by-Step Evaluation Workflow

  1. Prepare ground-truth data – Populate doc.gold[task_id] for each document before running the pipeline, ensuring keys match your task IDs.
  2. Run inference – Execute pipeline(docs) to fill doc.results with model predictions.
  3. Call evaluate – Invoke report = pipeline.evaluate(docs, judge=my_judge) to generate the evaluation report. Omit the judge parameter for deterministic comparison, or provide a dspy.LM instance for LLM-based scoring.
  4. Inspect results – Access report.task_reports[task_id] to view specific metrics like accuracy, F1, or exact-match scores.

Code Examples

Classification with Deterministic Scoring

For classification tasks, use deterministic exact-match evaluation by omitting the judge parameter:

from sieves import Doc, Pipeline
from sieves.tasks.predictive.classification import ClassificationTask

# 1. Build gold-label documents

docs = [
    Doc(text="I love this product!", gold={"cls": "positive"}),
    Doc(text="Terrible experience.", gold={"cls": "negative"}),
]

# 2. Create a classification task

cls_task = ClassificationTask(
    task_id="cls",
    label_set=["positive", "negative"],
)

# 3. Assemble and run the pipeline

pipe = Pipeline([cls_task])
pipe(docs)

# 4. Evaluate against ground truth (deterministic)

report = pipe.evaluate(docs)
print(report.summary())
print(report.task_reports["cls"])  # Detailed classification metrics

Information Extraction with an LLM Judge

For open-ended extraction tasks, supply a judge LLM to perform model-based rubric evaluation:

from sieves import Doc, Pipeline
from sieves.tasks.predictive.information_extraction import InformationExtractionTask
import dspy

# Documents with gold-standard entity annotations

docs = [
    Doc(
        text="John bought 3 apples on 2024-01-01.",
        gold={"ie": [{"entity": "John", "type": "PERSON"},
                     {"entity": "3", "type": "QUANTITY"},
                     {"entity": "2024-01-01", "type": "DATE"}]},
    ),
]

# Configure extraction task

ie_task = InformationExtractionTask(
    task_id="ie",
    schema={"entity": str, "type": str},
    mode="single",
)

pipe = Pipeline([ie_task])
pipe(docs)

# Instantiate judge LLM via DSPy

judge = dspy.OpenAI(model="gpt-4o-mini")

# Evaluate with model-based rubric

report = pipe.evaluate(docs, judge=judge)
print(report.task_reports["ie"].f1)  # F1 score computed by the judge

Programmatic Metric Access

Access specific metrics programmatically by iterating through the pipeline evaluation report:


# report is a PipelineEvaluationReport

for task_id, task_report in report.task_reports.items():
    print(f"Task {task_id}:")
    for metric, value in task_report.metrics.items():
        print(f"  {metric}: {value:.3f}")

Key Implementation Details

The evaluation API supports two comparison modes controlled by the optional judge parameter. When judge=None (the default), Sieves uses built-in deterministic logic such as exact string matching or set-based overlap. When a dspy.LM instance is provided, the system delegates comparison to the language model using task-specific rubrics, enabling evaluation of generative or open-ended tasks where exact matching is insufficient.

All tasks inherit the uniform evaluate interface from the base Task class, ensuring consistent benchmarking across different pipeline configurations. The PipelineEvaluationReport returned by pipeline.evaluate() aggregates these individual reports while maintaining access to per-task breakdowns via the task_reports dictionary attribute.

Summary

  • Populate Doc.gold with task-specific keys before processing to establish ground truth references.
  • Run pipeline.evaluate(docs) without a judge for deterministic scoring, or pass a dspy.LM instance for LLM-based evaluation.
  • Access detailed metrics through PipelineEvaluationReport.task_reports, which contains individual TaskEvaluationReport objects for each pipeline step.
  • Evaluation logic is centralized in sieves/tasks/core.py and aggregated in sieves/pipeline/core.py, providing a consistent API across classification, NER, and extraction tasks.

Frequently Asked Questions

What is the difference between Doc.results and Doc.gold?

Doc.results stores the structured predictions generated by pipeline tasks, while Doc.gold contains the reference annotations you provide for evaluation. Both are dictionaries keyed by task IDs (e.g., "cls", "ner"), allowing the evaluate method to compare predicted values against ground truth using deterministic logic or an LLM judge.

Can I use a custom LLM as a judge for evaluation?

Yes. The evaluate method accepts an optional judge parameter of type dspy.LM. Pass any DSPy language model instance (such as dspy.OpenAI or dspy.Ollama) to enable model-based rubric evaluation. When judge=None, Sieves falls back to deterministic comparison methods like exact match or set overlap.

How do I access specific metrics like F1 or accuracy?

Individual metrics are available through the TaskEvaluationReport objects stored in PipelineEvaluationReport.task_reports. For predictive tasks, you can access attributes directly (e.g., report.task_reports["ie"].f1) or iterate through the metrics dictionary to retrieve accuracy, precision, recall, or task-specific scores computed during evaluation.

Does evaluation work for all task types in Sieves?

Yes. All tasks inherit from the base Task class defined in sieves/tasks/core.py, which implements the core evaluate method. Predictive tasks (classification, NER, information extraction) extend this through PredictiveTask in sieves/tasks/predictive/core.py to provide specialized metrics, but the uniform API ensures every task in a pipeline can be evaluated against ground truth data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →