# How to Evaluate Pipeline Performance Against Ground Truth Data in Sieves

> Evaluate pipeline performance in Sieves by comparing predictions to ground truth. Populate Doc gold attribute, run pipeline, and use pipeline.evaluate to generate a report.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: performance
- Published: 2026-03-06

---

**To evaluate pipeline performance in Sieves, populate the `Doc.gold` attribute with ground-truth labels, run the pipeline to fill `Doc.results`, then call `pipeline.evaluate(docs)` to generate a `PipelineEvaluationReport` comparing predictions against references.**

The open-source Sieves library (mantisai/sieves) provides a unified evaluation framework that compares model predictions against annotated ground truth at both the task and pipeline levels. By storing references in `Doc.gold` and predictions in `Doc.results`, you can benchmark classification, NER, and information extraction tasks using either deterministic metrics or LLM-based judges.

## Understanding the Evaluation Architecture

### Document Structure: Gold vs. Results

Every document in Sieves is represented by a `Doc` object that carries dual dictionaries for evaluation:

- **`Doc.results`** – Stores structured outputs generated by the pipeline, keyed by task IDs (e.g., `"cls"`, `"ner"`, `"ie"`).
- **`Doc.gold`** – Holds the ground-truth labels or annotations supplied by the user, keyed by the same task IDs.

When you run a pipeline, tasks populate `doc.results[task_id]` with their predictions. During evaluation, the system compares these entries against `doc.gold[task_id]` to compute performance metrics.

### Task-Level Evaluation

Each task class implements an `evaluate` method defined in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) that performs the core comparison logic. The base `Task` class provides the signature:

```python
def evaluate(self, docs: Iterable[Doc], judge: dspy.LM | None = None) -> TaskEvaluationReport

```

According to the source code at [line 99](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py#L99), this method iterates through documents, compares `doc.results[task_id]` with `doc.gold[task_id]`, and produces a `TaskEvaluationReport` containing task-specific metrics. Predictive tasks (classification, NER, information extraction) inherit from `PredictiveTask` in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) and may enrich reports with additional fields such as confidence-score analysis or per-label breakdowns at [line 287](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py#L287).

### Pipeline-Level Aggregation

The `Pipeline` class orchestrates evaluation across all tasks. When you call `pipeline.evaluate(docs, judge)`, it executes the task-level `evaluate` method for every task in the sequence, then collates individual `TaskEvaluationReport` objects into a single `PipelineEvaluationReport`. This aggregation logic is implemented in [`sieves/pipeline/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/pipeline/core.py) at [line 284](https://github.com/mantisai/sieves/blob/main/sieves/pipeline/core.py#L284), providing a holistic view of performance across the entire workflow.

## Step-by-Step Evaluation Workflow

1. **Prepare ground-truth data** – Populate `doc.gold[task_id]` for each document before running the pipeline, ensuring keys match your task IDs.
2. **Run inference** – Execute `pipeline(docs)` to fill `doc.results` with model predictions.
3. **Call evaluate** – Invoke `report = pipeline.evaluate(docs, judge=my_judge)` to generate the evaluation report. Omit the `judge` parameter for deterministic comparison, or provide a `dspy.LM` instance for LLM-based scoring.
4. **Inspect results** – Access `report.task_reports[task_id]` to view specific metrics like accuracy, F1, or exact-match scores.

## Code Examples

### Classification with Deterministic Scoring

For classification tasks, use deterministic exact-match evaluation by omitting the judge parameter:

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.classification import ClassificationTask

# 1. Build gold-label documents

docs = [
    Doc(text="I love this product!", gold={"cls": "positive"}),
    Doc(text="Terrible experience.", gold={"cls": "negative"}),
]

# 2. Create a classification task

cls_task = ClassificationTask(
    task_id="cls",
    label_set=["positive", "negative"],
)

# 3. Assemble and run the pipeline

pipe = Pipeline([cls_task])
pipe(docs)

# 4. Evaluate against ground truth (deterministic)

report = pipe.evaluate(docs)
print(report.summary())
print(report.task_reports["cls"])  # Detailed classification metrics

```

### Information Extraction with an LLM Judge

For open-ended extraction tasks, supply a judge LLM to perform model-based rubric evaluation:

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.information_extraction import InformationExtractionTask
import dspy

# Documents with gold-standard entity annotations

docs = [
    Doc(
        text="John bought 3 apples on 2024-01-01.",
        gold={"ie": [{"entity": "John", "type": "PERSON"},
                     {"entity": "3", "type": "QUANTITY"},
                     {"entity": "2024-01-01", "type": "DATE"}]},
    ),
]

# Configure extraction task

ie_task = InformationExtractionTask(
    task_id="ie",
    schema={"entity": str, "type": str},
    mode="single",
)

pipe = Pipeline([ie_task])
pipe(docs)

# Instantiate judge LLM via DSPy

judge = dspy.OpenAI(model="gpt-4o-mini")

# Evaluate with model-based rubric

report = pipe.evaluate(docs, judge=judge)
print(report.task_reports["ie"].f1)  # F1 score computed by the judge

```

### Programmatic Metric Access

Access specific metrics programmatically by iterating through the pipeline evaluation report:

```python

# report is a PipelineEvaluationReport

for task_id, task_report in report.task_reports.items():
    print(f"Task {task_id}:")
    for metric, value in task_report.metrics.items():
        print(f"  {metric}: {value:.3f}")

```

## Key Implementation Details

The evaluation API supports two comparison modes controlled by the optional **judge** parameter. When `judge=None` (the default), Sieves uses built-in deterministic logic such as exact string matching or set-based overlap. When a `dspy.LM` instance is provided, the system delegates comparison to the language model using task-specific rubrics, enabling evaluation of generative or open-ended tasks where exact matching is insufficient.

All tasks inherit the uniform `evaluate` interface from the base `Task` class, ensuring consistent benchmarking across different pipeline configurations. The `PipelineEvaluationReport` returned by `pipeline.evaluate()` aggregates these individual reports while maintaining access to per-task breakdowns via the `task_reports` dictionary attribute.

## Summary

- Populate `Doc.gold` with task-specific keys before processing to establish ground truth references.
- Run `pipeline.evaluate(docs)` without a judge for deterministic scoring, or pass a `dspy.LM` instance for LLM-based evaluation.
- Access detailed metrics through `PipelineEvaluationReport.task_reports`, which contains individual `TaskEvaluationReport` objects for each pipeline step.
- Evaluation logic is centralized in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) and aggregated in [`sieves/pipeline/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/pipeline/core.py), providing a consistent API across classification, NER, and extraction tasks.

## Frequently Asked Questions

### What is the difference between `Doc.results` and `Doc.gold`?

`Doc.results` stores the structured predictions generated by pipeline tasks, while `Doc.gold` contains the reference annotations you provide for evaluation. Both are dictionaries keyed by task IDs (e.g., `"cls"`, `"ner"`), allowing the `evaluate` method to compare predicted values against ground truth using deterministic logic or an LLM judge.

### Can I use a custom LLM as a judge for evaluation?

Yes. The `evaluate` method accepts an optional `judge` parameter of type `dspy.LM`. Pass any DSPy language model instance (such as `dspy.OpenAI` or `dspy.Ollama`) to enable model-based rubric evaluation. When `judge=None`, Sieves falls back to deterministic comparison methods like exact match or set overlap.

### How do I access specific metrics like F1 or accuracy?

Individual metrics are available through the `TaskEvaluationReport` objects stored in `PipelineEvaluationReport.task_reports`. For predictive tasks, you can access attributes directly (e.g., `report.task_reports["ie"].f1`) or iterate through the `metrics` dictionary to retrieve accuracy, precision, recall, or task-specific scores computed during evaluation.

### Does evaluation work for all task types in Sieves?

Yes. All tasks inherit from the base `Task` class defined in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py), which implements the core `evaluate` method. Predictive tasks (classification, NER, information extraction) extend this through `PredictiveTask` in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) to provide specialized metrics, but the uniform API ensures every task in a pipeline can be evaluated against ground truth data.