# How to Use Predictive Tasks in Sieves: NER, Classification, IE, RE, and QA

> Learn how to use Sieves for NER, Classification, IE, RE, and QA. Unify predictive tasks with a common Pipeline interface for efficient execution.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Sieves unifies Named Entity Recognition, text classification, information extraction, relation extraction, and question answering into interchangeable `PredictiveTask` subclasses that execute through a common `Pipeline` interface with automatic batching, caching, and observability.**

Sieves is an open-source NLP framework that standardizes how predictive tasks are executed across different model backends. Whether you need to extract entities, classify documents, or answer questions, Sieves provides a consistent API through its `PredictiveTask` abstraction and `Pipeline` orchestration.

## Understanding the Predictive Task Architecture

All predictive tasks in Sieves share a common execution model defined in [[`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py)](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py). The architecture separates concerns into three layers: the **Task** itself, the **Bridge** that handles prompt engineering, and the **ModelWrapper** that executes inference.

### The PredictiveTask Base Class

The `PredictiveTask` abstract base class defines the interface that all concrete tasks implement. It handles:

- **Bridge selection**: Automatically chooses or accepts a custom `Bridge` that converts documents into prompts.
- **Result storage**: Places parsed outputs into `doc.results[task_id]` as Pydantic models.
- **Metadata tracking**: Records token usage and raw responses in `doc.meta[task_id]`.

### Bridges and Model Wrappers

Each task uses a **Bridge** (defined in task-specific [`bridges.py`](https://github.com/mantisai/sieves/blob/main/bridges.py) files) to:

1. Render Jinja2 prompt templates with few-shot examples formatted as XML.
2. Declare the output schema as a Pydantic model for structured generation.
3. Parse raw model outputs into the declared schema.

The **ModelWrapper** (e.g., [[`sieves/model_wrappers/outlines.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/outlines.py)](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/outlines.py)) executes the prompt on the chosen backend (Outlines, DSPy, LangChain, etc.) and returns the raw text for parsing.

## Named Entity Recognition (NER)

The NER task identifies entity spans and their types in unstructured text. Implemented in [[`sieves/tasks/predictive/ner/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/ner/core.py)](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/ner/core.py), it returns a list of `Entity` objects containing text, label, start/end offsets, and confidence scores.

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.ner import NERTask

# Create a document

doc = Doc(text="Alice works at OpenAI in San Francisco.")

# Initialise the NER task with custom entity types

ner = NERTask(
    model="gpt-4o-mini",
    label_set=["PERSON", "ORG", "LOC"],
)

# Build and run the pipeline

pipeline = Pipeline([ner])
result_docs = pipeline([doc])

# Access parsed entities

entities = result_docs[0].results[ner.id]
for ent in entities:
    print(f"{ent.text} → {ent.label}")

```

The bridge at [[`sieves/tasks/predictive/ner/bridges.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/ner/bridges.py)](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/ner/bridges.py) handles the prompt template that instructs the model to output structured entity lists.

## Text Classification

Classification assigns a single label from a predefined set to a document. The [`ClassificationTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/classification/core.py) supports few-shot learning through the `fewshot_examples` parameter, which automatically formats examples as XML in the prompt.

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.classification import ClassificationTask

doc = Doc(text="The movie was absolutely fantastic and thrilling!")

clf = ClassificationTask(
    model="gpt-4o-mini",
    label_set=["positive", "negative", "neutral"],
    fewshot_examples=[
        {"text": "I love this product!", "label": "positive"},
        {"text": "This is terrible.", "label": "negative"},
    ],
)

pipeline = Pipeline([clf])
out_doc = pipeline([doc])[0]

prediction = out_doc.results[clf.id]
print(f"Label: {prediction.label}, confidence: {prediction.confidence}")

```

The result object contains the predicted label and a confidence score derived from the model's output probabilities when available.

## Information Extraction (IE)

Information Extraction pulls structured key-value pairs from unstructured text. The [`InformationExtractionTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/information_extraction/core.py) accepts a schema dictionary mapping field names to types (string, date, money, etc.) and supports two modes: `single` (returns the first match) and `multiple` (returns a list of matches).

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.information_extraction import InformationExtractionTask

doc = Doc(text="""
Invoice #12345
Date: 2024‑02‑20
Total: $1,250.00
Due: 2024‑03‑20
""")

ie = InformationExtractionTask(
    model="gpt-4o-mini",
    schema={
        "invoice_number": "string",
        "date": "date",
        "total": "money",
        "due_date": "date"
    },
    mode="single",
)

pipeline = Pipeline([ie])
out = pipeline([doc])[0]

print(out.results[ie.id])

# → {'invoice_number': '12345', 'date': '2024‑02‑20', ...}

```

The schema types guide the Pydantic model generation in the bridge, ensuring type-safe extraction with validation.

## Relation Extraction (RE)

Relation Extraction identifies semantic relationships between entities. The [`RelationExtractionTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/relation_extraction/core.py) requires a `relation_schema` that defines valid subject-predicate-object patterns with entity types.

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.relation_extraction import RelationExtractionTask

doc = Doc(text="Alice, a researcher at OpenAI, presented her work in Paris.")

re_task = RelationExtractionTask(
    model="gpt-4o-mini",
    relation_schema=[
        {"subject": "PERSON", "predicate": "works_for", "object": "ORG"},
        {"subject": "PERSON", "predicate": "presented_at", "object": "LOC"},
    ],
)

pipeline = Pipeline([re_task])
out = pipeline([doc])[0]

for rel in out.results[re_task.id]:
    print(f"{rel.subject} –{rel.predicate}→ {rel.object}")

```

The task returns `Relation` objects containing the subject text, predicate, object text, and confidence scores.

## Question Answering (QA)

The Question Answering task extracts answers from a given context. The [`QuestionAnsweringTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/question_answering/core.py) requires a `question` parameter and processes the document text as the context.

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive.question_answering import QuestionAnsweringTask

context = """
The Eiffel Tower was completed in 1889 and stands 324 meters tall. It
attracts millions of visitors each year.
"""

doc = Doc(text=context)

qa = QuestionAnsweringTask(
    model="gpt-4o-mini",
    question="When was the Eiffel Tower completed?",
)

pipeline = Pipeline([qa])
out = pipeline([doc])[0]

answer = out.results[qa.id]
print(answer.answer)  # → "1889"

```

The result object contains the answer string and optional confidence metrics.

## Pipeline Orchestration and Caching

All predictive tasks integrate with the [`Pipeline`](https://github.com/mantisai/sieves/blob/main/sieves/pipeline.py) class for execution management. The pipeline handles:

- **Sequential execution**: Tasks run in order, with each task accessing `doc.results` from previous steps.
- **Automatic batching**: Controlled via `batch_size` in `ModelSettings`. Setting `batch_size=-1` processes all chunks together; otherwise, the pipeline respects the configured size.
- **Result caching**: Cached per-document hash, so repeated runs on identical `Doc` objects return instantly.
- **Observability**: Every model call records token usage (`input_tokens`, `output_tokens`) and the raw LLM response inside `doc.meta[task_id]`, enabling straightforward cost tracking and debugging.

```python
from sieves import Pipeline
from sieves.tasks.predictive.ner import NERTask
from sieves.tasks.predictive.classification import ClassificationTask

# Compose multiple predictive tasks

ner = NERTask(model="gpt-4o-mini", label_set=["PERSON", "ORG"])
clf = ClassificationTask(model="gpt-4o-mini", label_set=["news", "blog", "paper"])

pipeline = Pipeline([ner, clf])
results = pipeline(docs)

# Access results from both tasks

for doc in results:
    entities = doc.results[ner.id]
    category = doc.results[clf.id]

```

## Summary

- **Unified API**: All predictive tasks in Sieves inherit from `PredictiveTask` in [[`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py)](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py), providing a consistent interface for NER, classification, information extraction, relation extraction, and QA.
- **Bridge Pattern**: Each task uses a **Bridge** to handle Jinja2 prompt templating and Pydantic output parsing, separating model interaction from task logic.
- **Pipeline Orchestration**: The [`Pipeline`](https://github.com/mantisai/sieves/blob/main/sieves/pipeline.py) class sequences tasks, manages batching via `batch_size`, caches results by document hash, and tracks token usage in `doc.meta`.
- **Interchangeability**: Tasks can be stacked (e.g., NER → Relation Extraction) without glue code, with each task accessing previous results via `doc.results`.

## Frequently Asked Questions

### How do I add custom labels to a predictive task in Sieves?

Pass the `label_set` parameter when instantiating the task. For example, in [`NERTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/ner/core.py), set `label_set=["PRODUCT", "PRICE", "LOCATION"]` to constrain the model to your custom entity types. The bridge automatically incorporates these labels into the Jinja2 prompt template.

### Can I chain multiple predictive tasks together in a single pipeline?

Yes. The [`Pipeline`](https://github.com/mantisai/sieves/blob/main/sieves/pipeline.py) class accepts a list of tasks and executes them sequentially. Each task can access results from previous tasks via `doc.results[task_id]`. For example, you can pipe `NERTask` output into `RelationExtractionTask` to identify relationships between extracted entities without writing intermediate glue code.

### How does Sieves handle batching and caching for predictive tasks?

Batching is controlled through `ModelSettings` with the `batch_size` parameter. Setting `batch_size=-1` processes all document chunks together, while positive integers split work into smaller batches. Caching operates on document hashes—re-running a pipeline on identical `Doc` objects returns cached results instantly. Token usage metrics are stored in `doc.meta[task_id]` for cost tracking.

### What is the difference between Information Extraction and Relation Extraction in Sieves?

**Information Extraction** ([`InformationExtractionTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/information_extraction/core.py)) extracts flat key-value pairs from text using a defined schema (e.g., `{"date": "date", "amount": "money"}`). **Relation Extraction** ([`RelationExtractionTask`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/relation_extraction/core.py)) identifies semantic triples between entities (subject-predicate-object) using a `relation_schema` that constrains valid entity types for each argument.