# How to Export Results to HuggingFace Datasets Format in Sieves

> Easily export Sieves results to HuggingFace Datasets format. Use to_hf_dataset() to create standard datasets with text and labels for local saving or pushing to the Hub.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Call the `to_hf_dataset()` method on any Sieves predictive task instance, passing your processed `Doc` objects to generate a standard HuggingFace `datasets.Dataset` with `text` and `labels` columns that you can save locally or push to the Hub.**

The mantisai/sieves library provides native interoperability with the HuggingFace ecosystem through built-in dataset conversion utilities. When you need to export results to HuggingFace Datasets format in Sieves, the framework offers a seamless method that transforms your processed documents into the industry-standard format used for training and sharing NLP datasets.

## Understanding the `to_hf_dataset` Method

The export functionality is implemented in the base `Task` class located in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) (lines 373-398). This generic helper iterates over your documents, extracts task-specific predictions from `doc.results[task_id]`, and constructs a HuggingFace `Dataset` with two standard columns:

- **`text`**: The original document content
- **`labels`**: The model's predictions for that document

Most predictive tasks—including **Classification**, **NER**, and **QA**—inherit this implementation directly from the base class. Specific task types may extend this behavior; for example, classification tasks in [`sieves/tasks/predictive/classification/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/classification/core.py) (lines 403-422) utilize the standard export mechanism, while specialized tasks like `InformationExtraction` may add flags to handle multi-label formats.

## Step-by-Step Export Guide

### 1. Run a Predictive Task

First, process your documents through a Sieves pipeline to populate the results:

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive import Classification

docs = [
    Doc(text="I love this product!"),
    Doc(text="Terrible experience.")
]

classifier = Classification(model="gpt-4o-mini", labels=["positive", "negative"])
pipe = Pipeline([classifier])
docs = pipe(docs)  # doc.results["classification"] now contains predictions

```

### 2. Convert to HuggingFace Dataset Format

Call `to_hf_dataset()` on the task instance, passing the processed documents:

```python
hf_dataset = classifier.to_hf_dataset(docs)

```

This returns a `datasets.Dataset` object compatible with the entire HuggingFace ecosystem.

### 3. Apply Confidence Thresholds (Optional)

Filter out low-confidence predictions by specifying a threshold parameter. Only predictions with a confidence score greater than or equal to the threshold will be included in the export:

```python
hf_dataset = classifier.to_hf_dataset(docs, threshold=0.8)

```

### 4. Save or Push to the Hub

Persist your dataset using standard HuggingFace `datasets` methods:

```python

# Save locally

hf_dataset.save_to_disk("my_sieves_export")

# Or push to the HuggingFace Hub

hf_dataset.push_to_hub("username/my-sieves-dataset")

```

## Loading Datasets Back into Sieves

The reverse operation—converting a HuggingFace Dataset back into Sieves `Doc` objects—is handled by the `Doc.from_hf_dataset` class method in [`sieves/data/doc.py`](https://github.com/mantisai/sieves/blob/main/sieves/data/doc.py) (lines 82-98). This method maps the standard `text` column to `doc.text` and optionally uses an `id` column for document identifiers:

```python
from datasets import load_from_disk
from sieves import Doc

# Load the dataset

dataset = load_from_disk("my_sieves_export")

# Convert back to list of Doc objects

docs = Doc.from_hf_dataset(dataset)

```

## Complete Working Example

This end-to-end example demonstrates classification, export with threshold filtering, persistence, and reload:

```python
from sieves import Doc, Pipeline
from sieves.tasks.predictive import Classification
from datasets import load_from_disk

# 1. Prepare documents and run classification

docs = [
    Doc(text="The movie was fantastic!"),
    Doc(text="I did not enjoy the dinner."),
    Doc(text="Mediocre performance at best."),
]

classifier = Classification(
    model="gpt-4o-mini", 
    labels=["positive", "negative", "neutral"]
)
pipeline = Pipeline([classifier])
processed = pipeline(docs)

# 2. Export to HuggingFace Dataset with confidence threshold

hf_ds = classifier.to_hf_dataset(processed, threshold=0.6)

# 3. Verify structure

print("Features:", hf_ds.features)  # {'text': Value('string'), 'labels': ClassLabel(...)}

print("Sample row:", hf_ds[0])

# 4. Save locally

hf_ds.save_to_disk("exported_sieves_dataset")

# 5. Reload and restore documents

reloaded = load_from_disk("exported_sieves_dataset")
docs_back = Doc.from_hf_dataset(reloaded)
print(f"Restored {len(docs_back)} documents")

```

*Note: This functionality requires the `datasets` library, which is available as an optional dependency in the mantisai/sieves project.*

## Summary

- **Core method**: `Task.to_hf_dataset()` in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) (lines 373-398) handles the conversion logic for all predictive tasks.
- **Output format**: Standard HuggingFace `Dataset` with `text` and `labels` columns, ready for ML workflows.
- **Filtering**: Optional `threshold` parameter filters predictions by confidence score before export.
- **Persistence**: Use `save_to_disk()` for local storage or `push_to_hub()` for sharing on the HuggingFace Hub.
- **Round-trip support**: `Doc.from_hf_dataset()` in [`sieves/data/doc.py`](https://github.com/mantisai/sieves/blob/main/sieves/data/doc.py) (lines 82-98) enables loading standard datasets back into Sieves.

## Frequently Asked Questions

### What columns does the exported HuggingFace Dataset contain?

According to the implementation in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py), the exported dataset always contains a **`text`** column (holding the original document string) and a **`labels`** column (containing the task predictions). Specific task types may structure the labels differently—classification uses class labels, while NER tasks may export span-based annotations.

### Can I export results from any Sieves task type?

Yes. All predictive tasks in the `sieves.tasks.predictive` module inherit the `to_hf_dataset()` method from the base `Task` class defined in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py). This includes Classification, NER, QA, and Information Extraction tasks. Custom tasks that properly store results in `doc.results` can also utilize this export functionality.

### How do I filter predictions by confidence before exporting?

Pass the `threshold` parameter to `to_hf_dataset()`. For example, `task.to_hf_dataset(docs, threshold=0.8)` will only include predictions with a confidence score of 0.8 or higher in the final dataset. Predictions below this threshold are excluded from the `labels` column during the export process.

### Is it possible to reload a HuggingFace Dataset back into Sieves documents?

Yes. Use the `Doc.from_hf_dataset()` class method implemented in [`sieves/data/doc.py`](https://github.com/mantisai/sieves/blob/main/sieves/data/doc.py) (lines 82-98). This method reads a HuggingFace `Dataset`, maps the `text` column to document content and the `id` column to document identifiers, and returns a list of `Doc` objects that can be processed by Sieves pipelines.