How to Export Results to HuggingFace Datasets Format in Sieves
Call the to_hf_dataset() method on any Sieves predictive task instance, passing your processed Doc objects to generate a standard HuggingFace datasets.Dataset with text and labels columns that you can save locally or push to the Hub.
The mantisai/sieves library provides native interoperability with the HuggingFace ecosystem through built-in dataset conversion utilities. When you need to export results to HuggingFace Datasets format in Sieves, the framework offers a seamless method that transforms your processed documents into the industry-standard format used for training and sharing NLP datasets.
Understanding the to_hf_dataset Method
The export functionality is implemented in the base Task class located in sieves/tasks/predictive/core.py (lines 373-398). This generic helper iterates over your documents, extracts task-specific predictions from doc.results[task_id], and constructs a HuggingFace Dataset with two standard columns:
text: The original document contentlabels: The model's predictions for that document
Most predictive tasks—including Classification, NER, and QA—inherit this implementation directly from the base class. Specific task types may extend this behavior; for example, classification tasks in sieves/tasks/predictive/classification/core.py (lines 403-422) utilize the standard export mechanism, while specialized tasks like InformationExtraction may add flags to handle multi-label formats.
Step-by-Step Export Guide
1. Run a Predictive Task
First, process your documents through a Sieves pipeline to populate the results:
from sieves import Doc, Pipeline
from sieves.tasks.predictive import Classification
docs = [
Doc(text="I love this product!"),
Doc(text="Terrible experience.")
]
classifier = Classification(model="gpt-4o-mini", labels=["positive", "negative"])
pipe = Pipeline([classifier])
docs = pipe(docs) # doc.results["classification"] now contains predictions
2. Convert to HuggingFace Dataset Format
Call to_hf_dataset() on the task instance, passing the processed documents:
hf_dataset = classifier.to_hf_dataset(docs)
This returns a datasets.Dataset object compatible with the entire HuggingFace ecosystem.
3. Apply Confidence Thresholds (Optional)
Filter out low-confidence predictions by specifying a threshold parameter. Only predictions with a confidence score greater than or equal to the threshold will be included in the export:
hf_dataset = classifier.to_hf_dataset(docs, threshold=0.8)
4. Save or Push to the Hub
Persist your dataset using standard HuggingFace datasets methods:
# Save locally
hf_dataset.save_to_disk("my_sieves_export")
# Or push to the HuggingFace Hub
hf_dataset.push_to_hub("username/my-sieves-dataset")
Loading Datasets Back into Sieves
The reverse operation—converting a HuggingFace Dataset back into Sieves Doc objects—is handled by the Doc.from_hf_dataset class method in sieves/data/doc.py (lines 82-98). This method maps the standard text column to doc.text and optionally uses an id column for document identifiers:
from datasets import load_from_disk
from sieves import Doc
# Load the dataset
dataset = load_from_disk("my_sieves_export")
# Convert back to list of Doc objects
docs = Doc.from_hf_dataset(dataset)
Complete Working Example
This end-to-end example demonstrates classification, export with threshold filtering, persistence, and reload:
from sieves import Doc, Pipeline
from sieves.tasks.predictive import Classification
from datasets import load_from_disk
# 1. Prepare documents and run classification
docs = [
Doc(text="The movie was fantastic!"),
Doc(text="I did not enjoy the dinner."),
Doc(text="Mediocre performance at best."),
]
classifier = Classification(
model="gpt-4o-mini",
labels=["positive", "negative", "neutral"]
)
pipeline = Pipeline([classifier])
processed = pipeline(docs)
# 2. Export to HuggingFace Dataset with confidence threshold
hf_ds = classifier.to_hf_dataset(processed, threshold=0.6)
# 3. Verify structure
print("Features:", hf_ds.features) # {'text': Value('string'), 'labels': ClassLabel(...)}
print("Sample row:", hf_ds[0])
# 4. Save locally
hf_ds.save_to_disk("exported_sieves_dataset")
# 5. Reload and restore documents
reloaded = load_from_disk("exported_sieves_dataset")
docs_back = Doc.from_hf_dataset(reloaded)
print(f"Restored {len(docs_back)} documents")
Note: This functionality requires the datasets library, which is available as an optional dependency in the mantisai/sieves project.
Summary
- Core method:
Task.to_hf_dataset()insieves/tasks/predictive/core.py(lines 373-398) handles the conversion logic for all predictive tasks. - Output format: Standard HuggingFace
Datasetwithtextandlabelscolumns, ready for ML workflows. - Filtering: Optional
thresholdparameter filters predictions by confidence score before export. - Persistence: Use
save_to_disk()for local storage orpush_to_hub()for sharing on the HuggingFace Hub. - Round-trip support:
Doc.from_hf_dataset()insieves/data/doc.py(lines 82-98) enables loading standard datasets back into Sieves.
Frequently Asked Questions
What columns does the exported HuggingFace Dataset contain?
According to the implementation in sieves/tasks/predictive/core.py, the exported dataset always contains a text column (holding the original document string) and a labels column (containing the task predictions). Specific task types may structure the labels differently—classification uses class labels, while NER tasks may export span-based annotations.
Can I export results from any Sieves task type?
Yes. All predictive tasks in the sieves.tasks.predictive module inherit the to_hf_dataset() method from the base Task class defined in sieves/tasks/predictive/core.py. This includes Classification, NER, QA, and Information Extraction tasks. Custom tasks that properly store results in doc.results can also utilize this export functionality.
How do I filter predictions by confidence before exporting?
Pass the threshold parameter to to_hf_dataset(). For example, task.to_hf_dataset(docs, threshold=0.8) will only include predictions with a confidence score of 0.8 or higher in the final dataset. Predictions below this threshold are excluded from the labels column during the export process.
Is it possible to reload a HuggingFace Dataset back into Sieves documents?
Yes. Use the Doc.from_hf_dataset() class method implemented in sieves/data/doc.py (lines 82-98). This method reads a HuggingFace Dataset, maps the text column to document content and the id column to document identifiers, and returns a list of Doc objects that can be processed by Sieves pipelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →