How to Handle Image Inputs with PIL Images in Sieves: A Complete Guide

The Doc class in Sieves accepts a list of Pillow (PIL) Image objects via its images attribute, enabling multimodal pipelines to process raw visual data alongside text without automatic conversion or resizing.

The mantisai/sieves library provides native support for multimodal processing through its central data container. When you need to handle image inputs with PIL Images in Sieves, the framework stores native Pillow objects that any downstream task can manipulate using the full PIL API. This approach gives you complete control over image formats, sizes, and preprocessing while maintaining a consistent document representation throughout your pipeline.

The Doc Data Model for Images

In sieves/data/doc.py, the Doc class declares the images field on line 26 as an optional list of PIL Image objects:


# sieves/data/doc.py (line 26)

images: list[Image.Image] | None = None

When you instantiate a Doc with the images parameter, Sieves stores the list verbatim. The framework performs no automatic resizing, mode conversion, or caching. This design expects the caller to supply images in the target format required by your specific vision model or preprocessing logic.

Creating Documents with PIL Images

Attach Pillow Images to a document by passing them to the images parameter during instantiation:

from PIL import Image
from sieves import Doc

# Create or load PIL images

red_square = Image.new("RGB", (100, 100), color="red")
blue_square = Image.new("RGB", (100, 100), color="blue")

# Attach to document

doc = Doc(
    text="A red and blue composition",
    images=[red_square, blue_square]
)

# Access images downstream

print(doc.images[0].size)  # Output: (100, 100)

The doc.images list contains standard PIL Image objects, allowing you to call any Pillow method directly on the stored data.

Processing Images in Custom Tasks

Any task inheriting from sieves.Task can access and modify doc.images. Because Sieves exposes raw PIL objects, you can apply arbitrary transformations before passing data to a model:

from sieves import Task
from PIL import Image

class ResizeImages(Task):
    def __call__(self, docs):
        for doc in docs:
            if doc.images:
                doc.images = [img.resize((224, 224)) for img in doc.images]
        return docs

The Grayscale task demonstrates mode conversion:

class Grayscale(Task):
    def __call__(self, docs):
        for doc in docs:
            if doc.images:
                doc.images = [img.convert("L") for img in doc.images]
        return docs

These examples show how Sieves lets you leverage the entire Pillow API within your pipeline stages.

Image Equality and Comparison Logic

The Doc class implements deep equality semantics for images in sieves/data/doc.py. The _are_images_equal helper function (lines 41-47) compares image pairs using ImageChops.difference, checking:

  • Image dimensions (size)
  • Color mode (mode)
  • Pixel-wise equality

# Example equality check

from PIL import Image
from sieves import Doc

img_a = Image.new("RGB", (100, 100), color="red")
img_b = Image.new("RGB", (100, 100), color="red")
img_c = Image.new("RGB", (100, 100), color="blue")

doc_1 = Doc(images=[img_a])
doc_2 = Doc(images=[img_b])
doc_3 = Doc(images=[img_c])

assert doc_1 == doc_2  # True: pixel-wise identical

assert doc_1 != doc_3  # False: different pixels

The equality logic in lines 41-47 requires both documents to have the same number of images, matching sizes, modes, and pixel values. If one document has images=None and the other has a list, they are considered unequal.

Additionally, attempting to compare a Doc instance with a non-Doc object raises NotImplementedError (lines 55-56), preventing ambiguous equality checks in collections or assertions.

Loading Images from Hugging Face Datasets

While Doc.from_hf_dataset converts Hugging Face datasets into Doc objects, it currently populates only textual fields. To handle image columns, manually inject PIL objects after the initial conversion:

from datasets import load_dataset
from sieves import Doc
from PIL import Image

# Load dataset

dataset = load_dataset("my_multimodal_dataset", split="train")
docs = Doc.from_hf_dataset(dataset, column_map={"text": "caption"})

# Inject images manually

for doc, img_path in zip(docs, dataset["image_path"]):
    doc.images = [Image.open(img_path)]

This workflow applies when your dataset stores image file paths or byte data that requires conversion to PIL format before integration with Sieves.

Summary

  • The Doc class stores PIL Images in the images attribute as declared in sieves/data/doc.py (line 26).
  • Sieves performs no automatic image preprocessing, giving you full access to the Pillow API for resizing, mode conversion, and filtering.
  • Equality comparison uses pixel-wise analysis via ImageChops.difference in _are_images_equal (lines 41-47), validating size, mode, and content.
  • Non-Doc comparisons raise NotImplementedError (lines 55-56) to enforce type safety.
  • When using from_hf_dataset, manually populate the images attribute after text conversion to include visual data from Hugging Face datasets.

Frequently Asked Questions

Does Sieves automatically resize or convert image formats?

No. According to the source code in sieves/data/doc.py, the images field stores the provided list verbatim without automatic resizing, mode conversion, or color-space adjustments. You must preprocess images to your target specifications (e.g., 224x224 RGB) before or within your pipeline tasks.

How does Sieves compare images for equality between two Doc objects?

The Doc.__eq__ method delegates to _are_images_equal, which validates that both images have identical size and mode, then uses ImageChops.difference to verify pixel-wise equality. If any pair of corresponding images differs in dimensions, mode, or pixel values, the two Doc objects are considered unequal. This logic is tested in sieves/tests/test_doc.py (lines 21-55).

Can I load images directly using Doc.from_hf_dataset?

No. The from_hf_dataset method only maps text columns to the Doc instance. If your Hugging Face dataset contains image paths or binary image data, you must manually create PIL Image objects and assign them to doc.images after calling from_hf_dataset, as shown in the dataset integration examples.

What happens if I compare a Doc with a non-Doc object?

Comparing a Doc instance with any non-Doc object raises NotImplementedError (lines 55-56 in sieves/data/doc.py). This prevents ambiguous equality results when Doc objects are used in collections or comparison operations with incompatible types.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →