Working with ImageFileDataset in VERONA: A Complete Guide to Managing Image-Based Experiment Data

ImageFileDataset is a PyTorch-compatible dataset class in the VERONA framework that loads image tensors from disk alongside CSV labels, providing ID-based indexing, optional torchvision preprocessing, and efficient subsampling capabilities.

The VERONA library (ada-research/verona) provides specialized infrastructure for experimental machine learning workflows. At the center of its image handling capabilities sits the ImageFileDataset class, which inherits from the abstract ExperimentDataset base to deliver a standardized interface for folder-based image collections.

Architecture and Core Components

The ExperimentDataset Abstract Base

All dataset implementations in VERONA descend from the abstract class defined in ada_verona/database/dataset/experiment_dataset.py. This base enforces a strict contract requiring concrete classes to implement __len__, __getitem__, and get_subset, ensuring consistent behavior across the framework's data pipeline.

The DataPoint Container

Individual samples are encapsulated by the DataPoint class located in ada_verona/database/dataset/data_point.py. Each instance stores three attributes: id (a string identifier), label (an integer), and data (the image tensor). The class provides to_dict() and from_dict() methods for JSON-serializable representations.

ImageFileDataset Implementation

The concrete implementation in ada_verona/database/dataset/image_file_dataset.py orchestrates the mapping between image files and labels. It constructs an internal IDIndex structure that maps string identifiers (filenames without extensions) to numeric list positions, enabling deterministic lookups by original filename rather than arbitrary integer indices.

Loading and Preprocessing Image Data

When instantiating ImageFileDataset, the constructor invokes merge_label_file_with_images() to build a pandas DataFrame (image_label_df) that aligns each image filename with its corresponding label from the CSV. The method get_labeled_image_list() then converts this DataFrame into a sorted list of (Path, int) tuples to guarantee deterministic ordering.

from pathlib import Path
import torchvision.transforms as T
from ada_verona.database.dataset.image_file_dataset import ImageFileDataset

img_dir = Path("data/images")
label_csv = Path("data/labels.csv")

preprocess = T.Compose([
    T.ToTensor(),
    T.Normalize(mean=[0.5], std=[0.5])
])

dataset = ImageFileDataset(
    image_folder=img_dir,
    label_file=label_csv,
    preprocessing=preprocess
)

print(f"Dataset size: {len(dataset)}")
first_point = dataset[0]
print(f"ID: {first_point.id}, label: {first_point.label}, tensor shape: {first_point.data.shape}")

Key implementation details: The __init__ and __getitem__ methods in ada_verona/database/dataset/image_file_dataset.py (lines 48-63) handle the tensor loading via torch.load and optional transform application.

Accessing Data by Identifier

Unlike standard PyTorch datasets that restrict access to integer indices, ImageFileDataset maintains bidirectional mapping through the IDIndex object. The method get_id_index_from_value() retrieves the numeric index for a given string ID, returning -1 if the identifier is absent.

target_id = "cat_001"
idx_obj = dataset.get_id_index_from_value(target_id)

if idx_obj != -1:
    point = dataset[idx_obj.index]
    print(f"Found {target_id}: label={point.label}")
else:
    print("ID not present in dataset")

Source reference: The lookup logic resides in get_id_index_from_value within ada_verona/database/dataset/image_file_dataset.py (lines 28-41).

Creating Subsets with get_subset

The get_subset(values) method generates a new ImageFileDataset instance containing only entries whose string IDs appear in the provided values list. Rather than copying data, it reuses the original image folder and label file while filtering the internal _id_indices list, making memory-efficient dataset splits possible.

selected_ids = ["dog_007", "bird_023", "cat_001"]
sub_dataset = dataset.get_subset(selected_ids)

print(f"Subset size: {len(sub_dataset)}")
for dp in sub_dataset:
    print(dp.id, dp.label)

Source reference: The subset construction logic appears in ada_verona/database/dataset/image_file_dataset.py (lines 59-74).

Serializing DataPoint Objects

For checkpointing and metadata logging, DataPoint instances convert to plain dictionaries through to_dict(), which flattens the tensor to a serializable format. The class method from_dict() reconstructs the object, preserving the original ID and label.

from ada_verona.database.dataset.data_point import DataPoint

dp = dataset[0]
dp_dict = dp.to_dict()

# Later reconstruction

dp_restored = DataPoint.from_dict(dp_dict)
assert dp.id == dp_restored.id

Source reference: Serialization utilities are defined in ada_verona/database/dataset/data_point.py (lines 37-58).

Summary

  • ImageFileDataset in ada_verona/database/dataset/image_file_dataset.py provides a thin abstraction over folders of .pt image tensors and CSV label files.
  • The class implements the ExperimentDataset contract, ensuring compatibility with VERONA's verification and analysis modules.
  • IDIndex mapping enables O(1) lookup by filename (string ID) via get_id_index_from_value().
  • get_subset() creates efficient dataset views by filtering indices rather than reloading data.
  • DataPoint containers support serialization through to_dict() and from_dict() for persistent storage.

Frequently Asked Questions

What file formats does ImageFileDataset support?

ImageFileDataset expects image files saved as PyTorch serialized tensors (.pt files) and a CSV label file where the first column contains filenames (with extensions) and subsequent columns contain integer labels. The implementation uses torch.load() internally to deserialize individual images.

How does ImageFileDataset handle image preprocessing?

The optional preprocessing parameter accepts any callable, typically a torchvision.transforms.Compose pipeline. This transform is applied to the loaded tensor in __getitem__ before the data is wrapped in a DataPoint object, supporting normalization, resizing, or augmentation operations.

Can I use ImageFileDataset with PyTorch DataLoader?

Yes. Because ImageFileDataset fully implements the __len__ and __getitem__ interface required by PyTorch, it integrates seamlessly with torch.utils.data.DataLoader for batching, shuffling, and multi-process data loading. The deterministic filename sorting ensures reproducible ordering when shuffle is disabled.

How does ID-based indexing work under the hood?

During initialization, the dataset parses filenames to extract string IDs (text before the first period) and builds an IDIndex object storing both the list index and the string identifier. When you call get_id_index_from_value(), it queries this index to retrieve the numeric position required by __getitem__, enabling retrieval by original filename rather than arbitrary integer position.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →