# Working with ImageFileDataset in VERONA: A Complete Guide to Managing Image-Based Experiment Data

> Master VERONA's ImageFileDataset for efficient image experiment data management. Load image tensors and CSV labels with ID indexing, preprocessing, and subsampling. Your complete guide.

- Repository: [ADA research/verona](https://github.com/ada-research/verona)
- Tags: how-to-guide
- Published: 2026-02-23

---

**ImageFileDataset is a PyTorch-compatible dataset class in the VERONA framework that loads image tensors from disk alongside CSV labels, providing ID-based indexing, optional torchvision preprocessing, and efficient subsampling capabilities.**

The VERONA library (`ada-research/verona`) provides specialized infrastructure for experimental machine learning workflows. At the center of its image handling capabilities sits the **ImageFileDataset** class, which inherits from the abstract `ExperimentDataset` base to deliver a standardized interface for folder-based image collections.

## Architecture and Core Components

### The ExperimentDataset Abstract Base

All dataset implementations in VERONA descend from the abstract class defined in [`ada_verona/database/dataset/experiment_dataset.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/experiment_dataset.py). This base enforces a strict contract requiring concrete classes to implement `__len__`, `__getitem__`, and `get_subset`, ensuring consistent behavior across the framework's data pipeline.

### The DataPoint Container

Individual samples are encapsulated by the `DataPoint` class located in [`ada_verona/database/dataset/data_point.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/data_point.py). Each instance stores three attributes: `id` (a string identifier), `label` (an integer), and `data` (the image tensor). The class provides `to_dict()` and `from_dict()` methods for JSON-serializable representations.

### ImageFileDataset Implementation

The concrete implementation in [`ada_verona/database/dataset/image_file_dataset.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/image_file_dataset.py) orchestrates the mapping between image files and labels. It constructs an internal `IDIndex` structure that maps string identifiers (filenames without extensions) to numeric list positions, enabling deterministic lookups by original filename rather than arbitrary integer indices.

## Loading and Preprocessing Image Data

When instantiating **ImageFileDataset**, the constructor invokes `merge_label_file_with_images()` to build a pandas DataFrame (`image_label_df`) that aligns each image filename with its corresponding label from the CSV. The method `get_labeled_image_list()` then converts this DataFrame into a sorted list of `(Path, int)` tuples to guarantee deterministic ordering.

```python
from pathlib import Path
import torchvision.transforms as T
from ada_verona.database.dataset.image_file_dataset import ImageFileDataset

img_dir = Path("data/images")
label_csv = Path("data/labels.csv")

preprocess = T.Compose([
    T.ToTensor(),
    T.Normalize(mean=[0.5], std=[0.5])
])

dataset = ImageFileDataset(
    image_folder=img_dir,
    label_file=label_csv,
    preprocessing=preprocess
)

print(f"Dataset size: {len(dataset)}")
first_point = dataset[0]
print(f"ID: {first_point.id}, label: {first_point.label}, tensor shape: {first_point.data.shape}")

```

*Key implementation details*: The `__init__` and `__getitem__` methods in [`ada_verona/database/dataset/image_file_dataset.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/image_file_dataset.py) (lines 48-63) handle the tensor loading via `torch.load` and optional transform application.

## Accessing Data by Identifier

Unlike standard PyTorch datasets that restrict access to integer indices, **ImageFileDataset** maintains bidirectional mapping through the `IDIndex` object. The method `get_id_index_from_value()` retrieves the numeric index for a given string ID, returning `-1` if the identifier is absent.

```python
target_id = "cat_001"
idx_obj = dataset.get_id_index_from_value(target_id)

if idx_obj != -1:
    point = dataset[idx_obj.index]
    print(f"Found {target_id}: label={point.label}")
else:
    print("ID not present in dataset")

```

*Source reference*: The lookup logic resides in `get_id_index_from_value` within [`ada_verona/database/dataset/image_file_dataset.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/image_file_dataset.py) (lines 28-41).

## Creating Subsets with get_subset

The `get_subset(values)` method generates a new **ImageFileDataset** instance containing only entries whose string IDs appear in the provided `values` list. Rather than copying data, it reuses the original image folder and label file while filtering the internal `_id_indices` list, making memory-efficient dataset splits possible.

```python
selected_ids = ["dog_007", "bird_023", "cat_001"]
sub_dataset = dataset.get_subset(selected_ids)

print(f"Subset size: {len(sub_dataset)}")
for dp in sub_dataset:
    print(dp.id, dp.label)

```

*Source reference*: The subset construction logic appears in [`ada_verona/database/dataset/image_file_dataset.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/image_file_dataset.py) (lines 59-74).

## Serializing DataPoint Objects

For checkpointing and metadata logging, `DataPoint` instances convert to plain dictionaries through `to_dict()`, which flattens the tensor to a serializable format. The class method `from_dict()` reconstructs the object, preserving the original ID and label.

```python
from ada_verona.database.dataset.data_point import DataPoint

dp = dataset[0]
dp_dict = dp.to_dict()

# Later reconstruction

dp_restored = DataPoint.from_dict(dp_dict)
assert dp.id == dp_restored.id

```

*Source reference*: Serialization utilities are defined in [`ada_verona/database/dataset/data_point.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/data_point.py) (lines 37-58).

## Summary

- **ImageFileDataset** in [`ada_verona/database/dataset/image_file_dataset.py`](https://github.com/ada-research/verona/blob/main/ada_verona/database/dataset/image_file_dataset.py) provides a thin abstraction over folders of `.pt` image tensors and CSV label files.
- The class implements the **ExperimentDataset** contract, ensuring compatibility with VERONA's verification and analysis modules.
- **IDIndex** mapping enables O(1) lookup by filename (string ID) via `get_id_index_from_value()`.
- **get_subset()** creates efficient dataset views by filtering indices rather than reloading data.
- **DataPoint** containers support serialization through `to_dict()` and `from_dict()` for persistent storage.

## Frequently Asked Questions

### What file formats does ImageFileDataset support?

ImageFileDataset expects image files saved as PyTorch serialized tensors (`.pt` files) and a CSV label file where the first column contains filenames (with extensions) and subsequent columns contain integer labels. The implementation uses `torch.load()` internally to deserialize individual images.

### How does ImageFileDataset handle image preprocessing?

The optional `preprocessing` parameter accepts any callable, typically a `torchvision.transforms.Compose` pipeline. This transform is applied to the loaded tensor in `__getitem__` before the data is wrapped in a `DataPoint` object, supporting normalization, resizing, or augmentation operations.

### Can I use ImageFileDataset with PyTorch DataLoader?

Yes. Because ImageFileDataset fully implements the `__len__` and `__getitem__` interface required by PyTorch, it integrates seamlessly with `torch.utils.data.DataLoader` for batching, shuffling, and multi-process data loading. The deterministic filename sorting ensures reproducible ordering when shuffle is disabled.

### How does ID-based indexing work under the hood?

During initialization, the dataset parses filenames to extract string IDs (text before the first period) and builds an `IDIndex` object storing both the list index and the string identifier. When you call `get_id_index_from_value()`, it queries this index to retrieve the numeric position required by `__getitem__`, enabling retrieval by original filename rather than arbitrary integer position.