Working with PytorchExperimentDataset in VERONA: A Complete Guide

VERONA's PytorchExperimentDataset class wraps any standard PyTorch dataset to provide a unified DataPoint interface for verification workflows, enabling seamless integration with samplers, verification modules, and reporting tools.

The VERONA framework (available at ada-research/verona) standardizes how robustness and verification experiments handle data. At the heart of this system is the PytorchExperimentDataset class, which bridges the gap between arbitrary torch.utils.data.Dataset implementations and VERONA's internal verification pipeline. This guide explains how to use this wrapper effectively, from basic dataset wrapping to advanced subsetting and DataLoader integration.

What is PytorchExperimentDataset?

PytorchExperimentDataset is a thin wrapper located in ada_verona/database/dataset/pytorch_experiment_dataset.py. It implements Python's sequence protocol (__len__ and __getitem__) while maintaining an internal index mapping that allows for efficient subset creation without duplicating underlying data.

The class serves as the primary data interface for VERONA's VerificationModule and various samplers, ensuring that regardless of the original PyTorch dataset's structure, downstream components receive standardized DataPoint objects.

Key Design Features

Dataset Abstraction with Index Mapping

Rather than modifying the original dataset, PytorchExperimentDataset stores a reference to the underlying torch.utils.data.Dataset in self.dataset and maintains an internal list of valid indices in self._indices. When you access an item, the wrapper retrieves the data using the mapped index, allowing the framework to create arbitrary subsets by simply manipulating the index list.

This design is implemented in the __init__ method of PytorchExperimentDataset, where the full index range is generated by default: self._indices = list(range(len(dataset))).

Uniform DataPoint Interface

Every call to __getitem__ returns a DataPoint instance (defined in ada_verona/database/dataset/data_point.py). This dataclass standardizes access to:

  • id: The original dataset index (preserved even in subsets)
  • label: The classification label or target value
  • data: The raw tensor or feature vector

By enforcing this uniform interface, VERONA's verification modules can process data from MNIST, ImageNet, or custom datasets without modification.

Efficient Sub-setting

The get_subset(indices) method creates a new PytorchExperimentDataset instance that shares the same underlying PyTorch dataset but replaces self._indices with the provided list. This is memory-efficient—no data is copied—and is heavily used by the DatasetSampler and verification workflows when analyzing specific slices of data.

Working with PytorchExperimentDataset

Wrapping a Standard PyTorch Dataset

To use VERONA with existing PyTorch code, wrap your dataset immediately after instantiation:

from torchvision import datasets, transforms
from ada_verona.database.dataset.pytorch_experiment_dataset import PytorchExperimentDataset

# Load standard MNIST test set

mnist = datasets.MNIST(
    root="data",
    train=False,
    download=True,
    transform=transforms.ToTensor(),
)

# Wrap for VERONA compatibility

exp_dataset = PytorchExperimentDataset(mnist)
print(f"Total samples: {len(exp_dataset)}")  # Output: 10000

Accessing Individual DataPoints

Once wrapped, you access items through the DataPoint interface rather than raw tuples:


# Access first sample

dp = exp_dataset[0]

print(dp.id)      # 0 (original index preserved)

print(dp.label)   # 7 (example MNIST label)

print(dp.data.shape)  # torch.Size([1, 28, 28])

This standardized access pattern ensures that verification modules receive consistent input regardless of the original dataset's return format.

Creating Subsets for Targeted Analysis

Use get_subset() to create views into specific portions of your data without loading duplicates into memory:


# Create subset of first 100 samples

first_hundred = exp_dataset.get_subset(list(range(100)))
print(len(first_hundred))  # 100

# Indices reflect original positions

print(first_hundred[42].id)  # 42 (not 0)

# Create random sample subset

import random
random_indices = random.sample(range(len(exp_dataset)), 500)
random_subset = exp_dataset.get_subset(random_indices)

This functionality is essential when the VerificationModule needs to analyze specific adversarial examples or when samplers need to create stratified batches.

Integration with PyTorch DataLoader

Because PytorchExperimentDataset implements the sequence protocol, it works seamlessly with PyTorch's DataLoader for batch processing:

from torch.utils.data import DataLoader

loader = DataLoader(
    exp_dataset,
    batch_size=64,
    shuffle=True,
    num_workers=0  # Set >0 if dataset supports multiprocessing

)

for batch in loader:
    # batch is a list of DataPoint objects

    ids = [dp.id for dp in batch]
    data = torch.stack([dp.data for dp in batch])
    labels = torch.tensor([dp.label for dp in batch])
    
    # Process through model or verification routine

    # outputs = model(data)

Note that when using num_workers > 0, ensure your underlying PyTorch dataset is pickleable, as PytorchExperimentDataset itself supports standard pickling through its composition pattern.

Feeding Data to the VerificationModule

The primary use case for PytorchExperimentDataset is providing standardized data to VERONA's verification pipeline:

from ada_verona.verification_module.verification_module import VerificationModule
from ada_verona.verification_module.property_generator.one2any_property_generator import One2AnyPropertyGenerator

# Initialize verification with wrapped dataset

verifier = VerificationModule(
    dataset=exp_dataset,
    property_generator=One2AnyPropertyGenerator(),
    # Additional configuration for verification bounds, timeout, etc.

)

# Run verification analysis

results = verifier.run()
print(results.summary())

The VerificationModule relies on the DataPoint interface to consistently access id, label, and data attributes across different dataset types, making PytorchExperimentDataset essential for PyTorch-based experiments.

Summary

  • PytorchExperimentDataset in ada_verona/database/dataset/pytorch_experiment_dataset.py wraps any torch.utils.data.Dataset to provide VERONA-compatible data access.

  • The class maintains an internal index mapping (self._indices) that enables memory-efficient subset creation through the get_subset() method without copying underlying data.

  • All data access returns standardized DataPoint objects containing id, label, and data attributes, ensuring compatibility with VerificationModule and samplers.

  • The wrapper implements the Python sequence protocol, making it compatible with torch.utils.data.DataLoader for batch processing and multiprocessing.

  • Key integration points include the VerificationModule in ada_verona/verification_module/verification_module.py and various property generators that consume the standardized dataset interface.

Frequently Asked Questions

How does PytorchExperimentDataset handle dataset subsets without duplicating data in memory?

PytorchExperimentDataset uses an index mapping strategy where the get_subset(indices) method creates a new instance that shares the same underlying PyTorch dataset (self.dataset) but replaces self._indices with the provided list of integers. Because the actual data tensors remain stored in the original dataset and are only accessed by index during __getitem__ calls, creating subsets requires only O(n) memory for the index list rather than O(n) for the data itself.

Can I use PytorchExperimentDataset with custom PyTorch datasets that return multiple values?

Yes. PytorchExperimentDataset delegates item access to the underlying dataset and wraps the returned values in a DataPoint object. As long as your custom dataset returns a tuple or list where the first element is the data tensor and the second is the label (or follows the standard PyTorch dataset convention), the wrapper will correctly populate the DataPoint.data and DataPoint.label fields. If your dataset returns more complex structures, you may need to subclass PytorchExperimentDataset and override __getitem__ to properly map your custom returns to the DataPoint interface.

Is PytorchExperimentDataset compatible with PyTorch DataLoader multiprocessing?

Yes, PytorchExperimentDataset supports multiprocessing through PyTorch's DataLoader when the underlying PyTorch dataset is pickleable. Because PytorchExperimentDataset stores only a reference to the dataset and a list of indices (both pickleable if the base dataset is), it works with num_workers > 0 in DataLoader. However, ensure that your underlying dataset (e.g., custom torch.utils.data.Dataset implementations) properly implements __getstate__ and __setstate__ if it contains non-pickleable resources like file handles or database connections.

How does the DataPoint class standardize access across different dataset types?

The DataPoint class (defined in ada_verona/database/dataset/data_point.py) is a Python dataclass that enforces a uniform interface with three required attributes: id (the original dataset index), label (the ground truth or target), and data (the feature tensor). By wrapping all dataset returns in this structure, PytorchExperimentDataset ensures that downstream VERONA components—such as the VerificationModule in ada_verona/verification_module/verification_module.py and various property generators—can access data consistently without knowing the specifics of the original PyTorch dataset's return format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →