# Understanding the MegaWrapper Preprocessing Pipeline and Image Transforms in stable-worldmodel

> Explore the MegaWrapper preprocessing pipeline and image transforms in stable-worldmodel. Standardize environments to a pixel-based RL interface with efficient image preprocessing.

- Repository: [GalilAI-group/stable-worldmodel](https://github.com/galilai-group/stable-worldmodel)
- Tags: deep-dive
- Published: 2026-05-30

---

**The MegaWrapper class composes a sequence of specialized Gymnasium wrappers—AddPixelsWrapper, EverythingToInfoWrapper, EnsureInfoKeysWrapper, and ResizeGoalWrapper—to standardize any environment into a pixel-based reinforcement learning interface with configurable image preprocessing.**

The stable-worldmodel library provides a robust preprocessing stack for reinforcement learning researchers working with pixel observations. The **MegaWrapper preprocessing pipeline** serves as the central orchestrator, automatically rendering frames, applying transforms, and restructuring environment outputs into a consistent format suitable for world model training.

## The MegaWrapper Architecture

Located in [`stable_worldmodel/wrapper/default.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/default.py), the `MegaWrapper` class implements a four-stage wrapping strategy that transforms raw Gymnasium environments into fully pixel-based interfaces. Each stage handles a specific concern, from rendering to validation.

### AddPixelsWrapper and Image Rendering

The pipeline begins with `AddPixelsWrapper`, which calls `env.render()` to generate RGB arrays and resizes them to the target `image_shape`. According to the source code in [`default.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/default.py), this wrapper optionally applies a **torchvision-style transform** via the `pixels_transform` parameter before storing the result under the `pixels` key in the `info` dictionary.

```python

# Conceptual flow from AddPixelsWrapper (default.py lines 95-108)

pixels = env.render()  # RGB array

pixels = resize(pixels, image_shape, resample)  # bilinear/nearest

if pixels_transform:
    pixels = pixels_transform(pixels)
info['pixels'] = pixels

```

### EverythingToInfoWrapper for Unified Data Flow

Next, `EverythingToInfoWrapper` moves all observation data—including raw observations, rewards, and termination flags—into the `info` dictionary. This normalization allows downstream components to treat all environment outputs uniformly as metadata, simplifying the interface for pixel-based agents.

### EnsureInfoKeysWrapper Validation

Before completing the wrapping sequence, `EnsureInfoKeysWrapper` validates that required keys exist in the `info` dictionary. By default, it checks for keys matching the pattern `^pixels(?:\..*)?$`, raising a clear error if pixel data is missing. This prevents silent failures when preprocessing pipelines misconfigure.

### ResizeGoalWrapper for Goal-Conditioned Tasks

For environments that provide goal images, `ResizeGoalWrapper` (optional) processes these targets through the same resizing logic and an optional separate `goal_transform`. This ensures goal observations match the observation space dimensions, maintaining consistency in goal-conditioned reinforcement learning setups.

The `MegaWrapper.__init__` method wires these components together sequentially:

```python

# From MegaWrapper.__init__ in default.py

if add_pixels:
    env = AddPixelsWrapper(env, image_shape, pixels_transform, resample)
env = EverythingToInfoWrapper(env)
env = EnsureInfoKeysWrapper(env, req_keys)
if add_pixels:
    env = ResizeGoalWrapper(env, image_shape, goal_transform, resample)

```

## Image Transform Utilities and Visual Augmentations

The [`stable_worldmodel/wrapper/visual.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/visual.py) module extends the preprocessing pipeline with pixel-level augmentations. Each wrapper inherits from `_PixelTransform`, which automatically applies transformations to `env.render()` outputs and any `info['pixels*']` entries (supporting multi-camera setups like `pixels.camera0`).

### Available Augmentation Wrappers

The library includes the following visual disturbance wrappers, each implementing specific domain randomization techniques:

- **NoiseWrapper**: Adds step-dependent Gaussian noise with schedulable standard deviation (supports `linear`, `cosine`, `exponential`, `sinusoidal`, or custom callables)
- **ColorJitterWrapper**: Applies random brightness, contrast, saturation, and hue shifts sampled once per episode
- **BlurWrapper**: Applies Gaussian blur with configurable kernel size to simulate motion blur or low-quality sensors
- **OcclusionWrapper**: Renders random rectangular patches with solid colors to simulate missing data
- **MovingPatchWrapper**: Creates drifting colored patches that move smoothly across consecutive frames
- **RandomShiftWrapper**: Implements DrQ-style random translation with replicate padding
- **CutoutWrapper**: Masks random rectangles each frame for regularization
- **RandomConvWrapper**: Applies freshly sampled random convolutional filters per episode for extreme domain randomization
- **GrayscaleWrapper**: Converts RGB to grayscale with optional channel broadcasting
- **ResolutionWrapper**: Downsamples then upsamples to simulate low-resolution sensors
- **ChromaKeyWrapper**: Replaces keyed colors (e.g., green-screen) with static images or looping video backgrounds

All wrappers maintain the standard Gymnasium API (`reset()`, `step()`, `render()`), allowing unlimited compositional stacking either before `MegaWrapper` or within custom pipelines.

## Implementation Examples

### Basic MegaWrapper Configuration

The following example demonstrates wrapping a standard CartPole environment with pixel rendering and noise augmentation:

```python
import gymnasium as gym
from stable_worldmodel.wrapper.default import MegaWrapper
from stable_worldmodel.wrapper.visual import NoiseWrapper

# Create base environment

base_env = gym.make("CartPole-v1", render_mode="rgb_array")

# Add exponential decay noise schedule

noisy_env = NoiseWrapper(
    env=base_env,
    std=lambda step: 5.0 * (0.99 ** step),
    seed=42,
)

# Apply MegaWrapper with 84x84 resolution

env = MegaWrapper(
    env=noisy_env,
    image_shape=(84, 84),
    pixels_transform=None,
    goal_transform=None,
    required_keys=None,
    separate_goal=True,
    image_resample="bilinear",
    add_pixels=True,
)

obs, info = env.reset()
print(f"Pixel shape: {info['pixels'].shape}")  # Output: (84, 84, 3)

```

### Stacking Visual Augmentations

Visual wrappers compose sequentially. This example adds color jittering and blurring:

```python
from stable_worldmodel.wrapper.visual import ColorJitterWrapper, BlurWrapper

# Stack augmentations

env = ColorJitterWrapper(env, brightness=0.3, contrast=0.3, 
                         saturation=0.3, hue=0.1, seed=42)
env = BlurWrapper(env, kernel=3, sigma=0.0)

obs, info = env.reset()

# info["pixels"] now contains color-jittered, blurred frames

```

### Integrating Torchvision Transforms

For normalization and tensor conversion, pass `torchvision` transforms directly to `MegaWrapper`:

```python
import torchvision.transforms as T
from stable_worldmodel.wrapper.default import MegaWrapper

transform = T.Compose([
    T.RandomHorizontalFlip(p=0.5),
    T.ToTensor(),  # Converts to C×H×W float tensor

    T.Normalize(mean=[0.5]*3, std=[0.5]*3),
])

env = MegaWrapper(
    env=gym.make("Breakout-v0", render_mode="rgb_array"),
    image_shape=(84, 84),
    pixels_transform=transform,
    add_pixels=True,
)

```

### Background Substitution with ChromaKeyWrapper

Replace green-screen backgrounds with video loops for synthetic data generation:

```python
from stable_worldmodel.wrapper.visual import ChromaKeyWrapper

env = ChromaKeyWrapper(
    env=gym.make("Custom3DEnv", render_mode="rgb_array"),
    key_color=(0, 255, 0),
    media="background.mp4",
    tolerance=30.0,
)

```

## Summary

- The **MegaWrapper preprocessing pipeline** in [`stable_worldmodel/wrapper/default.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/default.py) standardizes Gymnasium environments through a four-stage wrapping process: pixel rendering, info restructuring, validation, and goal resizing.
- **AddPixelsWrapper** handles frame rendering and optional torchvision-compatible transforms via the `pixels_transform` parameter.
- **Visual augmentation wrappers** in [`stable_worldmodel/wrapper/visual.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/visual.py) inherit from `_PixelTransform` and support scheduled noise, color jittering, blurring, and chroma-key effects.
- All wrappers maintain Gymnasium API compatibility, allowing flexible stacking before or after the main preprocessing pipeline.

## Frequently Asked Questions

### What is the purpose of MegaWrapper in stable-worldmodel?

**MegaWrapper serves as the central preprocessing orchestrator** that converts any Gymnasium environment into a consistent pixel-based interface. It ensures all observations, rewards, and termination flags reside in the `info` dictionary while guaranteeing that pixel data exists under standardized keys like `pixels`, making downstream world model training code environment-agnostic.

### How does AddPixelsWrapper handle image transforms?

**AddPixelsWrapper applies transforms after resizing but before storage.** According to [`default.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/default.py), it first renders the environment, resizes the output to `image_shape` using the specified resampling method (bilinear or nearest), then optionally applies the `pixels_transform` callable (typically a torchvision transform) before inserting the result into `info['pixels']`.

### Can I stack multiple visual augmentation wrappers?

**Yes, visual wrappers compose arbitrarily.** Each augmentation in [`stable_worldmodel/wrapper/visual.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/visual.py) inherits from Gymnasium's wrapper base and the internal `_PixelTransform` class, allowing you to chain `NoiseWrapper`, `ColorJitterWrapper`, `BlurWrapper`, and others either before `MegaWrapper` (affecting raw renders) or after (affecting resized pixels).

### Where are the wrapper classes defined in the source code?

**The core preprocessing logic resides in two files:** `MegaWrapper`, `AddPixelsWrapper`, and the info-management wrappers are implemented in [`stable_worldmodel/wrapper/default.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/default.py), while pixel-level augmentations like `NoiseWrapper` and `ChromaKeyWrapper` are defined in [`stable_worldmodel/wrapper/visual.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/visual.py). Utility functions such as `get_in` for nested dictionary access are located in [`stable_worldmodel/utils.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/utils.py).