Implementing Custom Visual Wrappers for World Models: Occlusion, Chroma‑Key, and Moving‑Patch Guide

The stable‑worldmodel package provides a _PixelTransform base class in stable_worldmodel/wrapper/visual.py that enables custom visual wrappers to intercept RGB frames via render(), reset(), and step() methods, allowing you to apply occlusions, chroma‑key background replacement, and moving patches without modifying the underlying environment logic.

The stable‑worldmodel repository from the GalilAI Group ships with a modular augmentation system designed specifically for world model training pipelines. When you need to implement custom visual wrappers for world models—such as simulating sensor occlusion, green‑screen background substitution, or dynamic visual distractors—you can leverage the internal _PixelTransform abstraction that cleanly separates pixel‑space transformations from environment dynamics.

Architecture of the _PixelTransform Base Class

All visual wrappers inherit from _PixelTransform (defined in [stable_worldmodel/wrapper/visual.py](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/visual.py#L56-L78)). This base class overrides three core interaction points of a Gymnasium environment to guarantee every RGB frame passes through your custom transformation pipeline:

  • render() – Calls the wrapped environment’s render() method, then immediately pipes the returned frame through the wrapper’s _apply() logic.
  • reset() – Returns the standard (obs, info) tuple, but intercepts any info['pixels…'] entries and replaces them with transformed versions.
  • step() – Returns the standard (obs, reward, terminated, truncated, info) tuple, again post‑processing pixel data before it reaches the agent.

Concrete implementations only need to define _apply(frame), a method that receives an RGB array and returns the modified array. Wrappers that require stateful behavior (such as random sampling or motion physics) can additionally override reset() and step() to manage internal state.

OcclusionWrapper: Simulating Sensor Obstruction

OcclusionWrapper (visual.py#L29) draws static rectangular masks over random regions of the frame at the start of each episode. This simulates hardware failures or environmental obstructions that partially block the agent’s view.

Sampling Random Patches

During reset(), the wrapper calls _sample(h, w) to draw num_patches rectangles. Each rectangle’s height and width are sampled as random fractions of the frame dimensions, bounded by the size=(low, high) parameter:

  • y and x define the top‑left origin.
  • ph and pw define the patch height and width in pixels.

These coordinates are cached so every frame in the episode receives the same occlusion pattern.

Applying Static Occlusions

The _apply(frame) method executes once per frame after reset(). It copies the input frame, fills the cached rectangles with a constant color (default 0 for black), and returns the masked image. Because the patch list is only regenerated when the environment resets, the occlusion remains stable throughout the episode, consistent with a fixed sensor failure.

ChromaKeyWrapper: Background Replacement

ChromaKeyWrapper (visual.py#L81) implements green‑screen style substitution. Pixels within a Euclidean distance tolerance of a key_color are replaced by pixels from a static image or a looping video.

Loading Media Assets

The constructor delegates to _load_media(path), which detects the file extension. If the path ends with a video extension, the method loads all frames into a stacked NumPy array (np.stack). Otherwise, it loads a single image. The wrapper stores this media and tracks a frame index for video looping.

Color Distance Masking

Inside _apply(frame), the wrapper computes a Euclidean distance between each pixel and the key_color (specified as an RGB list such as [0, 255, 0]). Pixels falling within tolerance are masked and replaced by the corresponding pixels from the background media. For video backgrounds, _next_frame(h, w) advances the internal index after each retrieval, creating seamless looping footage behind the foreground agent.

MovingPatchWrapper: Dynamic Visual Distractors

MovingPatchWrapper (visual_worldmodel/wrapper/visual.py#L66) differs from OcclusionWrapper by introducing continuous motion. Solid‑color patches drift across the frame at a fixed speed, bouncing off image borders to create synthetic visual noise.

Physics and Border Reflection

During reset, _init_patches(h, w) samples initial positions, sizes, and random motion angles. Velocities are derived from the speed parameter (pixels per step) and the angle.

The _advance() method updates positions using these velocities. When a patch hits a border, the method reflects the position and inverts the corresponding velocity component, ensuring the patch remains within bounds while maintaining constant speed.

Step Synchronization

Unlike static wrappers, MovingPatchWrapper overrides step(action). It first forwards the action to the wrapped environment, then calls _advance() to update patch positions so that motion is synchronized with environment dynamics. Finally, it calls _apply() to render the patches at their new coordinates before returning the observation tuple.

Practical Implementation Examples

The following snippets demonstrate how to instantiate and compose these custom visual wrappers for world models with any Gymnasium environment that returns RGB frames via render().

Basic Occlusion

import gymnasium as gym
from stable_worldmodel.wrapper.visual import OcclusionWrapper

env = gym.make("CartPole-v1", render_mode="rgb_array")
env = OcclusionWrapper(
    env,
    num_patches=2,
    size=(0.05, 0.15),  # 5% to 15% of frame height/width

    color=0,            # Black occlusion

    seed=42
)

Chroma‑Key Background

from stable_worldmodel.wrapper.visual import ChromaKeyWrapper

env = gym.make("HandManipulateBlock-v0", render_mode="rgb_array")
env = ChromaKeyWrapper(
    env,
    key_color=[0, 255, 0],                # Target green

    media="assets/backgrounds/office_scene.mp4",
    tolerance=30
)

Moving Distractor Patches

from stable_worldmodel.wrapper.visual import MovingPatchWrapper

env = gym.make("MiniGrid-Empty-5x5-v0", render_mode="rgb_array")
env = MovingPatchWrapper(
    env,
    num_patches=3,
    size=(0.08, 0.12),
    color=[255, 0, 0],  # Red patches

    speed=4.0,
    seed=123
)

Stacking Multiple Wrappers

Because each wrapper inherits from gym.Wrapper and only touches pixel data after the underlying environment produces it, you can compose them in any order:

env = gym.make("MiniGrid-Empty-8x8-v0", render_mode="rgb_array")

# Apply occlusion first

env = OcclusionWrapper(env, num_patches=1, size=(0.1, 0.2), color=0, seed=1)

# Add drifting patches on top

env = MovingPatchWrapper(env, num_patches=2, size=(0.05, 0.1), color=200, speed=3, seed=2)

# Finally replace green background

env = ChromaKeyWrapper(env, key_color=[0, 255, 0], media="assets/bg.png", tolerance=20)

The order matters: each wrapper processes the frame produced by the previous layer in the stack.

Summary

  • _PixelTransform in stable_worldmodel/wrapper/visual.py provides the architectural contract for all visual augmentations, ensuring wrappers intercept frames at render(), reset(), and step() without altering environment logic.
  • OcclusionWrapper generates static random rectangles at episode boundaries, simulating fixed sensor obstructions.
  • ChromaKeyWrapper replaces pixels matching a target color with frames from an image or video stream, enabling dynamic background substitution.
  • MovingPatchWrapper implements physics‑based patch movement synchronized with environment steps, creating continuous visual distractors.
  • All wrappers can be stacked arbitrarily because they respect the standard Gymnasium API and only modify pixel buffers.

Frequently Asked Questions

Can I stack these visual wrappers with other Gymnasium wrappers?

Yes. Because OcclusionWrapper, ChromaKeyWrapper, and MovingPatchWrapper inherit from gym.Wrapper and only transform pixel data returned by render(), they compose safely with observation‑preprocessing wrappers like FrameStack or GrayScaleObservation. Ensure visual wrappers are applied after environment creation but before observation‑space‑altering wrappers if those wrappers inspect pixel shapes.

How do I specify custom colors for occlusions or moving patches?

Both OcclusionWrapper and MovingPatchWrapper accept a color parameter. For grayscale occlusions, pass an integer such as 0 (black) or 255 (white). For colored moving patches, pass an RGB list such as [255, 0, 0] for red or [0, 0, 255] for blue. The wrapper broadcasts this color across the entire patch area.

Does the wrapper modify the observation space returned by the environment?

No. The visual wrappers modify the pixel data stored in info dictionaries (typically under keys like info['pixels'] or info['rgb']) and the array returned by render(). They do not change the observation_space attribute of the base environment, ensuring compatibility with world models that expect the original observation structure.

What video formats are supported by ChromaKeyWrapper?

The _load_media method uses standard image I/O libraries to detect file extensions. It explicitly checks for video extensions to trigger np.stack frame loading. While the source code handles generic video paths, you should verify that your specific format (commonly .mp4, .avi, or .mov) is readable by the underlying NumPy and imageio stack used in your Python environment. For production stability, .mp4 with H.264 encoding is recommended.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →