Configuring Action Spaces for Model-Based Planning Solvers in Stable‑WorldModel
All model‑based planning solvers in Stable‑WorldModel require explicit initialization via the configure() method to validate Gymnasium action spaces and compute dimensionality using the PlanConfig.action_block parameter.
Stable‑WorldModel provides several planning solvers—including PGDSolver, MPPIDolver, and CEMSolver—that must be told how to interpret environment actions before generating plans. Configuring action spaces correctly ensures the solver handles discrete or continuous actions according to the environment specification. This registration happens through a unified contract defined in the base solver interface, where action_space, parallel environment counts, and planning parameters are formally bound.
The Solver Configuration Contract
Every solver implements the configure method defined in stable_worldmodel/solver/solver.py. This method serves as the initialization gate that validates compatibility between the solver algorithm and the environment's action representation.
Method Signature and Parameters
The configuration interface enforces keyword‑only arguments to prevent ordering errors:
def configure(self, *, action_space: gym.Space, n_envs: int, config: Any) -> None
action_space– A Gymnasium space (gymnasium.spaces.Boxorgymnasium.spaces.Discrete) describing the environment's action structure.n_envs– Integer specifying the number of parallel environments to plan for simultaneously.config– APlanConfiginstance (defined instable_worldmodel/policy.py) containinghorizon,receding_horizon, andaction_block.
Action Space Compatibility by Solver Type
Each solver extracts dimensionality differently and enforces specific constraints on the action space type. Passing an incompatible space results in immediate assertions or warnings.
Discrete-Only Planning with PGDSolver
PGDSolver (projected‑gradient descent) strictly requires discrete action spaces. In stable_worldmodel/solver/pgd.py (lines 55‑68), the configure method raises an AssertionError if a continuous Box space is supplied:
assert isinstance(action_space, Discrete)
self._action_dim = int(np.prod(action_space.shape[1:]))
self._action_simplex_dim = int(action_space.n)
The solver calculates the raw action dimension from the shape and stores the number of discrete categories (action_space.n) for simplex projection during optimization.
Continuous Action Support in MPPIDolver
MPPIDSolver (model‑predictive path integral) primarily targets continuous control. As implemented in stable_worldmodel/solver/mppi.py (lines 57‑70), it accepts Box spaces natively and logs a warning when Discrete is passed:
self._action_dim = int(np.prod(action_space.shape[1:]))
if not isinstance(action_space, Box):
logging.warning("MPPI works best with continuous action spaces")
While the solver proceeds with discrete spaces, performance degrades because MPPI is designed for Gaussian action sampling in continuous domains.
Universal Solvers: CEM and Variants
CEMSolver, CategoricalCEMSolver, LagrangianSolver, and GDSolver handle both Box and Discrete spaces. In stable_worldmodel/solver/cem.py (lines 57‑66), the configuration extracts dimensionality without type restrictions:
self._action_dim = int(np.prod(action_space.shape[1:]))
These solvers rely on the PlanConfig object to finalize the effective planning dimension.
Computing Flattened Action Dimensions
All solvers expose an action_dim property that returns the total flattened size of the action sequence. The calculation combines the raw action dimension with the block size specified in the configuration:
total_dim = self._action_dim * self._config.action_block
self._action_dim– ForBoxspaces, this equals the product of the shape dimensions; forDiscretespaces, it represents the category count.config.action_block– Groups consecutive timesteps into a single macro‑action, effectively scaling the planning dimension.
The PlanConfig class in stable_worldmodel/policy.py (lines 18‑23) defines these parameters:
class PlanConfig:
def __init__(self, horizon: int, receding_horizon: int, action_block: int = 1):
self.horizon = horizon
self.receding_horizon = receding_horizon
self.action_block = action_block
Step-by-Step Configuration Example
The following pattern demonstrates how to wire a solver into a Stable‑WorldModel pipeline:
import gymnasium as gym
from stable_worldmodel.solver.pgd import PGDSolver
from stable_worldmodel.policy import PlanConfig
# 1. Create environment and extract action space
env = gym.make("CartPole-v1")
action_space = env.action_space # Discrete(2)
# 2. Define planning configuration
cfg = PlanConfig(horizon=10, receding_horizon=5, action_block=2)
# 3. Instantiate solver with a trained world model
solver = PGDSolver(model=my_world_model, n_steps=5)
# 4. Configure action space validation and dimensionality
solver.configure(action_space=action_space, n_envs=1, config=cfg)
# 5. Execute planning step
plan = solver.solve(info_dict=env_info)
If action_space were a Box (continuous), the configure call would raise an AssertionError, requiring you to switch to MPPIDolver or another compatible solver.
Summary
- Unified Interface – All solvers implement
configure(action_space, n_envs, config)defined instable_worldmodel/solver/solver.py. - Type Validation – PGDSolver requires
Discretespaces and asserts this inpgd.py; MPPIDolver prefersBoxspaces and warns on discrete inputs. - Dimensionality Scaling – Solvers compute
action_dim * action_blockto handle macro‑actions over planning horizons. - Configuration Object –
PlanConfiginstable_worldmodel/policy.pysupplieshorizon,receding_horizon, andaction_blockparameters. - Error Handling – Incompatible action spaces trigger immediate, clear errors during configuration rather than during planning.
Frequently Asked Questions
What happens if I pass a Box action space to PGDSolver?
The configure method in stable_worldmodel/solver/pgd.py raises an AssertionError immediately because PGDSolver only supports discrete action spaces for its simplex projection algorithm. You must use a continuous‑compatible solver like MPPIDolver or CEMSolver instead.
How does the action_block parameter affect planning?
action_block groups multiple consecutive timesteps into a single macro‑action, effectively multiplying the action dimensionality by this factor. According to stable_worldmodel/policy.py, this allows planners to optimize over coarser temporal resolutions while maintaining the same horizon length.
Can I use discrete action spaces with MPPIDolver?
Yes, but it is not recommended. While MPPIDolver.configure() in stable_worldmodel/solver/mppi.py accepts discrete spaces, it logs a warning because MPPI relies on continuous Gaussian sampling. Performance typically degrades compared to native discrete solvers like PGDSolver or CategoricalCEMSolver.
Where is the solver configuration interface defined?
The base contract resides in stable_worldmodel/solver/solver.py. Each specific implementation—such as pgd.py, mppi.py, and cem.py—overrides this interface to extract action_dim and validate space types according to the algorithmic requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →