Execution Flow of World.collect vs World.evaluate in stable-worldmodel
World.collect records raw trajectories to disk or memory buffers using auto-reset mode, while World.evaluate computes aggregated success metrics and optionally records videos using either auto-reset or wait-mode depending on whether evaluation is episodic or dataset-driven. Both methods leverage the same internal stepping loop but differ fundamentally in their callbacks, reset behavior, and output destinations.
stable-worldmodel (galilai-group/stable-worldmodel) provides a vectorized environment driver in stable_worldmodel.world.World that supports both data collection and policy evaluation. Understanding how collect and evaluate share common infrastructure while serving different purposes is essential for efficient pipeline design.
Common Groundwork: The Core Stepping Loop
Both methods rely on a unified execution foundation implemented in stable_worldmodel/world/world.py.
Environment Initialization and Policy Attachment
When you instantiate World, the constructor (World.__init__) creates a list of identical environment factories, wraps each with MegaWrapper ([stable_worldmodel/wrapper/mega_wrapper.py](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/mega_wrapper.py)), and initializes an EnvPool for vectorized execution. The set_policy method attaches your policy to the environment pool and propagates seeds if the policy exposes them.
The Internal Run Mechanism
The heart of both collect and evaluate is the _run_iter → _run pipeline:
_run_iter: Drives the policy by computing actions via_get_actions, stepping the env pool (self.envs.step), invoking an optionalon_stepcallback, and detecting finished episodes. It yields(env_idx, ep_count)tuples and handles reset logic based on themodeparameter._run: Wraps_run_iterand adds anon_donehook that fires when episodes terminate.
This shared infrastructure ensures consistent stepping behavior while allowing flexible callback injection for different use cases.
World.collect: Recording Raw Trajectories
World.collect is designed to generate and persist interaction data. According to the source code in [stable_worldmodel/world/world.py](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py), its execution flow prioritizes data throughput and memory efficiency.
Setup and Validation
The method first validates that exactly one of path or writer is provided. It then creates a writer context using get_format(format).open_writer(path) and pre-allocates per-environment buffers:
buffers = [defaultdict(list) for _ in range(self.num_envs)]
These buffers temporarily hold episode data before streaming to the writer.
The Collection Callback
collect registers an on_step callback that processes the infos dict from each environment step. For every key containing an np.ndarray or torch.Tensor (ignoring private keys), the callback detaches and copies values into the corresponding buffer list. This columnar buffering strategy converts streamed environment data into episode-wise trajectories.
Execution Flow
- Iteration: Calls
_run_iterwithmode='auto'(automatic reset after each episode termination) and the collection callback. - Yielding: When an episode finishes, the generator yields
(env_idx, _). The buffers for that environment are converted to a plain dict:ep = {k: list(v) for k, v in buffers[env_idx].items()}. - Writing: The episode generator streams to
writer.write_episodes(), ensuring bounded memory usage even with thousands of episodes. - Reset Behavior: Always uses
'auto'mode—terminated environments reset immediately to begin fresh episodes.
World.evaluate: Computing Success Metrics
World.evaluate focuses on measuring policy performance rather than persisting full trajectories. The method signature dispatches to specialized private helpers based on whether you provide a dataset.
Evaluation Mode Dispatch
Located in [world.py](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py), evaluate selects between two modes:
- Episodic evaluation (default): Uses
_evaluatewithreset_mode='auto'. - Dataset-driven evaluation: Uses
_evaluate_from_datasetwithreset_mode='wait'(default), allowing fixed initial states from existing data.
Episodic Evaluation (_evaluate)
This path pre-allocates a results dictionary with zeroed arrays for tracking outcomes. If video recording is requested, it initializes a frames dictionary and registers callbacks:
on_step: Captures the latest pixel frame per environment.on_done: Stores success flags, seeds, and writes per-episode video files.
After calling _run with the selected mode, it computes success_rate from the accumulated results.
Dataset-Driven Evaluation (_evaluate_from_dataset)
Designed for replaying test sets or fixed initial conditions:
- State Extraction: Calls
_extract_init_goalto pull initial and goal states from the dataset. - Environment Seeding: Resets environments to extracted seeds and optionally applies user-provided callables.
- Info Overwrite: Broadcasts initial/goal observations into the current
infos. - Goal Snapshot: Stores a fixed goal reference merged on every step.
- Execution: Runs
_runwithmax_steps=eval_budgetandmode='wait'(environments freeze when terminated rather than auto-resetting). - Metrics: Updates
episode_successeswhen environments reach goals or exhaust budget, then calculates finalsuccess_rate.
Key Differences Between collect and evaluate
| Aspect | World.collect |
World.evaluate |
|---|---|---|
| Primary Goal | Persist raw trajectories to disk or memory. | Compute aggregated performance metrics. |
| Output | Episodes written via Writer (Lance files or ReplayBuffer). |
Dictionary with success_rate, episode_successes, and seeds. |
| Reset Mode | Always 'auto' (immediate reset after termination). |
'auto' for episodic, 'wait' for dataset-driven (frozen terminated envs). |
| Callbacks | on_step buffers all info columns. |
on_step captures video frames; on_done records success and writes video. |
| Underlying Call | _run_iter(..., mode='auto', on_step=on_step) |
_run(..., mode=chosen_mode, on_step=..., on_done=...) |
| Dataset Support | None—purely environment-driven generation. | Full support via _evaluate_from_dataset with custom init/goal states. |
Usage Examples
import stable_worldmodel as swm
# Record 200 expert episodes to a Lance file
world = swm.World('swm/PushT-v1', num_envs=4, image_shape=(64, 64))
world.set_policy(swm.Policy.expert())
world.collect('data/push_expert.lance', episodes=200, seed=42)
# Evaluate a learned policy with video recording
policy = swm.Policy.load('my_policy.pt')
world.set_policy(policy)
metrics = world.evaluate(episodes=100, seed=7, video='videos/')
print(f"Success rate: {metrics['success_rate']:.1f}%")
# Dataset-driven evaluation (replay test set)
test_ds = swm.data.load('test_dataset.lance')
metrics = world.evaluate(
dataset=test_ds,
episodes_idx=[0, 1, 2, 3],
start_steps=[0, 10, 20, 30],
goal_offset=30,
eval_budget=50,
video='eval_videos/'
)
Summary
- Both
collectandevaluateuse the internal_run_iter→_runpipeline instable_worldmodel/world/world.py, but differ in callbacks and reset modes. collectalways runs withmode='auto', buffers trajectory data viaon_step, and streams episodes to aWriterfor persistence.evaluatereturns metrics dictionaries, supports video recording throughon_stepandon_donehooks, and usesmode='wait'for dataset-driven evaluation to preserve fixed initial states.MegaWrapperandEnvPoolprovide the normalized observations and vectorized stepping that both methods consume.
Frequently Asked Questions
What is the difference between auto and wait reset modes?
Auto mode ('auto') immediately resets terminated environments, ensuring continuous data generation. Wait mode ('wait') keeps terminated environments frozen until all environments in the pool finish, which is crucial for dataset-driven evaluation where you need synchronized episode boundaries or fixed initial conditions.
Can I call collect and evaluate on the same World instance sequentially?
Yes. Both methods are stateless with respect to the World instance's configuration, though you should call world.reset() between operations if you need to clear internal buffers or change the policy. The environment pool remains intact between calls, but episode counters and temporary buffers are fresh for each invocation.
How does dataset-driven evaluation differ from episodic evaluation?
Episodic evaluation samples new initial states from the environment's reset distribution and uses auto-reset. Dataset-driven evaluation extracts specific initial states and goals from a provided dataset, overwrites the environment's infos with these fixed states, and uses wait-mode to prevent premature resetting, allowing you to measure success rates against specific test conditions.
Where are trajectories stored when using World.collect?
Trajectories are written through the Writer abstraction defined in stable_worldmodel/data/format.py. By default, collect writes to Lance files on disk when a path is provided, or to an in-memory ReplayBuffer when passed directly. The data includes all numpy array or torch tensor entries from the environment's infos dict, excluding private keys (those starting with underscores).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →