# Execution Flow of World.collect vs World.evaluate in stable-worldmodel

> Understand the difference between World collect and World evaluate execution flow in stable-worldmodel. Learn how they record data and compute metrics for effective RL experimentation.

- Repository: [GalilAI-group/stable-worldmodel](https://github.com/galilai-group/stable-worldmodel)
- Tags: internals
- Published: 2026-05-30

---

**`World.collect` records raw trajectories to disk or memory buffers using auto-reset mode, while `World.evaluate` computes aggregated success metrics and optionally records videos using either auto-reset or wait-mode depending on whether evaluation is episodic or dataset-driven.** Both methods leverage the same internal stepping loop but differ fundamentally in their callbacks, reset behavior, and output destinations.

`stable-worldmodel` (galilai-group/stable-worldmodel) provides a vectorized environment driver in `stable_worldmodel.world.World` that supports both data collection and policy evaluation. Understanding how `collect` and `evaluate` share common infrastructure while serving different purposes is essential for efficient pipeline design.

## Common Groundwork: The Core Stepping Loop

Both methods rely on a unified execution foundation implemented in [`stable_worldmodel/world/world.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py).

### Environment Initialization and Policy Attachment

When you instantiate `World`, the constructor (`World.__init__`) creates a list of identical environment factories, wraps each with `MegaWrapper` ([[`stable_worldmodel/wrapper/mega_wrapper.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/mega_wrapper.py)](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/wrapper/mega_wrapper.py)), and initializes an `EnvPool` for vectorized execution. The `set_policy` method attaches your policy to the environment pool and propagates seeds if the policy exposes them.

### The Internal Run Mechanism

The heart of both `collect` and `evaluate` is the `_run_iter` → `_run` pipeline:

- **`_run_iter`**: Drives the policy by computing actions via `_get_actions`, stepping the env pool (`self.envs.step`), invoking an optional `on_step` callback, and detecting finished episodes. It yields `(env_idx, ep_count)` tuples and handles reset logic based on the `mode` parameter.
- **`_run`**: Wraps `_run_iter` and adds an `on_done` hook that fires when episodes terminate.

This shared infrastructure ensures consistent stepping behavior while allowing flexible callback injection for different use cases.

## World.collect: Recording Raw Trajectories

`World.collect` is designed to generate and persist interaction data. According to the source code in [[`stable_worldmodel/world/world.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py)](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py), its execution flow prioritizes data throughput and memory efficiency.

### Setup and Validation

The method first validates that exactly one of `path` or `writer` is provided. It then creates a writer context using `get_format(format).open_writer(path)` and pre-allocates per-environment buffers:

```python
buffers = [defaultdict(list) for _ in range(self.num_envs)]

```

These buffers temporarily hold episode data before streaming to the writer.

### The Collection Callback

`collect` registers an `on_step` callback that processes the `infos` dict from each environment step. For every key containing an `np.ndarray` or `torch.Tensor` (ignoring private keys), the callback detaches and copies values into the corresponding buffer list. This columnar buffering strategy converts streamed environment data into episode-wise trajectories.

### Execution Flow

1. **Iteration**: Calls `_run_iter` with `mode='auto'` (automatic reset after each episode termination) and the collection callback.
2. **Yielding**: When an episode finishes, the generator yields `(env_idx, _)`. The buffers for that environment are converted to a plain dict: `ep = {k: list(v) for k, v in buffers[env_idx].items()}`.
3. **Writing**: The episode generator streams to `writer.write_episodes()`, ensuring bounded memory usage even with thousands of episodes.
4. **Reset Behavior**: Always uses `'auto'` mode—terminated environments reset immediately to begin fresh episodes.

## World.evaluate: Computing Success Metrics

`World.evaluate` focuses on measuring policy performance rather than persisting full trajectories. The method signature dispatches to specialized private helpers based on whether you provide a dataset.

### Evaluation Mode Dispatch

Located in [[`world.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/world.py)](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py), `evaluate` selects between two modes:

- **Episodic evaluation** (default): Uses `_evaluate` with `reset_mode='auto'`.
- **Dataset-driven evaluation**: Uses `_evaluate_from_dataset` with `reset_mode='wait'` (default), allowing fixed initial states from existing data.

### Episodic Evaluation (`_evaluate`)

This path pre-allocates a `results` dictionary with zeroed arrays for tracking outcomes. If video recording is requested, it initializes a `frames` dictionary and registers callbacks:

- **`on_step`**: Captures the latest pixel frame per environment.
- **`on_done`**: Stores success flags, seeds, and writes per-episode video files.

After calling `_run` with the selected mode, it computes `success_rate` from the accumulated results.

### Dataset-Driven Evaluation (`_evaluate_from_dataset`)

Designed for replaying test sets or fixed initial conditions:

1. **State Extraction**: Calls `_extract_init_goal` to pull initial and goal states from the dataset.
2. **Environment Seeding**: Resets environments to extracted seeds and optionally applies user-provided callables.
3. **Info Overwrite**: Broadcasts initial/goal observations into the current `infos`.
4. **Goal Snapshot**: Stores a fixed goal reference merged on every step.
5. **Execution**: Runs `_run` with `max_steps=eval_budget` and `mode='wait'` (environments freeze when terminated rather than auto-resetting).
6. **Metrics**: Updates `episode_successes` when environments reach goals or exhaust budget, then calculates final `success_rate`.

## Key Differences Between collect and evaluate

| Aspect | `World.collect` | `World.evaluate` |
|--------|----------------|------------------|
| **Primary Goal** | Persist raw trajectories to disk or memory. | Compute aggregated performance metrics. |
| **Output** | Episodes written via `Writer` (Lance files or `ReplayBuffer`). | Dictionary with `success_rate`, `episode_successes`, and `seeds`. |
| **Reset Mode** | Always `'auto'` (immediate reset after termination). | `'auto'` for episodic, `'wait'` for dataset-driven (frozen terminated envs). |
| **Callbacks** | `on_step` buffers all info columns. | `on_step` captures video frames; `on_done` records success and writes video. |
| **Underlying Call** | `_run_iter(..., mode='auto', on_step=on_step)` | `_run(..., mode=chosen_mode, on_step=..., on_done=...)` |
| **Dataset Support** | None—purely environment-driven generation. | Full support via `_evaluate_from_dataset` with custom init/goal states. |

## Usage Examples

```python
import stable_worldmodel as swm

# Record 200 expert episodes to a Lance file

world = swm.World('swm/PushT-v1', num_envs=4, image_shape=(64, 64))
world.set_policy(swm.Policy.expert())
world.collect('data/push_expert.lance', episodes=200, seed=42)

# Evaluate a learned policy with video recording

policy = swm.Policy.load('my_policy.pt')
world.set_policy(policy)
metrics = world.evaluate(episodes=100, seed=7, video='videos/')
print(f"Success rate: {metrics['success_rate']:.1f}%")

# Dataset-driven evaluation (replay test set)

test_ds = swm.data.load('test_dataset.lance')
metrics = world.evaluate(
    dataset=test_ds,
    episodes_idx=[0, 1, 2, 3],
    start_steps=[0, 10, 20, 30],
    goal_offset=30,
    eval_budget=50,
    video='eval_videos/'
)

```

## Summary

- Both `collect` and `evaluate` use the internal `_run_iter` → `_run` pipeline in [`stable_worldmodel/world/world.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/world/world.py), but differ in callbacks and reset modes.
- `collect` always runs with `mode='auto'`, buffers trajectory data via `on_step`, and streams episodes to a `Writer` for persistence.
- `evaluate` returns metrics dictionaries, supports video recording through `on_step` and `on_done` hooks, and uses `mode='wait'` for dataset-driven evaluation to preserve fixed initial states.
- `MegaWrapper` and `EnvPool` provide the normalized observations and vectorized stepping that both methods consume.

## Frequently Asked Questions

### What is the difference between auto and wait reset modes?

**Auto mode** (`'auto'`) immediately resets terminated environments, ensuring continuous data generation. **Wait mode** (`'wait'`) keeps terminated environments frozen until all environments in the pool finish, which is crucial for dataset-driven evaluation where you need synchronized episode boundaries or fixed initial conditions.

### Can I call collect and evaluate on the same World instance sequentially?

Yes. Both methods are stateless with respect to the World instance's configuration, though you should call `world.reset()` between operations if you need to clear internal buffers or change the policy. The environment pool remains intact between calls, but episode counters and temporary buffers are fresh for each invocation.

### How does dataset-driven evaluation differ from episodic evaluation?

Episodic evaluation samples new initial states from the environment's reset distribution and uses auto-reset. Dataset-driven evaluation extracts specific initial states and goals from a provided dataset, overwrites the environment's infos with these fixed states, and uses wait-mode to prevent premature resetting, allowing you to measure success rates against specific test conditions.

### Where are trajectories stored when using World.collect?

Trajectories are written through the `Writer` abstraction defined in [`stable_worldmodel/data/format.py`](https://github.com/galilai-group/stable-worldmodel/blob/main/stable_worldmodel/data/format.py). By default, `collect` writes to Lance files on disk when a `path` is provided, or to an in-memory `ReplayBuffer` when passed directly. The data includes all numpy array or torch tensor entries from the environment's `infos` dict, excluding private keys (those starting with underscores).