# How to Debug Memory Issues with the ErrorBuffer Utility in LongLive

> Debug memory issues in LongLive using the ErrorBuffer utility. Monitor stats, verify shard size, and ensure buffer warmup for efficient memory management.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**To debug memory issues with the ErrorBuffer utility in LongLive, monitor bucket statistics via `stats()`, verify that `shard_size` matches your sequence-parallel size, and ensure `er_buffer_warmup_iter` is greater than zero to allow buffer population.**

The `ErrorBuffer` utility in the NVlabs/LongLive repository implements error-recycling for diffusion training by storing prediction errors on CPU memory and re-injecting them during SVI-style self-forcing. When configured incorrectly, this bucketed ring buffer can trigger CPU out-of-memory errors or fail to populate entirely, stalling the error-recycling mechanism. Understanding how to debug memory issues with the ErrorBuffer utility requires tracing its construction in [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py), its sharding logic across sequence-parallel ranks, and its warm-up phase behavior.

## Understanding ErrorBuffer Architecture and Sharding

### 1-D vs. 2-D Bucket Modes

`ErrorBuffer` operates in two distinct layout modes defined in [`utils/error_buffer.py`](https://github.com/NVlabs/LongLive/blob/main/utils/error_buffer.py). In **1-D mode** (enabled when `num_blocks ≤ 0`), buckets are keyed solely by diffusion timestep, creating a flat structure across the temporal dimension. In **2-D mode** (enabled when `num_blocks > 0`), buckets are keyed by both global block position and timestep, creating a grid layout that maps specific spatial positions to specific noise levels.

### Sequence-Parallel Sharding Strategy

To prevent CPU memory exhaustion, the buffer shards buckets across sequence-parallel (SP) ranks. Each rank only allocates timestep buckets where `t_bucket % shard_size == shard_rank`. During `CausalDiffusion.__init__` in [`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py) (lines 101-106), the buffer is constructed via `build_error_buffer()` with specific sharding parameters:

```python
self.error_buffer = build_error_buffer(
    cfg_dict,
    num_blocks=self.er_num_blocks,
    global_block_offset=self.er_block_offset,
    shard_rank=er_shard_rank,
    shard_size=er_shard_size,
)

```

The `er_num_blocks` represents the local block count for the rank (global blocks divided by `sp_size`), while `global_block_offset` tracks the absolute block range for logging purposes.

## Identifying Memory-Related Symptoms

**CPU RAM Spikes Leading to OOM**
Excessive memory consumption typically occurs when `max_size_per_bucket` (default 50) or `num_buckets` is configured too high for available CPU RAM. Verify memory pressure by printing `self.error_buffer.stats()` and inspecting the `total_entries` count.

**Few or Zero Filled Buckets**
If `stats()` reports `filled_buckets` as `0/40` or similar low ratios, the warm-up phase may be too short or `er_buffer_warmup_iter` is set to 0. Confirm that `global_step` reaches `er_buffer_warmup_iter` before the training loop transitions from warm-up to regular operation.

**Rank Mismatches in Buffer Population**
When some ranks report "covers GLOBAL blocks" in logs while others show empty buffers, inspect the sharding configuration. Ensure `shard_size` equals the sequence-parallel size and that `sp_size` divides the global block count evenly. Check the startup logs generated in [`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py) (lines 151-162) for per-rank block ranges.

**Checkpoint Loading Failures**
`ErrorBuffer.load_state_dict` raises "Refusing to load" errors when checkpoints contain mismatched `global_block_offset` values. Resolve by ensuring each SP rank loads its own per-rank checkpoint, or pass `strict_offset=False` to the load function.

**"Bucket Not Owned" Warnings**
Repeated warnings indicate distributed initialization failures where `shard_size > 1` but ranks are not properly initialized. Verify `torch.distributed.is_initialized()` returns `True` before training begins.

## Debugging Checklist and Procedures

1. **Print Buffer Statistics Every N Steps**
   ```python
   if step % 100 == 0:
       print("[Debug] ErrorBuffer stats:", self.error_buffer.stats())
   ```

   Monitor `filled_buckets` and `total_entries` to confirm population.

2. **Validate Sharding Configuration**
   ```python
   assert self.error_buffer.shard_size == self.sequence_parallel_size, \
          f"Shard size {self.error_buffer.shard_size} != SP size {self.sequence_parallel_size}"
   ```

3. **Check Warm-up Execution Status**
   ```python
   print("[Debug] Warm-up phase active:",
         self.error_buffer and global_step <= self.er_buffer_warmup_iter)
   ```

4. **Inspect Per-Rank Block Offsets**
   Review the startup logs from [`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py) (lines 151-162) to verify each rank's `er_num_blocks` and `er_block_offset` align with expectations.

5. **Force a Minimal Buffer for Testing**
   Temporarily reduce memory footprint in your config to isolate issues:
   ```yaml
   error_recycling:
     enabled: true
     num_buckets: 8
     buffer_size_per_bucket: 5
   ```

   Run a short training iteration and verify that `filled_buckets` grows without triggering OOM.

## Practical Code Examples

### Creating a Minimal Buffer for Debugging

```python
from utils.error_buffer import build_error_buffer

# Small, deterministic buffer for unit tests

debug_buf = build_error_buffer(
    config={"num_buckets": 4, "buffer_size_per_bucket": 2},
    num_blocks=0,                # 1-D mode

    global_block_offset=0,
    shard_rank=0,
    shard_size=1,
)

```

### Manually Adding and Sampling Errors

```python
import torch

# Create dummy error tensor matching expected shape

dummy_err = torch.randn(1, 5, 3, 8, 8)   # (block, C, H, W)

# Store in bucket corresponding to timestep 123

debug_buf.add(dummy_err, timestep_index=123)

# Sample an entry on CPU with specific dtype

sampled = debug_buf.sample(timestep_index=123, device="cpu", dtype=torch.float32)
print("Sampled shape:", sampled.shape)

```

### Inspecting Runtime Statistics

```python
print("Current stats:", debug_buf.stats())

# Example output: {'total_added': 1, 'filled_buckets': '1/4', 'total_entries': 1}

```

### Integration Context in Training Loops

During the warm-up phase in `generator_loss`, the buffer is populated via `_gather_errors_for_buffer` (which performs `all_gather` across DP or WORLD groups) and `_apply_gathered_items` (which randomly selects entries for local storage). During training, errors are injected at three points:
- **Latent prefix** via `_inject_error_buffer` (`E_img`)
- **Latent before noise** via `_inject_latent_error_buffer` (`E_vid`)
- **Noise tensor** via `_inject_noise_error_buffer` (`E_noise`)

## Key Source Files

- **[`utils/error_buffer.py`](https://github.com/NVlabs/LongLive/blob/main/utils/error_buffer.py)**: Implements `ErrorBuffer`, bucket logic, sharding, serialization, and `build_error_buffer`.
- **[`model/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/model/diffusion.py)**: Contains `CausalDiffusion` class, buffer instantiation, injection helpers, and warm-up logic.
- **[`trainer/diffusion.py`](https://github.com/NVlabs/LongLive/blob/main/trainer/diffusion.py)**: Logs per-rank block coverage and orchestrates distributed groups for sequence-parallel and data-parallel training.
- **[`utils/distributed.py`](https://github.com/NVlabs/LongLive/blob/main/utils/distributed.py)**: Provides helper functions for establishing DP/SP groups used during buffer all-gather operations.

## Summary

- **Monitor `stats()`** to verify buffer population and identify memory pressure through `total_entries`.
- **Match sharding parameters** by ensuring `shard_size` equals your sequence-parallel size and `global_block_offset` aligns across checkpoints.
- **Enable warm-up** by setting `er_buffer_warmup_iter > 0` to allow the buffer to collect initial errors via `_gather_errors_for_buffer`.
- **Reduce bucket dimensions** (`num_buckets` and `buffer_size_per_bucket`) if encountering CPU OOM errors.
- **Load checkpoints carefully** by verifying `global_block_offset` matches or using `strict_offset=False` when resuming distributed training.

## Frequently Asked Questions

### What causes ErrorBuffer to consume excessive CPU memory?

Excessive memory usage occurs when `max_size_per_bucket` (default 50) or `num_buckets` is configured too high relative to available CPU RAM and the number of sequence-parallel ranks. Each rank allocates buckets for its shard of timesteps, so increasing SP size reduces per-rank memory, but large bucket capacities can still overwhelm high-resolution video training. Reduce `buffer_size_per_bucket` or increase sharding to mitigate this.

### Why are my ErrorBuffer buckets not filling during training?

Buckets remain empty if the warm-up phase is disabled or too short. The buffer only receives entries during the first `er_buffer_warmup_iter` steps via `_apply_gathered_items`. Verify that `global_step` reaches this threshold and that `_gather_errors_for_buffer` is actually being called in `generator_loss`. Also check that `sample_pos_any_t` or `sample` is not being called before warm-up completes.

### How do I fix "Refusing to load" checkpoint errors in ErrorBuffer?

This error occurs in `load_state_dict` when the saved checkpoint's `global_block_offset` does not match the current rank's offset. This typically happens when loading a single checkpoint across multiple SP ranks that were saved with different offsets. Solution: ensure each SP rank loads its own per-rank checkpoint file, or instantiate the buffer with `strict_offset=False` to override the safety check.

### Can I disable ErrorBuffer sharding to simplify debugging?

While you cannot entirely disable sharding without modifying the source, you can simulate a non-sharded configuration by setting `shard_size=1` and `shard_rank=0` when calling `build_error_buffer`. This forces the rank to own all timestep buckets, effectively disabling the distribution of memory across ranks. Note that this will increase CPU memory consumption on that single rank, so reduce `num_buckets` and `buffer_size_per_bucket` proportionally when using this debugging approach.