How to Debug Memory Issues with the ErrorBuffer Utility in LongLive

To debug memory issues with the ErrorBuffer utility in LongLive, monitor bucket statistics via stats(), verify that shard_size matches your sequence-parallel size, and ensure er_buffer_warmup_iter is greater than zero to allow buffer population.

The ErrorBuffer utility in the NVlabs/LongLive repository implements error-recycling for diffusion training by storing prediction errors on CPU memory and re-injecting them during SVI-style self-forcing. When configured incorrectly, this bucketed ring buffer can trigger CPU out-of-memory errors or fail to populate entirely, stalling the error-recycling mechanism. Understanding how to debug memory issues with the ErrorBuffer utility requires tracing its construction in model/diffusion.py, its sharding logic across sequence-parallel ranks, and its warm-up phase behavior.

Understanding ErrorBuffer Architecture and Sharding

1-D vs. 2-D Bucket Modes

ErrorBuffer operates in two distinct layout modes defined in utils/error_buffer.py. In 1-D mode (enabled when num_blocks ≤ 0), buckets are keyed solely by diffusion timestep, creating a flat structure across the temporal dimension. In 2-D mode (enabled when num_blocks > 0), buckets are keyed by both global block position and timestep, creating a grid layout that maps specific spatial positions to specific noise levels.

Sequence-Parallel Sharding Strategy

To prevent CPU memory exhaustion, the buffer shards buckets across sequence-parallel (SP) ranks. Each rank only allocates timestep buckets where t_bucket % shard_size == shard_rank. During CausalDiffusion.__init__ in model/diffusion.py (lines 101-106), the buffer is constructed via build_error_buffer() with specific sharding parameters:

self.error_buffer = build_error_buffer(
    cfg_dict,
    num_blocks=self.er_num_blocks,
    global_block_offset=self.er_block_offset,
    shard_rank=er_shard_rank,
    shard_size=er_shard_size,
)

The er_num_blocks represents the local block count for the rank (global blocks divided by sp_size), while global_block_offset tracks the absolute block range for logging purposes.

CPU RAM Spikes Leading to OOM Excessive memory consumption typically occurs when max_size_per_bucket (default 50) or num_buckets is configured too high for available CPU RAM. Verify memory pressure by printing self.error_buffer.stats() and inspecting the total_entries count.

Few or Zero Filled Buckets If stats() reports filled_buckets as 0/40 or similar low ratios, the warm-up phase may be too short or er_buffer_warmup_iter is set to 0. Confirm that global_step reaches er_buffer_warmup_iter before the training loop transitions from warm-up to regular operation.

Rank Mismatches in Buffer Population When some ranks report "covers GLOBAL blocks" in logs while others show empty buffers, inspect the sharding configuration. Ensure shard_size equals the sequence-parallel size and that sp_size divides the global block count evenly. Check the startup logs generated in trainer/diffusion.py (lines 151-162) for per-rank block ranges.

Checkpoint Loading Failures ErrorBuffer.load_state_dict raises "Refusing to load" errors when checkpoints contain mismatched global_block_offset values. Resolve by ensuring each SP rank loads its own per-rank checkpoint, or pass strict_offset=False to the load function.

"Bucket Not Owned" Warnings Repeated warnings indicate distributed initialization failures where shard_size > 1 but ranks are not properly initialized. Verify torch.distributed.is_initialized() returns True before training begins.

Debugging Checklist and Procedures

  1. Print Buffer Statistics Every N Steps

    if step % 100 == 0:
        print("[Debug] ErrorBuffer stats:", self.error_buffer.stats())

    Monitor filled_buckets and total_entries to confirm population.

  2. Validate Sharding Configuration

    assert self.error_buffer.shard_size == self.sequence_parallel_size, \
           f"Shard size {self.error_buffer.shard_size} != SP size {self.sequence_parallel_size}"
  3. Check Warm-up Execution Status

    print("[Debug] Warm-up phase active:",
          self.error_buffer and global_step <= self.er_buffer_warmup_iter)
  4. Inspect Per-Rank Block Offsets Review the startup logs from trainer/diffusion.py (lines 151-162) to verify each rank's er_num_blocks and er_block_offset align with expectations.

  5. Force a Minimal Buffer for Testing Temporarily reduce memory footprint in your config to isolate issues:

    error_recycling:
      enabled: true
      num_buckets: 8
      buffer_size_per_bucket: 5

    Run a short training iteration and verify that filled_buckets grows without triggering OOM.

Practical Code Examples

Creating a Minimal Buffer for Debugging

from utils.error_buffer import build_error_buffer

# Small, deterministic buffer for unit tests

debug_buf = build_error_buffer(
    config={"num_buckets": 4, "buffer_size_per_bucket": 2},
    num_blocks=0,                # 1-D mode

    global_block_offset=0,
    shard_rank=0,
    shard_size=1,
)

Manually Adding and Sampling Errors

import torch

# Create dummy error tensor matching expected shape

dummy_err = torch.randn(1, 5, 3, 8, 8)   # (block, C, H, W)

# Store in bucket corresponding to timestep 123

debug_buf.add(dummy_err, timestep_index=123)

# Sample an entry on CPU with specific dtype

sampled = debug_buf.sample(timestep_index=123, device="cpu", dtype=torch.float32)
print("Sampled shape:", sampled.shape)

Inspecting Runtime Statistics

print("Current stats:", debug_buf.stats())

# Example output: {'total_added': 1, 'filled_buckets': '1/4', 'total_entries': 1}

Integration Context in Training Loops

During the warm-up phase in generator_loss, the buffer is populated via _gather_errors_for_buffer (which performs all_gather across DP or WORLD groups) and _apply_gathered_items (which randomly selects entries for local storage). During training, errors are injected at three points:

  • Latent prefix via _inject_error_buffer (E_img)
  • Latent before noise via _inject_latent_error_buffer (E_vid)
  • Noise tensor via _inject_noise_error_buffer (E_noise)

Key Source Files

  • utils/error_buffer.py: Implements ErrorBuffer, bucket logic, sharding, serialization, and build_error_buffer.
  • model/diffusion.py: Contains CausalDiffusion class, buffer instantiation, injection helpers, and warm-up logic.
  • trainer/diffusion.py: Logs per-rank block coverage and orchestrates distributed groups for sequence-parallel and data-parallel training.
  • utils/distributed.py: Provides helper functions for establishing DP/SP groups used during buffer all-gather operations.

Summary

  • Monitor stats() to verify buffer population and identify memory pressure through total_entries.
  • Match sharding parameters by ensuring shard_size equals your sequence-parallel size and global_block_offset aligns across checkpoints.
  • Enable warm-up by setting er_buffer_warmup_iter > 0 to allow the buffer to collect initial errors via _gather_errors_for_buffer.
  • Reduce bucket dimensions (num_buckets and buffer_size_per_bucket) if encountering CPU OOM errors.
  • Load checkpoints carefully by verifying global_block_offset matches or using strict_offset=False when resuming distributed training.

Frequently Asked Questions

What causes ErrorBuffer to consume excessive CPU memory?

Excessive memory usage occurs when max_size_per_bucket (default 50) or num_buckets is configured too high relative to available CPU RAM and the number of sequence-parallel ranks. Each rank allocates buckets for its shard of timesteps, so increasing SP size reduces per-rank memory, but large bucket capacities can still overwhelm high-resolution video training. Reduce buffer_size_per_bucket or increase sharding to mitigate this.

Why are my ErrorBuffer buckets not filling during training?

Buckets remain empty if the warm-up phase is disabled or too short. The buffer only receives entries during the first er_buffer_warmup_iter steps via _apply_gathered_items. Verify that global_step reaches this threshold and that _gather_errors_for_buffer is actually being called in generator_loss. Also check that sample_pos_any_t or sample is not being called before warm-up completes.

How do I fix "Refusing to load" checkpoint errors in ErrorBuffer?

This error occurs in load_state_dict when the saved checkpoint's global_block_offset does not match the current rank's offset. This typically happens when loading a single checkpoint across multiple SP ranks that were saved with different offsets. Solution: ensure each SP rank loads its own per-rank checkpoint file, or instantiate the buffer with strict_offset=False to override the safety check.

Can I disable ErrorBuffer sharding to simplify debugging?

While you cannot entirely disable sharding without modifying the source, you can simulate a non-sharded configuration by setting shard_size=1 and shard_rank=0 when calling build_error_buffer. This forces the rank to own all timestep buckets, effectively disabling the distribution of memory across ranks. Note that this will increase CPU memory consumption on that single rank, so reduce num_buckets and buffer_size_per_bucket proportionally when using this debugging approach.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →