How to Debug Memory Issues with the ErrorBuffer Utility in LongLive
To debug memory issues with the ErrorBuffer utility in LongLive, monitor bucket statistics via stats(), verify that shard_size matches your sequence-parallel size, and ensure er_buffer_warmup_iter is greater than zero to allow buffer population.
The ErrorBuffer utility in the NVlabs/LongLive repository implements error-recycling for diffusion training by storing prediction errors on CPU memory and re-injecting them during SVI-style self-forcing. When configured incorrectly, this bucketed ring buffer can trigger CPU out-of-memory errors or fail to populate entirely, stalling the error-recycling mechanism. Understanding how to debug memory issues with the ErrorBuffer utility requires tracing its construction in model/diffusion.py, its sharding logic across sequence-parallel ranks, and its warm-up phase behavior.
Understanding ErrorBuffer Architecture and Sharding
1-D vs. 2-D Bucket Modes
ErrorBuffer operates in two distinct layout modes defined in utils/error_buffer.py. In 1-D mode (enabled when num_blocks ≤ 0), buckets are keyed solely by diffusion timestep, creating a flat structure across the temporal dimension. In 2-D mode (enabled when num_blocks > 0), buckets are keyed by both global block position and timestep, creating a grid layout that maps specific spatial positions to specific noise levels.
Sequence-Parallel Sharding Strategy
To prevent CPU memory exhaustion, the buffer shards buckets across sequence-parallel (SP) ranks. Each rank only allocates timestep buckets where t_bucket % shard_size == shard_rank. During CausalDiffusion.__init__ in model/diffusion.py (lines 101-106), the buffer is constructed via build_error_buffer() with specific sharding parameters:
self.error_buffer = build_error_buffer(
cfg_dict,
num_blocks=self.er_num_blocks,
global_block_offset=self.er_block_offset,
shard_rank=er_shard_rank,
shard_size=er_shard_size,
)
The er_num_blocks represents the local block count for the rank (global blocks divided by sp_size), while global_block_offset tracks the absolute block range for logging purposes.
Identifying Memory-Related Symptoms
CPU RAM Spikes Leading to OOM
Excessive memory consumption typically occurs when max_size_per_bucket (default 50) or num_buckets is configured too high for available CPU RAM. Verify memory pressure by printing self.error_buffer.stats() and inspecting the total_entries count.
Few or Zero Filled Buckets
If stats() reports filled_buckets as 0/40 or similar low ratios, the warm-up phase may be too short or er_buffer_warmup_iter is set to 0. Confirm that global_step reaches er_buffer_warmup_iter before the training loop transitions from warm-up to regular operation.
Rank Mismatches in Buffer Population
When some ranks report "covers GLOBAL blocks" in logs while others show empty buffers, inspect the sharding configuration. Ensure shard_size equals the sequence-parallel size and that sp_size divides the global block count evenly. Check the startup logs generated in trainer/diffusion.py (lines 151-162) for per-rank block ranges.
Checkpoint Loading Failures
ErrorBuffer.load_state_dict raises "Refusing to load" errors when checkpoints contain mismatched global_block_offset values. Resolve by ensuring each SP rank loads its own per-rank checkpoint, or pass strict_offset=False to the load function.
"Bucket Not Owned" Warnings
Repeated warnings indicate distributed initialization failures where shard_size > 1 but ranks are not properly initialized. Verify torch.distributed.is_initialized() returns True before training begins.
Debugging Checklist and Procedures
-
Print Buffer Statistics Every N Steps
if step % 100 == 0: print("[Debug] ErrorBuffer stats:", self.error_buffer.stats())Monitor
filled_bucketsandtotal_entriesto confirm population. -
Validate Sharding Configuration
assert self.error_buffer.shard_size == self.sequence_parallel_size, \ f"Shard size {self.error_buffer.shard_size} != SP size {self.sequence_parallel_size}" -
Check Warm-up Execution Status
print("[Debug] Warm-up phase active:", self.error_buffer and global_step <= self.er_buffer_warmup_iter) -
Inspect Per-Rank Block Offsets Review the startup logs from
trainer/diffusion.py(lines 151-162) to verify each rank'ser_num_blocksander_block_offsetalign with expectations. -
Force a Minimal Buffer for Testing Temporarily reduce memory footprint in your config to isolate issues:
error_recycling: enabled: true num_buckets: 8 buffer_size_per_bucket: 5Run a short training iteration and verify that
filled_bucketsgrows without triggering OOM.
Practical Code Examples
Creating a Minimal Buffer for Debugging
from utils.error_buffer import build_error_buffer
# Small, deterministic buffer for unit tests
debug_buf = build_error_buffer(
config={"num_buckets": 4, "buffer_size_per_bucket": 2},
num_blocks=0, # 1-D mode
global_block_offset=0,
shard_rank=0,
shard_size=1,
)
Manually Adding and Sampling Errors
import torch
# Create dummy error tensor matching expected shape
dummy_err = torch.randn(1, 5, 3, 8, 8) # (block, C, H, W)
# Store in bucket corresponding to timestep 123
debug_buf.add(dummy_err, timestep_index=123)
# Sample an entry on CPU with specific dtype
sampled = debug_buf.sample(timestep_index=123, device="cpu", dtype=torch.float32)
print("Sampled shape:", sampled.shape)
Inspecting Runtime Statistics
print("Current stats:", debug_buf.stats())
# Example output: {'total_added': 1, 'filled_buckets': '1/4', 'total_entries': 1}
Integration Context in Training Loops
During the warm-up phase in generator_loss, the buffer is populated via _gather_errors_for_buffer (which performs all_gather across DP or WORLD groups) and _apply_gathered_items (which randomly selects entries for local storage). During training, errors are injected at three points:
- Latent prefix via
_inject_error_buffer(E_img) - Latent before noise via
_inject_latent_error_buffer(E_vid) - Noise tensor via
_inject_noise_error_buffer(E_noise)
Key Source Files
utils/error_buffer.py: ImplementsErrorBuffer, bucket logic, sharding, serialization, andbuild_error_buffer.model/diffusion.py: ContainsCausalDiffusionclass, buffer instantiation, injection helpers, and warm-up logic.trainer/diffusion.py: Logs per-rank block coverage and orchestrates distributed groups for sequence-parallel and data-parallel training.utils/distributed.py: Provides helper functions for establishing DP/SP groups used during buffer all-gather operations.
Summary
- Monitor
stats()to verify buffer population and identify memory pressure throughtotal_entries. - Match sharding parameters by ensuring
shard_sizeequals your sequence-parallel size andglobal_block_offsetaligns across checkpoints. - Enable warm-up by setting
er_buffer_warmup_iter > 0to allow the buffer to collect initial errors via_gather_errors_for_buffer. - Reduce bucket dimensions (
num_bucketsandbuffer_size_per_bucket) if encountering CPU OOM errors. - Load checkpoints carefully by verifying
global_block_offsetmatches or usingstrict_offset=Falsewhen resuming distributed training.
Frequently Asked Questions
What causes ErrorBuffer to consume excessive CPU memory?
Excessive memory usage occurs when max_size_per_bucket (default 50) or num_buckets is configured too high relative to available CPU RAM and the number of sequence-parallel ranks. Each rank allocates buckets for its shard of timesteps, so increasing SP size reduces per-rank memory, but large bucket capacities can still overwhelm high-resolution video training. Reduce buffer_size_per_bucket or increase sharding to mitigate this.
Why are my ErrorBuffer buckets not filling during training?
Buckets remain empty if the warm-up phase is disabled or too short. The buffer only receives entries during the first er_buffer_warmup_iter steps via _apply_gathered_items. Verify that global_step reaches this threshold and that _gather_errors_for_buffer is actually being called in generator_loss. Also check that sample_pos_any_t or sample is not being called before warm-up completes.
How do I fix "Refusing to load" checkpoint errors in ErrorBuffer?
This error occurs in load_state_dict when the saved checkpoint's global_block_offset does not match the current rank's offset. This typically happens when loading a single checkpoint across multiple SP ranks that were saved with different offsets. Solution: ensure each SP rank loads its own per-rank checkpoint file, or instantiate the buffer with strict_offset=False to override the safety check.
Can I disable ErrorBuffer sharding to simplify debugging?
While you cannot entirely disable sharding without modifying the source, you can simulate a non-sharded configuration by setting shard_size=1 and shard_rank=0 when calling build_error_buffer. This forces the rank to own all timestep buckets, effectively disabling the distribution of memory across ranks. Note that this will increase CPU memory consumption on that single rank, so reduce num_buckets and buffer_size_per_bucket proportionally when using this debugging approach.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →