How to Configure the Offline Rendering Pipeline for Very Long Videos (25,000+ Frames)
To configure the offline rendering pipeline for very long videos in LingBot-Map, implement chunked rendering with --chunk-size limits, stream frames directly to the encoder to bypass PNG intermediates, and tune FPS and resolution in the YAML configuration to handle sequences exceeding 25,000 frames without exhausting RAM or disk resources.
LingBot-Map is an open-source framework for large-scale RGB-D scene reconstruction and novel view synthesis. When processing massive video sequences spanning 25,000 frames or more—potentially reaching millions of frames—you must configure the offline rendering pipeline to prevent memory overflow and filesystem saturation. The repository provides specific mechanisms in demo_render/process_videos.sh and the Python rendering modules to handle these extreme workloads through segmentation, streaming, and parallelization strategies.
Understanding the Core Rendering Components
The offline rendering pipeline in Robbyant/lingbot-map consists of three tightly integrated components that handle different stages of video generation.
Scene construction (demo_render/rgbd_render/scene.py) loads the dataset, builds the point-cloud representation, and prepares camera trajectories for the entire sequence. Frame rendering (demo_render/rgbd_render/renderer.py) executes the GPU-accelerated rendering loop, generating individual RGB frames, depth maps, or combined visualizations. Video encoding (demo_render/rgbd_render/video.py) consumes these frames—either from disk or memory—and compresses them into MP4 containers using FFmpeg or OpenCV backends.
For sequences exceeding 25,000 frames, keeping all rendered frames in RAM or writing millions of individual PNG files creates unsustainable memory pressure and filesystem load. The pipeline addresses this through configurable chunking and zero-copy streaming mechanisms.
Implementing Chunk-Wise Rendering
The most reliable method for handling extremely long videos is processing them in manageable segments rather than attempting to render the entire sequence in a single pass. The repository provides demo_render/process_videos.sh to automate this segmentation.
Configure the chunk size based on your available VRAM and storage throughput. For a 2-million-frame video, process 20,000-frame chunks:
./demo_render/process_videos.sh \
--scene-file path/to/scene.yaml \
--output-dir ./rendered_chunks \
--chunk-size 20000 \
--fps 30
The --chunk-size parameter determines how many frames the renderer processes before invoking the encoder. After each chunk finishes encoding to MP4, the script deletes temporary PNG files to maintain modest disk usage. This approach caps memory consumption to the size of one chunk rather than the entire video duration.
Zero-Copy Streaming to Bypass Disk I/O
When your GPU can generate frames faster than the storage subsystem can write PNG files, you can eliminate the intermediate filesystem operations entirely. The demo_render/rgbd_render/video.py module provides the foundation for this optimization through the encode_video function, but you must implement a memory-based variant for streaming workflows.
Create a new utility module at demo_render/rgbd_render/streaming_utils.py with the following implementation:
import cv2
import numpy as np
from demo_render.rgbd_render.video import colorize_depth
def encode_video_from_memory(frames: np.ndarray, out_path: str, fps: int = 30):
"""
Encode a NumPy array of RGB frames (S, H, W, 3) to MP4 without writing PNGs.
"""
h, w = frames[0].shape[:2]
writer = cv2.VideoWriter(
out_path,
cv2.VideoWriter_fourcc(*'mp4v'),
fps,
(w, h)
)
for img in frames:
bgr = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)
writer.write(bgr)
writer.release()
def encode_depth_from_memory(
depths: np.ndarray,
out_path: str,
fps: int,
depth_range: tuple,
colormap: str = 'turbo'
):
"""
Colour-map depth frames on-the-fly and encode them.
"""
h, w = depths[0].shape[:2]
writer = cv2.VideoWriter(
out_path,
cv2.VideoWriter_fourcc(*'mp4v'),
fps,
(w, h)
)
for d in depths:
bgr = colorize_depth(d, *depth_range, colormap=colormap)
writer.write(bgr)
writer.release()
Integrate these helpers with the viser_wrapper renderer in a batch script:
from lingbot_map.vis import viser_wrapper
from demo_render.rgbd_render.streaming_utils import (
encode_video_from_memory,
encode_depth_from_memory
)
CHUNK_SIZE = 20_000
FPS = 30
total_frames = 2_000_000
num_chunks = total_frames // CHUNK_SIZE
for chunk_idx in range(num_chunks):
start = chunk_idx * CHUNK_SIZE
end = min(start + CHUNK_SIZE, total_frames)
# Render RGB and depth for this chunk directly to memory
rgb_frames, depth_frames = viser_wrapper.render_chunk(start, end)
# Encode without intermediate PNGs
encode_video_from_memory(
rgb_frames,
f'rendered_chunks/rgb_{chunk_idx:04d}.mp4',
fps=FPS
)
encode_depth_from_memory(
depth_frames,
f'rendered_chunks/depth_{chunk_idx:04d}.mp4',
fps=FPS,
depth_range=(0.5, 10.0)
)
This zero-copy approach prevents filesystem thrashing and allows the pipeline to sustain high throughput for days-long rendering jobs.
Optimizing Frame Rate and Resolution
Very long videos often do not require full native resolution or frame rate to convey spatial relationships effectively. Reduce the total frame count and processing load by adjusting parameters in demo_render/config/default.yaml:
fps: 30 # Desired output frame rate
width: 1280 # Set to 0 to keep original width
height: 720 # Set to 0 to keep original height
downsample_factor: 2 # Render every N-th frame only
These values propagate to demo_render/rgbd_render/renderer.py during scene initialization. Setting downsample_factor: 2 halves the frame count while maintaining visual continuity, effectively doubling the chunk size capacity for the same memory footprint.
Parallelizing Across Multiple GPUs
When hardware resources permit, distribute chunk rendering across multiple GPUs or CPU cores to reduce wall-clock time. Use GNU Parallel to orchestrate workers:
parallel -j 4 ./demo_render/process_videos.sh \
--scene-file scene.yaml \
--chunk-index {} \
--chunk-size 20000 \
--output-dir ./rendered_chunks \
::: $(seq 0 99)
Each worker generates independent MP4 segments. Concatenate these into a final video without re-encoding using FFmpeg:
ffmpeg -f concat -safe 0 \
-i <(for f in rendered_chunks/*.mp4; do echo "file '$f'"; done) \
-c copy final_long_video.mp4
This method preserves quality while combining hundreds of segments into a single playable file.
Memory-Efficient Depth Visualization
Depth sequence encoding (encode_depth_video) typically colorizes each frame using the Turbo colormap before compression, which doubles memory usage for RGB buffers. For massive sequences, use the streaming encode_depth_from_memory implementation shown above, which processes and writes one frame at a time rather than holding the entire colorized sequence in RAM.
Alternatively, adjust the depth_range parameter to use narrower bounds, reducing the numerical precision required for the visualization and improving compression ratios.
Summary
- Segment long sequences using
--chunk-sizeindemo_render/process_videos.shto cap memory usage at 10k–100k frames per iteration. - Stream frames directly to the encoder using
encode_video_from_memoryto eliminate intermediate PNG storage and reduce filesystem load. - Tune rendering parameters in
demo_render/config/default.yaml—specificallydownsample_factorandfps—to reduce the total frame count for sequences exceeding 25,000 frames. - Parallelize chunk processing across multiple GPUs with GNU Parallel, then concatenate segments using FFmpeg's concat demuxer.
- Process depth maps efficiently by colorizing on-the-fly in the encoding loop rather than holding full RGB sequences in memory.
Frequently Asked Questions
What is the maximum recommended chunk size for 25,000+ frame videos?
The optimal chunk size depends on your GPU VRAM and system RAM. For systems with 24GB VRAM and 64GB system RAM, chunks of 20,000 to 50,000 frames provide a balance between efficiency and stability. For 2-million-frame sequences, 20,000-frame chunks prevent memory exhaustion while minimizing encoder invocation overhead.
Can I render videos longer than 2 million frames without intermediate files?
Yes, by implementing the encode_video_from_memory streaming pattern demonstrated in the streaming_utils.py code example. This approach feeds NumPy arrays directly from viser_wrapper.render_chunk() into OpenCV's VideoWriter, bypassing the PNG serialization in demo_render/rgbd_render/renderer.py entirely.
How do I adjust the rendering resolution without modifying source code?
Edit demo_render/config/default.yaml and set the width and height parameters to your desired output dimensions, or set them to 0 to preserve original sensor resolution. These values are read during scene initialization in demo_render/rgbd_render/scene.py and enforced throughout the rendering pipeline.
Why does the depth video encoding consume more memory than RGB encoding?
The encode_depth_video function in demo_render/rgbd_render/video.py applies a colorization colormap (typically Turbo) to single-channel depth arrays, converting them to three-channel RGB representations. This triples the memory footprint. Use the streaming encode_depth_from_memory helper to process depth frames iteratively, or reduce the depth_range interval to minimize precision requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →