How to Use torch.compile for Accelerated Inference in Ling Bot-Map
Enabling torch.compile in Ling Bot-Map involves passing the --compile flag to entry-point scripts or manually wrapping the model's hot sub-modules with torch.compile(..., mode="reduce-overhead"), followed by a CUDA graph warm-up to eliminate first-frame overhead.
Ling Bot-Map leverages torch.compile to accelerate its streaming inference pipeline for the GCT model. This optimization compiles fixed-shape transformer blocks while preserving dynamic control flow for KV-cache management, delivering significant latency reduction without increasing memory consumption. The implementation follows an opt-in pattern found in demo.py and supporting utilities that targets specific sub-modules and warms them with CUDA graphs.
Why the Reduce-Overhead Mode?
torch.compile supports several compilation modes, but Ling Bot-Map specifically uses mode="reduce-overhead" according to the source code in demo.py and scripts/benchmark_gct_memory.py. This mode is ideal for streaming workloads because the model contains many small, dynamic control-flow steps such as KV-cache management and per-frame skips. The reduce-overhead mode minimizes JIT compilation costs while still fusing heavy linear and attention kernels, making it optimal for latency-sensitive inference.
Prerequisites and Device Configuration
torch.compile only functions on CUDA devices. The repository automatically configures PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True in entry-point scripts to prevent out-of-memory errors during the compilation phase. This setting allows PyTorch to expand CUDA memory segments dynamically rather than pre-allocating large contiguous blocks.
Step-by-Step Implementation
1. Enable the Compile Flag in Entry Points
All primary entry-point scripts—including demo.py, scripts/benchmark_gct_memory.py, and gct_profile.py—accept a --compile argument. When present, the script passes the model through the compile_model helper before streaming begins:
python demo.py \
--image_folder path/to/images \
--backend flashinfer \
--dtype auto \
--compile
2. Compile Hot Sub-Modules Programmatically
The compile_model helper (implemented in demo.py and scripts/benchmark_gct_memory.py) iterates over fixed-shape sub-modules of the GCT aggregator. It replaces critical blocks with their compiled versions:
# Compile frame blocks
agg.frame_blocks[i] = torch.compile(block, mode="reduce-overhead")
# Compile patch embedding blocks
agg.patch_embed.blocks[i] = torch.compile(block, mode="reduce-overhead")
# Compile global block components
if hasattr(block, "attn_pre"):
block.attn_pre = torch.compile(block.attn_pre, mode="reduce-overhead")
if hasattr(block, "ffn_residual"):
block.ffn_residual = torch.compile(block.ffn_residual, mode="reduce-overhead")
block.attn.proj = torch.compile(block.attn.proj, mode="reduce-overhead")
These targets reside in lingbot_map/models/gct_stream.py and lingbot_map/layers/, specifically within the transformer block definitions (block.py, attention.py).
3. Warm-Up with CUDA Graph Capture
After compilation, the demo executes _warm_streaming to prepare CUDA graphs. This routine calls torch.compiler.cudagraph_mark_step_begin() before each forward pass during a short synthetic run, capturing the exact tensor shapes the real inference will use:
def warmup(model, images, dtype):
torch.compiler.cudagraph_mark_step_begin()
with torch.no_grad(), torch.amp.autocast("cuda", dtype=dtype):
model.forward(images[:1]) # Capture graph with dummy input
warmup(model, synthetic_images, torch.float16)
This step eliminates first-frame overhead and ensures compiled kernels remain "hot" in memory. The warm-up is called automatically when using --compile with demo.py.
Complete Programmatic Example
For custom integrations outside the CLI tools, implement the compilation workflow manually:
import torch
from lingbot_map.models.gct_stream import GCTStream
# Initialize model
model = GCTStream(...)
model.eval().cuda()
agg = model.aggregator
# Compile fixed-shape blocks
for i, blk in enumerate(agg.frame_blocks):
agg.frame_blocks[i] = torch.compile(blk, mode="reduce-overhead")
for i, blk in enumerate(agg.patch_embed.blocks):
agg.patch_embed.blocks[i] = torch.compile(blk, mode="reduce-overhead")
for blk in agg.global_blocks:
if hasattr(blk, "attn_pre"):
blk.attn_pre = torch.compile(blk.attn_pre, mode="reduce-overhead")
if hasattr(blk, "ffn_residual"):
blk.ffn_residual = torch.compile(blk.ffn_residual, mode="reduce-overhead")
blk.attn.proj = torch.compile(blk.attn.proj, mode="reduce-overhead")
# Warm-up CUDA graphs
torch.compiler.cudagraph_mark_step_begin()
with torch.no_grad(), torch.amp.autocast("cuda", dtype=torch.float16):
model.forward(dummy_input)
# Run streaming inference
with torch.no_grad():
for frame in stream_of_frames:
output = model.forward(frame)
Benchmarking the Acceleration
Measure your speed-up using the provided benchmark script:
python scripts/benchmark_gct_memory.py \
--height 384 --width 518 \
--frame-counts 64 128 256 \
--compile
This script prints both peak memory usage and runtime for each frame count, enabling direct comparison between eager and compiled execution modes. For detailed profiling statistics, use gct_profile.py with the same --compile flag to see per-layer compilation overhead and kernel fusion results.
Compatibility with Other Features
The compilation path is safe to use alongside other optimization flags. You can combine --compile with:
--backend flashinfer(alternative attention implementation)--dtype auto(automatic mixed precision)- Various frame resolutions and batch configurations
No source-code modifications are required when toggling compilation on or off.
Summary
- Activation: Add
--compiletodemo.py,benchmark_gct_memory.py, orgct_profile.pyto enable the workflow. - Targets: The compiler focuses on
frame_blocks,patch_embed.blocks, and global block components (attn_pre,ffn_residual,attn.proj). - Mode: Always uses
mode="reduce-overhead"to balance dynamic control flow with kernel fusion. - Warm-up: CUDA graph capture via
_warm_streamingis required to eliminate first-frame overhead in streaming mode. - Hardware: CUDA-only, with automatic memory configuration to prevent OOM during compilation.
Frequently Asked Questions
Does torch.compile increase memory usage in Ling Bot-Map?
No, memory usage remains unchanged according to the repository implementation. The reduce-overhead mode optimizes kernel execution without altering tensor allocations. The scripts set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True specifically to manage compilation memory efficiently without increasing the overall footprint.
Why is the warm-up step necessary when using torch.compile?
The warm-up step captures a CUDA graph that matches the exact input shapes and data types of real inference. This eliminates the latency penalty of graph capture on the first real frame. Without warm-up, the first few inference calls would execute slower as PyTorch dynamically compiles and optimizes kernels on-demand.
Can I use torch.compile with CPU-only inference?
No, the Ling Bot-Map implementation requires CUDA. The torch.compile workflow in this repository specifically targets CUDA graph capture and GPU kernel optimization. CPU inference falls back to standard eager mode execution regardless of the --compile flag setting.
What happens if I omit the --compile flag?
The model runs in standard eager mode. All sub-modules execute their original Python implementations without JIT compilation. While functionally identical, eager mode exhibits higher latency per frame due to Python overhead and lacks the kernel fusion benefits provided by the compiled graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →