How to Use torch.compile for Performance Tuning in LingBot-Map: Reducing Inference Overhead

Use the --compile flag in demo.py or benchmark scripts to enable torch.compile with mode="reduce-overhead" on hot GCT sub-modules, followed by CUDA graph warm-up to eliminate first-frame overhead in streaming inference.

The LingBot-Map repository (Robbyant/lingbot-map) implements a high-performance streaming inference pipeline for the GCT model. By leveraging torch.compile performance tuning techniques, you can significantly reduce JIT compilation overhead and fuse heavy linear operations while maintaining the dynamic KV-cache management required for streaming workloads.

Why torch.compile Uses reduce-overhead Mode

torch.compile supports several compilation modes, but LingBot-Map specifically fixes mode="reduce-overhead" in its optimization workflow. This mode is optimal because the GCT model contains many small, dynamic control-flow steps for KV-cache management and per-frame skips. The reduce-overhead setting minimizes JIT compilation costs while still fusing heavy linear and attention kernels, making it ideal for low-latency streaming inference.

Step-by-Step: Enabling torch.compile in LingBot-Map

Step 1: Enable the Compile Flag

All entry-point scripts in the repository accept a --compile argument. When present, the model is passed through the compile_model helper before streaming starts:

python demo.py \
    --image_folder path/to/images \
    --backend flashinfer \
    --dtype auto \
    --compile

This flag is also available in scripts/benchmark_gct_memory.py and gct_profile.py for benchmarking and profiling workflows.

Step 2: Compile Hot Sub-Modules

The compile_model function (located in demo.py and scripts/benchmark_gct_memory.py) iterates over specific high-traffic modules within the GCT aggregator. It compiles frame blocks, patch-embedding blocks, and attention components:


# Compile frame blocks

agg.frame_blocks[i] = torch.compile(block, mode="reduce-overhead")

# Compile patch embedding blocks  

agg.patch_embed.blocks[i] = torch.compile(block, mode="reduce-overhead")

# Compile attention pre-processing and FFN residuals if present

if hasattr(block, "attn_pre"):
    block.attn_pre = torch.compile(block.attn_pre, mode="reduce-overhead")
if hasattr(block, "ffn_residual"):
    block.ffn_residual = torch.compile(block.ffn_residual, mode="reduce-overhead")

# Compile attention projections

block.attn.proj = torch.compile(block.attn.proj, mode="reduce-overhead")

This selective compilation targets the heavy computational paths in lingbot_map/layers/block.py and lingbot_map/layers/attention.py without modifying the underlying source code.

Step 3: Warm-Up with CUDA Graph Capture

After compilation, the demo executes _warm_streaming to capture a CUDA graph that matches the exact tensor shapes used in production. This eliminates the first-frame overhead of graph capture:

def warmup(model, images, dtype):
    torch.compiler.cudagraph_mark_step_begin()
    with torch.no_grad(), torch.amp.autocast("cuda", dtype=dtype):
        model.forward(images[:1])  # Dummy forward to capture graph

warmup(model, synthetic_images, torch.float16)

The torch.compiler.cudagraph_mark_step_begin() call ensures that subsequent forward passes in the streaming loop execute entirely within the compiled CUDA graph.

Programmatic Usage Without CLI

For custom integrations, compile the model manually after loading:

import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize model

model = GCTStream(...)
model.eval().cuda()

# Compile hot paths

agg = model.aggregator
for i, blk in enumerate(agg.frame_blocks):
    agg.frame_blocks[i] = torch.compile(blk, mode="reduce-overhead")
for i, blk in enumerate(agg.patch_embed.blocks):
    agg.patch_embed.blocks[i] = torch.compile(blk, mode="reduce-overhead")
for blk in agg.global_blocks:
    if hasattr(blk, "attn_pre"):
        blk.attn_pre = torch.compile(blk.attn_pre, mode="reduce-overhead")
    if hasattr(blk, "ffn_residual"):
        blk.ffn_residual = torch.compile(blk.ffn_residual, mode="reduce-overhead")
    blk.attn.proj = torch.compile(blk.attn.proj, mode="reduce-overhead")

# Warm-up and run

torch.compiler.cudagraph_mark_step_begin()
with torch.no_grad():
    for frame in stream_of_frames:
        out = model.forward(frame)

Benchmarking Compilation Benefits

Measure the performance gains using the built-in benchmark script:

python scripts/benchmark_gct_memory.py \
    --height 384 --width 518 \
    --frame-counts 64 128 256 \
    --compile

This outputs both peak memory usage and runtime for each frame count, allowing direct comparison between eager mode and compiled execution.

Key Configuration Details

  • Device Requirement: torch.compile only operates on CUDA devices. The scripts automatically configure PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to prevent out-of-memory errors during compilation.
  • Compatibility: The compilation path is safe to use with other flags including --backend flashinfer and --dtype auto.
  • Source Files: The implementation resides in demo.py (entry point), scripts/benchmark_gct_memory.py (benchmarking), and lingbot_map/models/gct_stream.py (core streaming model).

Summary

  • Add --compile to any CLI command to activate the torch.compile workflow in LingBot-Map.
  • The system targets frame blocks, patch embeddings, and attention projections using mode="reduce-overhead".
  • CUDA graph warm-up is required to capture tensor shapes and eliminate first-frame latency penalties.
  • The optimization is fully compatible with FlashInfer backends and automatic mixed precision.
  • Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid OOM issues during the compilation phase.

Frequently Asked Questions

What compilation mode does LingBot-Map use and why?

LingBot-Map uses mode="reduce-overhead" because the GCT model's streaming inference involves frequent dynamic control-flow decisions for KV-cache management. This mode minimizes JIT compilation overhead while still fusing the heavy linear and attention kernels that dominate execution time.

Is torch.compile compatible with FlashInfer backends?

Yes, the compilation workflow is fully compatible with the --backend flashinfer flag. You can enable both optimizations simultaneously without source code modifications, allowing the compiled kernels to work alongside FlashInfer's optimized attention implementations.

Do I need to modify source code to enable compilation?

No, the repository provides an opt-in workflow. Simply add the --compile argument when running demo.py, scripts/benchmark_gct_memory.py, or gct_profile.py. The compile_model helper function handles all module wrapping and configuration automatically.

Why is CUDA graph warm-up necessary for streaming inference?

The warm-up routine executes torch.compiler.cudagraph_mark_step_begin() before dummy forward passes to record a CUDA graph with the exact tensor shapes used in production. This eliminates the runtime overhead of graph capture on the first real inference frame, ensuring consistent low latency throughout the streaming session.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →