# How torch.compile Improves LingBot-Map Inference Speed: A Complete Guide

> Discover how torch.compile accelerates LingBot-Map inference by optimizing transformer sub-modules via CUDA graph capture. Learn to boost your model's speed.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-23

---

**torch.compile reduces LingBot-Map streaming inference latency by compiling hot transformer sub-modules with "reduce-overhead" mode and warming them up with CUDA graph capture, triggered via the `--compile` flag in demo.py.**

LingBot-Map accelerates its streaming GCT model inference through an opt-in optimization workflow using PyTorch 2.0's `torch.compile`. According to the Robbyant/lingbot-map source code, this implementation targets fixed-shape sub-modules within the aggregator and applies CUDA graph warm-up to eliminate first-frame overhead, delivering significant latency reduction without increasing memory consumption.

## Understanding the Compilation Strategy

### Why "reduce-overhead" Mode

`torch.compile` supports multiple compilation modes, but LingBot-Map explicitly uses `mode="reduce-overhead"` to handle the GCT model's architectural constraints. The streaming pipeline contains many small, dynamic control-flow steps including KV-cache management and per-frame skip connections. The "reduce-overhead" mode minimizes JIT compilation costs while still fusing heavy linear and attention kernels, making it ideal for streaming workloads where consistent frame rates matter more than single-batch throughput.

### Targeting Hot Sub-Modules

Rather than compiling the entire model, LingBot-Map selectively compiles only the computationally intensive blocks. In [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), the `compile_model` helper iterates over specific module lists to minimize compilation time and memory footprint:

- **Frame blocks**: `agg.frame_blocks[i]`
- **Patch embedding blocks**: `agg.patch_embed.blocks[i]`
- **Attention pre-processing**: `block.attn_pre` (if present)
- **FFN residual connections**: `block.ffn_residual` (if present)
- **Attention projections**: `block.attn.proj`

This granular approach ensures only the fixed-shape, high-execution-frequency paths receive optimization.

## Step-by-Step Implementation

### 1. Enable Compilation via CLI

All entry-point scripts accept a `--compile` argument that activates the optimization pipeline. When present, the model passes through the `compile_model` helper before streaming begins.

```bash
python demo.py \
    --image_folder path/to/images \
    --backend flashinfer \
    --dtype auto \
    --compile

```

The scripts automatically configure `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` to prevent out-of-memory errors during the compilation phase.

### 2. Compile Core Transformer Blocks

The `compile_model` function in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) applies `torch.compile` to each target module:

```python

# Compile frame blocks

agg.frame_blocks[i] = torch.compile(block, mode="reduce-overhead")

# Compile patch embedding blocks  

agg.patch_embed.blocks[i] = torch.compile(block, mode="reduce-overhead")

# Compile attention and FFN sub-components

if hasattr(block, "attn_pre"):
    block.attn_pre = torch.compile(block.attn_pre, mode="reduce-overhead")
if hasattr(block, "ffn_residual"):
    block.fnn_residual = torch.compile(block.ffn_residual, mode="reduce-overhead")
block.attn.proj = torch.compile(block.attn.proj, mode="reduce-overhead")

```

This code replaces the original modules with compiled versions that execute optimized CUDA kernels.

### 3. CUDA Graph Warm-Up

After compilation, the demo executes a warm-up routine (`_warm_streaming`) that captures a CUDA graph matching the exact inference shapes. The implementation calls `torch.compiler.cudagraph_mark_step_begin()` before each forward pass during warm-up:

```python
def warmup(model, images, dtype):
    torch.compiler.cudagraph_mark_step_begin()
    with torch.no_grad(), torch.amp.autocast("cuda", dtype=dtype):
        model.forward(images[:1])  # Dummy forward for graph capture

warmup(model, synthetic_images, torch.float16)

```

This step eliminates the first-frame overhead of graph capture and ensures compiled kernels remain hot in GPU memory.

### 4. Execute Streaming Inference

Once warmed up, the standard `inference_streaming` method in the GCT model runs entirely under the compiled graph. The normal streaming loop executes without source code modifications:

```python
with torch.no_grad():
    for frame in stream_of_frames:
        out = model.forward(frame)  # Executes compiled kernels

```

## Code Implementation Examples

### Minimal Programmatic Usage

For custom integrations, manually invoke the compilation workflow after model initialization:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize model

model = GCTStream(...)
model.eval().cuda()

# Access aggregator

agg = model.aggregator

# Compile frame blocks

for i, blk in enumerate(agg.frame_blocks):
    agg.frame_blocks[i] = torch.compile(blk, mode="reduce-overhead")

# Compile patch embeddings

for i, blk in enumerate(agg.patch_embed.blocks):
    agg.patch_embed.blocks[i] = torch.compile(blk, mode="reduce-overhead")

# Compile global block components

for blk in agg.global_blocks:
    if hasattr(blk, "attn_pre"):
        blk.attn_pre = torch.compile(blk.attn_pre, mode="reduce-overhead")
    if hasattr(blk, "ffn_residual"):
        blk.ffn_residual = torch.compile(blk.ffn_residual, mode="reduce-overhead")
    blk.attn.proj = torch.compile(blk.attn.proj, mode="reduce-overhead")

# Warm-up with CUDA graph

torch.compiler.cudagraph_mark_step_begin()
with torch.no_grad(), torch.amp.autocast("cuda", dtype=torch.float16):
    model.forward(dummy_input)

```

### Benchmarking the Speed-Up

Use the built-in benchmarking script to measure improvements:

```bash
python scripts/benchmark_gct_memory.py \
    --height 384 --width 518 \
    --frame-counts 64 128 256 \
    --compile

```

The script reports both peak memory usage and runtime for each frame count, allowing direct comparison between compiled and eager execution modes.

## Performance and Compatibility

**Device Requirements**: `torch.compile` requires CUDA devices. CPU inference falls back to standard eager mode execution.

**Memory Characteristics**: The compilation maintains the same peak memory usage as eager mode while reducing latency, as the optimization focuses on kernel fusion rather than activation checkpointing.

**Flag Compatibility**: The `--compile` flag works safely with other performance flags including `--backend flashinfer` and `--dtype auto`. No source code modifications are required to enable these combinations.

## Summary

- **torch.compile** accelerates LingBot-Map by fusing attention and linear kernels in the GCT model's hot paths.
- Use the `--compile` CLI flag in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) or [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) to activate the workflow.
- The implementation targets specific sub-modules (frame blocks, patch embeddings, attention projections) with `mode="reduce-overhead"`.
- CUDA graph warm-up via `torch.compiler.cudagraph_mark_step_begin()` is required to eliminate first-frame overhead.
- Compatible with FlashInfer backends and automatic mixed precision (`dtype auto`).

## Frequently Asked Questions

### What hardware is required to use torch.compile in LingBot-Map?

`torch.compile` requires a CUDA-capable GPU. The LingBot-Map scripts automatically configure CUDA memory allocation settings (`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`) to prevent OOM errors during compilation, but CPU-only environments will execute the standard eager mode without compilation benefits.

### Which specific modules get compiled when I use the --compile flag?

The `compile_model` function compiles five specific module types: frame blocks (`agg.frame_blocks`), patch embedding blocks (`agg.patch_embed.blocks`), attention pre-processing layers (`attn_pre`), FFN residual connections (`ffn_residual`), and attention projection layers (`attn.proj`). This selective approach optimizes the fixed-shape computational bottlenecks while avoiding dynamic control-flow sections.

### Is the warm-up step necessary for every inference session?

Yes, the warm-up step is required for streaming mode to achieve optimal performance. The `_warm_streaming` routine in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) executes a few dummy frames while calling `torch.compiler.cudagraph_mark_step_begin()` to capture the CUDA graph. This eliminates compilation overhead from the actual inference loop and ensures the compiled kernels remain resident in GPU memory.

### Can I use torch.compile with FlashInfer or other backend optimizations?

Yes, the compilation path is fully compatible with the `--backend flashinfer` flag and other optimizations like `--dtype auto`. You can combine these flags safely: `python demo.py --backend flashinfer --dtype auto --compile`. The `torch.compile` transformation occurs after model initialization but before the backend-specific optimizations take effect.