# How to Use torch.compile for Performance Tuning in LingBot-Map: Reducing Inference Overhead

> Boost LingBot-Map performance by enabling torch.compile with reduce-overhead mode. Learn how to slash inference overhead and eliminate first-frame latency for smoother streaming.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-30

---

**Use the `--compile` flag in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) or benchmark scripts to enable `torch.compile` with `mode="reduce-overhead"` on hot GCT sub-modules, followed by CUDA graph warm-up to eliminate first-frame overhead in streaming inference.**

The LingBot-Map repository (Robbyant/lingbot-map) implements a high-performance streaming inference pipeline for the GCT model. By leveraging `torch.compile` performance tuning techniques, you can significantly reduce JIT compilation overhead and fuse heavy linear operations while maintaining the dynamic KV-cache management required for streaming workloads.

## Why torch.compile Uses reduce-overhead Mode

`torch.compile` supports several compilation modes, but LingBot-Map specifically fixes `mode="reduce-overhead"` in its optimization workflow. This mode is optimal because the GCT model contains many small, dynamic control-flow steps for KV-cache management and per-frame skips. The **reduce-overhead** setting minimizes JIT compilation costs while still fusing heavy linear and attention kernels, making it ideal for low-latency streaming inference.

## Step-by-Step: Enabling torch.compile in LingBot-Map

### Step 1: Enable the Compile Flag

All entry-point scripts in the repository accept a `--compile` argument. When present, the model is passed through the `compile_model` helper before streaming starts:

```bash
python demo.py \
    --image_folder path/to/images \
    --backend flashinfer \
    --dtype auto \
    --compile

```

This flag is also available in [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) and [`gct_profile.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_profile.py) for benchmarking and profiling workflows.

### Step 2: Compile Hot Sub-Modules

The `compile_model` function (located in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py)) iterates over specific high-traffic modules within the GCT aggregator. It compiles frame blocks, patch-embedding blocks, and attention components:

```python

# Compile frame blocks

agg.frame_blocks[i] = torch.compile(block, mode="reduce-overhead")

# Compile patch embedding blocks  

agg.patch_embed.blocks[i] = torch.compile(block, mode="reduce-overhead")

# Compile attention pre-processing and FFN residuals if present

if hasattr(block, "attn_pre"):
    block.attn_pre = torch.compile(block.attn_pre, mode="reduce-overhead")
if hasattr(block, "ffn_residual"):
    block.ffn_residual = torch.compile(block.ffn_residual, mode="reduce-overhead")

# Compile attention projections

block.attn.proj = torch.compile(block.attn.proj, mode="reduce-overhead")

```

This selective compilation targets the heavy computational paths in [`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py) and [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) without modifying the underlying source code.

### Step 3: Warm-Up with CUDA Graph Capture

After compilation, the demo executes `_warm_streaming` to capture a CUDA graph that matches the exact tensor shapes used in production. This eliminates the first-frame overhead of graph capture:

```python
def warmup(model, images, dtype):
    torch.compiler.cudagraph_mark_step_begin()
    with torch.no_grad(), torch.amp.autocast("cuda", dtype=dtype):
        model.forward(images[:1])  # Dummy forward to capture graph

warmup(model, synthetic_images, torch.float16)

```

The `torch.compiler.cudagraph_mark_step_begin()` call ensures that subsequent forward passes in the streaming loop execute entirely within the compiled CUDA graph.

## Programmatic Usage Without CLI

For custom integrations, compile the model manually after loading:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize model

model = GCTStream(...)
model.eval().cuda()

# Compile hot paths

agg = model.aggregator
for i, blk in enumerate(agg.frame_blocks):
    agg.frame_blocks[i] = torch.compile(blk, mode="reduce-overhead")
for i, blk in enumerate(agg.patch_embed.blocks):
    agg.patch_embed.blocks[i] = torch.compile(blk, mode="reduce-overhead")
for blk in agg.global_blocks:
    if hasattr(blk, "attn_pre"):
        blk.attn_pre = torch.compile(blk.attn_pre, mode="reduce-overhead")
    if hasattr(blk, "ffn_residual"):
        blk.ffn_residual = torch.compile(blk.ffn_residual, mode="reduce-overhead")
    blk.attn.proj = torch.compile(blk.attn.proj, mode="reduce-overhead")

# Warm-up and run

torch.compiler.cudagraph_mark_step_begin()
with torch.no_grad():
    for frame in stream_of_frames:
        out = model.forward(frame)

```

## Benchmarking Compilation Benefits

Measure the performance gains using the built-in benchmark script:

```bash
python scripts/benchmark_gct_memory.py \
    --height 384 --width 518 \
    --frame-counts 64 128 256 \
    --compile

```

This outputs both peak memory usage and runtime for each frame count, allowing direct comparison between eager mode and compiled execution.

## Key Configuration Details

- **Device Requirement**: `torch.compile` only operates on CUDA devices. The scripts automatically configure `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` to prevent out-of-memory errors during compilation.
- **Compatibility**: The compilation path is safe to use with other flags including `--backend flashinfer` and `--dtype auto`.
- **Source Files**: The implementation resides in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (entry point), [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) (benchmarking), and [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) (core streaming model).

## Summary

- Add `--compile` to any CLI command to activate the `torch.compile` workflow in LingBot-Map.
- The system targets frame blocks, patch embeddings, and attention projections using `mode="reduce-overhead"`.
- CUDA graph warm-up is required to capture tensor shapes and eliminate first-frame latency penalties.
- The optimization is fully compatible with FlashInfer backends and automatic mixed precision.
- Set `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` to avoid OOM issues during the compilation phase.

## Frequently Asked Questions

### What compilation mode does LingBot-Map use and why?

LingBot-Map uses `mode="reduce-overhead"` because the GCT model's streaming inference involves frequent dynamic control-flow decisions for KV-cache management. This mode minimizes JIT compilation overhead while still fusing the heavy linear and attention kernels that dominate execution time.

### Is torch.compile compatible with FlashInfer backends?

Yes, the compilation workflow is fully compatible with the `--backend flashinfer` flag. You can enable both optimizations simultaneously without source code modifications, allowing the compiled kernels to work alongside FlashInfer's optimized attention implementations.

### Do I need to modify source code to enable compilation?

No, the repository provides an opt-in workflow. Simply add the `--compile` argument when running [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py), or [`gct_profile.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_profile.py). The `compile_model` helper function handles all module wrapping and configuration automatically.

### Why is CUDA graph warm-up necessary for streaming inference?

The warm-up routine executes `torch.compiler.cudagraph_mark_step_begin()` before dummy forward passes to record a CUDA graph with the exact tensor shapes used in production. This eliminates the runtime overhead of graph capture on the first real inference frame, ensuring consistent low latency throughout the streaming session.