RuntimeError with torch.compile and expandable_segments CUDA allocator in LingBot-Map

LingBot-Map throws a RuntimeError when both expandable_segments and torch.compile are enabled simultaneously because the dynamic KV-cache allocator conflicts with static CUDA graph capture requirements.

LingBot-Map implements a streaming inference architecture that supports memory-efficient KV-cache growth through the expandable_segments feature. However, this dynamic CUDA allocator is fundamentally incompatible with torch.compile's graph optimization, which expects static memory shapes during capture. When both features are active, the system fails with a CUDA allocation error.

Root Cause: Dynamic vs. Static Memory

The conflict stems from opposing architectural requirements. The expandable_segments allocator reserves a single large buffer designed to dynamically resize as the KV-cache grows during streaming inference. Conversely, torch.compile (particularly with mode="reduce-overhead" and CUDA graph warm-up) attempts to capture all memory allocations in a static CUDA graph at compilation time.

When the JIT compiler encounters the allocator's growth operations, the capture fails because the buffer size changes between graph recording and execution. This triggers the error:


RuntimeError: CUDA error: invalid allocation size …

Source Code Locations

The incompatibility is explicitly documented throughout the LingBot-Map codebase. In demo.py lines 32-33, the developers note: "Caveat: expandable_segments:True is incompatible with torch.compile's …". This warning is repeated in the argument parser at line 385 for the --compile flag and in gct_profile.py at line 240 for profiling runs.

The actual compilation occurs in lingbot_map/models/gct_stream.py at line 420, where hot modules are wrapped with torch.compile. The benchmark script scripts/benchmark_gct_memory.py at line 165 also applies compilation, making these locations the trigger points for the error when expandable_segments is active.

Resolution Strategies

Choose one of the following configurations based on your performance requirements:

Option 1: Streaming with Expandable Cache (No Compilation)

Use this configuration when processing large sequences that require dynamic memory growth. Disable torch.compile to allow the allocator to resize freely.

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --image_folder example/courthouse \
    --mask_sky \
    --expandable_segments True

Option 2: Compiled Inference with Static Cache

Use this configuration for maximum throughput with fixed-size sequences. Disable expandable_segments to use a fixed-size KV-cache buffer compatible with CUDA graph capture.

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --image_folder example/courthouse \
    --mask_sky \
    --compile \
    --expandable_segments False

Option 3: Attempting Both (Unsupported)

Currently, LingBot-Map does not support simultaneous use of both features. The repository's main branch lacks a custom allocator that can satisfy torch.compile's static requirements while maintaining dynamic growth capabilities. Attempting to enable both will result in an immediate RuntimeError with a clear incompatibility message.

Key Files Involved

Understanding the repository structure helps diagnose related issues:

Summary

  • The RuntimeError occurs because expandable_segments dynamically resizes CUDA memory while torch.compile requires static allocations for graph capture.
  • File references: The conflict is documented in demo.py (lines 32-33, 385) and gct_profile.py (line 240), with compilation triggers in gct_stream.py (line 420).
  • Fix: Choose either dynamic caching without compilation (--expandable_segments True without --compile) or static caching with compilation (--compile with --expandable_segments False).
  • Current limitation: Both features cannot be enabled simultaneously in the current release.

Frequently Asked Questions

Can I use torch.compile with expandable_segments if I disable CUDA graphs?

No. Even without explicit CUDA graph warm-up, torch.compile with mode="reduce-overhead" in LingBot-Map relies on static memory assumptions that conflict with the dynamic allocator. The source code explicitly marks these features as incompatible regardless of graph configuration.

Where exactly does the error originate in the codebase?

The error surfaces in lingbot_map/models/gct_stream.py at line 420 where modules are compiled, and in scripts/benchmark_gct_memory.py at line 165. However, the root cause is the allocator initialization that occurs earlier when expandable_segments=True is parsed in demo.py or gct_profile.py.

Will future versions of LingBot-Map support both features together?

According to the current main branch analysis, a custom graph-friendly allocator is not yet implemented. Future releases may introduce this capability, but as of the latest commit, you must choose between dynamic KV-cache expansion and torch compilation.

How do I verify which mode is currently active in my runtime?

Check your command-line arguments for --expandable_segments and --compile flags. If both are set to True, the argument parser in demo.py (line 385) will print a warning or raise an error before model execution begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →