# How `expandable_segments` Affects CUDA Memory Allocation in LingBot-Map

> Learn how expandable_segments improves CUDA memory allocation in LingBot-Map. Enable dynamic memory growth the PyTorch CUDA allocator to fit larger workloads on your GPU.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-25

---

**Setting `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` configures PyTorch’s CUDA allocator to grow memory segments dynamically on demand, eliminating the large gap between allocated and reserved memory and allowing larger streaming workloads to fit on limited GPU memory.**

LingBot-Map leverages PyTorch’s caching allocator for all CUDA tensor operations. By configuring the `expandable_segments` option, the repository optimizes memory usage for streaming inference scenarios where the KV-cache grows gradually over time, preventing out-of-memory errors on GPUs with constrained capacity.

## What `expandable_segments` Changes in the CUDA Allocator

PyTorch’s CUDA allocator pre-reserves memory pools to speed up allocation and deallocation. The `expandable_segments` flag fundamentally alters how these pools expand to match workload demands.

### Without Expandable Segments (Default Behavior)

When `expandable_segments` is disabled, the allocator creates **fixed-size segments** up front. This strategy maintains fast allocation speeds but creates a persistent gap between `memory_allocated` (actually used) and `memory_reserved` (pre-reserved pool). The reserved pool stays at its maximum size throughout execution, often leaving several gigabytes unused yet blocked from other applications.

### With Expandable Segments Enabled

When enabled via the environment variable, segments start small and **grow on demand** only when forward passes require additional space. This dynamic expansion ensures that `memory_reserved` tracks closely with `memory_allocated`, significantly reducing the memory footprint and allowing longer sequences or higher-resolution inputs to process without triggering out-of-memory errors.

| Behavior | Fixed Segments (Default) | Expandable Segments |
|----------|-------------------------|---------------------|
| **Segment Growth** | Fixed size allocated upfront | Grows dynamically as needed |
| **Reserved vs Allocated Gap** | Large (often several GB) | Minimal (tracks actual usage) |
| **Peak Memory Footprint** | Higher reserved memory | Lower reserved memory |
| **Compatibility** | Universal | Incompatible with `torch.compile` + CUDA graphs |

## Implementation and Early Configuration in LingBot-Map

LingBot-Map sets this configuration **early**, before any CUDA context initializes, ensuring the allocator adopts expandable mode for the entire runtime. In [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) lines 28-31, the code establishes the environment variable prior to importing `torch`:

```python
import os

# Must be set before any torch import

os.environ.setdefault(
    "PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True"
)
import torch

```

The benchmark script [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py) applies the same setting at lines 38-40 for synthetic memory profiling. This early initialization is critical because PyTorch’s allocator determines its strategy at context creation; changing the configuration after importing `torch` has no effect on the active session.

## Memory Impact on Streaming Inference

LingBot-Map’s **streaming inference** processes frames one-by-one while maintaining a KV-cache that grows as more frames stream through. Without expandable segments, the allocator would reserve a massive upfront pool to accommodate the maximum potential cache size, limiting the total number of frames processable on a given GPU.

With `expandable_segments:True`, the allocator expands only as the cache grows. Console logs demonstrate the tight correlation between allocated and reserved memory:

```text
GPU mem after load: alloc=2.31 GB, reserved=2.31 GB
...
GPU peak during inference: 4.27 GB (reserved peak 4.35 GB)

```

Without the flag, the reserved figures would be noticeably higher than the allocated figures, effectively wasting GPU memory that could otherwise accommodate additional frames or higher resolution inputs.

## The `torch.compile` Incompatibility

There is a critical caveat when using `expandable_segments` with modern PyTorch optimization features. The flag is **incompatible** with `torch.compile` when using CUDA-graph warm-up (`cudagraph_trees`). The compiled path assumes the classic fixed-segment topology, leading to a runtime error.

When users request `--compile`, LingBot-Map deliberately **skips** setting the flag to avoid this error. In [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) lines 32-38, the code implements a conditional check:

```python
import sys

if "--compile" not in sys.argv:
    os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")

```

Attempting to use both features simultaneously produces the error:

```

RuntimeError: Expected curr_block->next == nullptr

```

This occurs because CUDA-graph replay expects the fixed-segment memory layout that `expandable_segments` modifies.

## Practical Code Examples

### Enable Expandable Segments (Standard Configuration)

Execute this before any PyTorch import to activate dynamic memory growth:

```python
import os

os.environ.setdefault(
    "PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True"
)
import torch

```

### Inspect Allocated vs Reserved Memory

Monitor the effectiveness of the setting by comparing these statistics:

```python
device = torch.device("cuda")
torch.cuda.empty_cache()

# After model load

print(f"Allocated: {torch.cuda.memory_allocated(device)/1e9:.2f} GB")
print(f"Reserved : {torch.cuda.memory_reserved(device)/1e9:.2f} GB")

```

When `expandable_segments` is active, the output shows a tight gap between allocated and reserved figures.

### Disable for Compiled Runs

Mirror LingBot-Map’s compatibility logic in your own scripts:

```python
import os
import sys

if "--compile" not in sys.argv:
    os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")

```

This conditional prevents the configuration when `torch.compile` is requested.

## Summary

- **Dynamic Growth**: `expandable_segments:True` allows PyTorch’s CUDA allocator to grow memory pools on demand rather than pre-reserving large fixed blocks.
- **Memory Efficiency**: The setting eliminates the gap between allocated and reserved memory, crucial for LingBot-Map’s streaming inference where KV-caches expand gradually.
- **Early Configuration**: Set `PYTORCH_CUDA_ALLOC_CONF` before importing `torch` in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (lines 28-31) or [`benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark_gct_memory.py) (lines 38-40).
- **Compatibility Trade-off**: Disable the flag when using `torch.compile` with CUDA graphs to avoid `RuntimeError: Expected curr_block->next == nullptr`.

## Frequently Asked Questions

### What is the exact environment variable format for enabling expandable segments?

Set `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` in your environment before importing PyTorch. According to the LingBot-Map source code in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) lines 28-31, you must use `os.environ.setdefault()` prior to `import torch` to ensure the allocator initializes with this configuration.

### Why does LingBot-Map need expandable segments for streaming inference?

Streaming inference processes video frames sequentially while maintaining a KV-cache that grows with each new frame. Without expandable segments, PyTorch reserves the maximum potential memory upfront, wasting GPU capacity. The dynamic growth model allows the allocator to match the actual cache size, fitting longer sequences on limited hardware.

### When should I disable expandable_segments in LingBot-Map?

Disable the flag when using the `--compile` command-line argument. As implemented in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) lines 32-38, LingBot-Map skips setting the environment variable during compiled runs because `torch.compile` with CUDA-graph warm-up assumes a fixed-segment memory topology. Using both together triggers a `RuntimeError` regarding block linkage expectations.

### How can I verify that expandable segments is working?

Check the relationship between allocated and reserved memory using `torch.cuda.memory_allocated()` and `torch.cuda.memory_reserved()`. When the feature is active, these values remain close together (e.g., 4.27 GB allocated vs 4.35 GB reserved). Without the flag, reserved memory typically exceeds allocated memory by several gigabytes.