# How to Build Custom CUDA Extensions for Voxelization and Frustum Culling in LingBot-Map

> Learn to build custom CUDA extensions for voxelization and frustum culling in LingBot-Map. Accelerate 3D reconstruction by compiling kernels into Python extensions with PyTorch.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-30

---

**LingBot-Map accelerates 3-D reconstruction by compiling CUDA kernels for voxelization and frustum culling into Python-accessible extensions using PyTorch's `cpp_extension` utilities.**

LingBot-Map is an open-source spatial mapping framework that leverages GPU acceleration for real-time dense reconstruction. To build custom CUDA extensions for voxelization and frustum culling in LingBot-Map, you will compile kernels from `preprocess/points_visibility/` into loadable modules that enable zero-copy processing of point cloud data.

## Prerequisites for Building LingBot-Map CUDA Extensions

Before compiling the extensions, ensure your environment meets the hardware and software requirements for CUDA development.

### System Requirements

You need an NVIDIA GPU with compute capability compatible with your CUDA toolkit version. The build process requires the CUDA toolkit to be installed system-wide, along with a C++ compiler (GCC on Linux or MSVC on Windows) that matches your PyTorch CUDA version.

### Python Dependencies

Install PyTorch with CUDA support and NVIDIA Kaolin, which provides helper geometry utilities used by the kernels.

```bash
pip install torch torchvision
pip install --index-url https://pypi.org/simple \
    kaolin -f https://nvidia-kaolin.s3.us-east-2.amazonaws.com/torch-2.8.0_cu128.html

```

## Compiling the Voxelization and Frustum Culling Kernels

The build process converts raw CUDA source files into loadable Python modules through PyTorch's JIT compilation infrastructure.

### Understanding the Source Files

Two CUDA kernels provide the core functionality according to the LingBot-Map source code:

- **`preprocess/points_visibility/visibility_kernel.cu`** – Implements sparse voxel grid construction and **Morton code** generation for spatial sorting.
- **`preprocess/points_visibility/frustum_cull.cu`** – Implements view-frustum intersection tests to filter points outside the camera view.

These files contain kernel implementations that use shared memory tiling and warp-level reductions for optimal performance.

### Running the Build Script

Navigate to the extension directory and execute the setup script. In [`demo_render/render_cuda_ext/setup.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/render_cuda_ext/setup.py), PyTorch's `cpp_extension` utilities define two modules: `voxel_morton_ext` and `frustum_cull_ext`.

```bash
cd demo_render/render_cuda_ext
python setup.py build_ext --inplace
cd ../..

```

This command compiles the kernels into shared objects (`voxel_morton_ext.*.so` and `frustum_cull_ext.*.so`) located next to the setup script. The `--inplace` flag ensures the extensions are importable from the demo_render directory structure.

## Integrating the Extensions into the Rendering Pipeline

Once built, the extensions are imported by the rendering system to perform GPU-accelerated preprocessing. The main consumer is **[`lingbot_map/vis/rgbd_render.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/rgbd_render.py)**, which imports the modules as:

```python
import voxel_morton_ext as vm
import frustum_cull_ext as fc

```

### Voxelization with voxel_morton_ext

The voxelization kernel inserts points into a sparse grid and returns sorted Morton codes. This enables deterministic spatial ordering required by the transformer-based map encoder.

```python
import torch
import voxel_morton_ext as vm

points = torch.randn(1_000_000, 3, device='cuda')
voxel_size = 0.05

morton_codes, voxel_points = vm.voxelize(points, voxel_size)
print(f"Reduced {points.shape[0]} points to {voxel_points.shape[0]} voxels")

```

The `voxelize` function returns a tuple containing `morton_codes` (sorted int64 keys) and `voxel_points` (representative coordinates per voxel).

### Frustum Culling with frustum_cull_ext

Frustum culling removes points outside the camera view before rendering, reducing memory bandwidth for the neural mapper.

```python
import frustum_cull_ext as fc

intrinsics = torch.tensor([[525.0, 0.0, 319.5],
                           [0.0, 525.0, 239.5],
                           [0.0, 0.0, 1.0]], device='cuda')
pose = torch.eye(4, device='cuda')
near, far = 0.1, 10.0

visible_points = fc.cull(voxel_points, intrinsics, pose, near, far)

```

The `cull` function accepts camera intrinsics, pose matrices, and clipping planes to return a filtered point tensor.

## Performance Benefits of GPU-Native Processing

By building these custom CUDA extensions for voxelization and frustum culling in LingBot-Map, the pipeline achieves **zero-copy GPU processing** where point data remains on device memory throughout preprocessing. This eliminates costly host-device transfers that would otherwise bottleneck real-time SLAM operations.

The Morton ordering generated by `visibility_kernel.cu` creates spatially coherent data layouts that improve cache utilization during downstream transformer attention computations. Because the kernels are pure CUDA, they automatically leverage hardware-specific optimizations including shared memory tiling and warp-level primitives.

## Summary

- **Build location**: Run `python setup.py build_ext --inplace` from `demo_render/render_cuda_ext/`
- **Source kernels**: `preprocess/points_visibility/visibility_kernel.cu` (voxelization) and `frustum_cull.cu` (culling)
- **Python modules**: Import as `voxel_morton_ext` and `frustum_cull_ext`
- **Integration**: Used in [`lingbot_map/vis/rgbd_render.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/vis/rgbd_render.py) for real-time rendering
- **Key benefit**: Zero-copy processing with deterministic Morton ordering for scalable 3-D reconstruction

## Frequently Asked Questions

### Do I need a specific CUDA version to build these extensions?

You must use a CUDA toolkit version compatible with your installed PyTorch distribution. The setup script uses `torch.utils.cpp_extension`, which automatically detects the CUDA version used to compile PyTorch and applies matching compiler flags. Mismatched versions will raise errors during the compilation of `visibility_kernel.cu` or `frustum_cull.cu`.

### Can I modify the voxel size parameters without rebuilding?

Yes. The `voxel_size` parameter in `vm.voxelize()` is a runtime argument, so you can adjust grid resolution without recompiling the extension. However, if you modify the kernel logic itself—such as changing the Morton code bit depth or culling algorithms—you must rerun `setup.py build_ext --inplace` to regenerate the shared objects.

### How do I verify the extensions built correctly?

After compilation, check that `voxel_morton_ext.*.so` and `frustum_cull_ext.*.so` exist in `demo_render/render_cuda_ext/`. Then verify Python can load them:

```python
import voxel_morton_ext, frustum_cull_ext
print("Extensions loaded successfully")

```

If import errors occur, ensure your `LD_LIBRARY_PATH` includes the CUDA toolkit libraries and that PyTorch can locate the compiled modules.

### Are these extensions compatible with torch.compile?

The extensions expose kernels through `torch.autograd.Function` wrappers, making them compatible with PyTorch's eager mode and `torch.compile` graphs. Since the CUDA implementations are opaque to the Python tracer, they are treated as black-box operators during compilation, preserving their performance characteristics within optimized execution graphs.