How to Build Custom CUDA Extensions for Voxelization and Frustum Culling in LingBot-Map

LingBot-Map accelerates 3-D reconstruction by compiling CUDA kernels for voxelization and frustum culling into Python-accessible extensions using PyTorch's cpp_extension utilities.

LingBot-Map is an open-source spatial mapping framework that leverages GPU acceleration for real-time dense reconstruction. To build custom CUDA extensions for voxelization and frustum culling in LingBot-Map, you will compile kernels from preprocess/points_visibility/ into loadable modules that enable zero-copy processing of point cloud data.

Prerequisites for Building LingBot-Map CUDA Extensions

Before compiling the extensions, ensure your environment meets the hardware and software requirements for CUDA development.

System Requirements

You need an NVIDIA GPU with compute capability compatible with your CUDA toolkit version. The build process requires the CUDA toolkit to be installed system-wide, along with a C++ compiler (GCC on Linux or MSVC on Windows) that matches your PyTorch CUDA version.

Python Dependencies

Install PyTorch with CUDA support and NVIDIA Kaolin, which provides helper geometry utilities used by the kernels.

pip install torch torchvision
pip install --index-url https://pypi.org/simple \
    kaolin -f https://nvidia-kaolin.s3.us-east-2.amazonaws.com/torch-2.8.0_cu128.html

Compiling the Voxelization and Frustum Culling Kernels

The build process converts raw CUDA source files into loadable Python modules through PyTorch's JIT compilation infrastructure.

Understanding the Source Files

Two CUDA kernels provide the core functionality according to the LingBot-Map source code:

  • preprocess/points_visibility/visibility_kernel.cu – Implements sparse voxel grid construction and Morton code generation for spatial sorting.
  • preprocess/points_visibility/frustum_cull.cu – Implements view-frustum intersection tests to filter points outside the camera view.

These files contain kernel implementations that use shared memory tiling and warp-level reductions for optimal performance.

Running the Build Script

Navigate to the extension directory and execute the setup script. In demo_render/render_cuda_ext/setup.py, PyTorch's cpp_extension utilities define two modules: voxel_morton_ext and frustum_cull_ext.

cd demo_render/render_cuda_ext
python setup.py build_ext --inplace
cd ../..

This command compiles the kernels into shared objects (voxel_morton_ext.*.so and frustum_cull_ext.*.so) located next to the setup script. The --inplace flag ensures the extensions are importable from the demo_render directory structure.

Integrating the Extensions into the Rendering Pipeline

Once built, the extensions are imported by the rendering system to perform GPU-accelerated preprocessing. The main consumer is lingbot_map/vis/rgbd_render.py, which imports the modules as:

import voxel_morton_ext as vm
import frustum_cull_ext as fc

Voxelization with voxel_morton_ext

The voxelization kernel inserts points into a sparse grid and returns sorted Morton codes. This enables deterministic spatial ordering required by the transformer-based map encoder.

import torch
import voxel_morton_ext as vm

points = torch.randn(1_000_000, 3, device='cuda')
voxel_size = 0.05

morton_codes, voxel_points = vm.voxelize(points, voxel_size)
print(f"Reduced {points.shape[0]} points to {voxel_points.shape[0]} voxels")

The voxelize function returns a tuple containing morton_codes (sorted int64 keys) and voxel_points (representative coordinates per voxel).

Frustum Culling with frustum_cull_ext

Frustum culling removes points outside the camera view before rendering, reducing memory bandwidth for the neural mapper.

import frustum_cull_ext as fc

intrinsics = torch.tensor([[525.0, 0.0, 319.5],
                           [0.0, 525.0, 239.5],
                           [0.0, 0.0, 1.0]], device='cuda')
pose = torch.eye(4, device='cuda')
near, far = 0.1, 10.0

visible_points = fc.cull(voxel_points, intrinsics, pose, near, far)

The cull function accepts camera intrinsics, pose matrices, and clipping planes to return a filtered point tensor.

Performance Benefits of GPU-Native Processing

By building these custom CUDA extensions for voxelization and frustum culling in LingBot-Map, the pipeline achieves zero-copy GPU processing where point data remains on device memory throughout preprocessing. This eliminates costly host-device transfers that would otherwise bottleneck real-time SLAM operations.

The Morton ordering generated by visibility_kernel.cu creates spatially coherent data layouts that improve cache utilization during downstream transformer attention computations. Because the kernels are pure CUDA, they automatically leverage hardware-specific optimizations including shared memory tiling and warp-level primitives.

Summary

  • Build location: Run python setup.py build_ext --inplace from demo_render/render_cuda_ext/
  • Source kernels: preprocess/points_visibility/visibility_kernel.cu (voxelization) and frustum_cull.cu (culling)
  • Python modules: Import as voxel_morton_ext and frustum_cull_ext
  • Integration: Used in lingbot_map/vis/rgbd_render.py for real-time rendering
  • Key benefit: Zero-copy processing with deterministic Morton ordering for scalable 3-D reconstruction

Frequently Asked Questions

Do I need a specific CUDA version to build these extensions?

You must use a CUDA toolkit version compatible with your installed PyTorch distribution. The setup script uses torch.utils.cpp_extension, which automatically detects the CUDA version used to compile PyTorch and applies matching compiler flags. Mismatched versions will raise errors during the compilation of visibility_kernel.cu or frustum_cull.cu.

Can I modify the voxel size parameters without rebuilding?

Yes. The voxel_size parameter in vm.voxelize() is a runtime argument, so you can adjust grid resolution without recompiling the extension. However, if you modify the kernel logic itself—such as changing the Morton code bit depth or culling algorithms—you must rerun setup.py build_ext --inplace to regenerate the shared objects.

How do I verify the extensions built correctly?

After compilation, check that voxel_morton_ext.*.so and frustum_cull_ext.*.so exist in demo_render/render_cuda_ext/. Then verify Python can load them:

import voxel_morton_ext, frustum_cull_ext
print("Extensions loaded successfully")

If import errors occur, ensure your LD_LIBRARY_PATH includes the CUDA toolkit libraries and that PyTorch can locate the compiled modules.

Are these extensions compatible with torch.compile?

The extensions expose kernels through torch.autograd.Function wrappers, making them compatible with PyTorch's eager mode and torch.compile graphs. Since the CUDA implementations are opaque to the Python tracer, they are treated as black-box operators during compilation, preserving their performance characteristics within optimized execution graphs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →