How to Build Custom CUDA Extensions for Voxelization and Frustum Culling in LingBot-Map
LingBot-Map accelerates 3-D reconstruction by compiling CUDA kernels for voxelization and frustum culling into Python-accessible extensions using PyTorch's cpp_extension utilities.
LingBot-Map is an open-source spatial mapping framework that leverages GPU acceleration for real-time dense reconstruction. To build custom CUDA extensions for voxelization and frustum culling in LingBot-Map, you will compile kernels from preprocess/points_visibility/ into loadable modules that enable zero-copy processing of point cloud data.
Prerequisites for Building LingBot-Map CUDA Extensions
Before compiling the extensions, ensure your environment meets the hardware and software requirements for CUDA development.
System Requirements
You need an NVIDIA GPU with compute capability compatible with your CUDA toolkit version. The build process requires the CUDA toolkit to be installed system-wide, along with a C++ compiler (GCC on Linux or MSVC on Windows) that matches your PyTorch CUDA version.
Python Dependencies
Install PyTorch with CUDA support and NVIDIA Kaolin, which provides helper geometry utilities used by the kernels.
pip install torch torchvision
pip install --index-url https://pypi.org/simple \
kaolin -f https://nvidia-kaolin.s3.us-east-2.amazonaws.com/torch-2.8.0_cu128.html
Compiling the Voxelization and Frustum Culling Kernels
The build process converts raw CUDA source files into loadable Python modules through PyTorch's JIT compilation infrastructure.
Understanding the Source Files
Two CUDA kernels provide the core functionality according to the LingBot-Map source code:
preprocess/points_visibility/visibility_kernel.cu– Implements sparse voxel grid construction and Morton code generation for spatial sorting.preprocess/points_visibility/frustum_cull.cu– Implements view-frustum intersection tests to filter points outside the camera view.
These files contain kernel implementations that use shared memory tiling and warp-level reductions for optimal performance.
Running the Build Script
Navigate to the extension directory and execute the setup script. In demo_render/render_cuda_ext/setup.py, PyTorch's cpp_extension utilities define two modules: voxel_morton_ext and frustum_cull_ext.
cd demo_render/render_cuda_ext
python setup.py build_ext --inplace
cd ../..
This command compiles the kernels into shared objects (voxel_morton_ext.*.so and frustum_cull_ext.*.so) located next to the setup script. The --inplace flag ensures the extensions are importable from the demo_render directory structure.
Integrating the Extensions into the Rendering Pipeline
Once built, the extensions are imported by the rendering system to perform GPU-accelerated preprocessing. The main consumer is lingbot_map/vis/rgbd_render.py, which imports the modules as:
import voxel_morton_ext as vm
import frustum_cull_ext as fc
Voxelization with voxel_morton_ext
The voxelization kernel inserts points into a sparse grid and returns sorted Morton codes. This enables deterministic spatial ordering required by the transformer-based map encoder.
import torch
import voxel_morton_ext as vm
points = torch.randn(1_000_000, 3, device='cuda')
voxel_size = 0.05
morton_codes, voxel_points = vm.voxelize(points, voxel_size)
print(f"Reduced {points.shape[0]} points to {voxel_points.shape[0]} voxels")
The voxelize function returns a tuple containing morton_codes (sorted int64 keys) and voxel_points (representative coordinates per voxel).
Frustum Culling with frustum_cull_ext
Frustum culling removes points outside the camera view before rendering, reducing memory bandwidth for the neural mapper.
import frustum_cull_ext as fc
intrinsics = torch.tensor([[525.0, 0.0, 319.5],
[0.0, 525.0, 239.5],
[0.0, 0.0, 1.0]], device='cuda')
pose = torch.eye(4, device='cuda')
near, far = 0.1, 10.0
visible_points = fc.cull(voxel_points, intrinsics, pose, near, far)
The cull function accepts camera intrinsics, pose matrices, and clipping planes to return a filtered point tensor.
Performance Benefits of GPU-Native Processing
By building these custom CUDA extensions for voxelization and frustum culling in LingBot-Map, the pipeline achieves zero-copy GPU processing where point data remains on device memory throughout preprocessing. This eliminates costly host-device transfers that would otherwise bottleneck real-time SLAM operations.
The Morton ordering generated by visibility_kernel.cu creates spatially coherent data layouts that improve cache utilization during downstream transformer attention computations. Because the kernels are pure CUDA, they automatically leverage hardware-specific optimizations including shared memory tiling and warp-level primitives.
Summary
- Build location: Run
python setup.py build_ext --inplacefromdemo_render/render_cuda_ext/ - Source kernels:
preprocess/points_visibility/visibility_kernel.cu(voxelization) andfrustum_cull.cu(culling) - Python modules: Import as
voxel_morton_extandfrustum_cull_ext - Integration: Used in
lingbot_map/vis/rgbd_render.pyfor real-time rendering - Key benefit: Zero-copy processing with deterministic Morton ordering for scalable 3-D reconstruction
Frequently Asked Questions
Do I need a specific CUDA version to build these extensions?
You must use a CUDA toolkit version compatible with your installed PyTorch distribution. The setup script uses torch.utils.cpp_extension, which automatically detects the CUDA version used to compile PyTorch and applies matching compiler flags. Mismatched versions will raise errors during the compilation of visibility_kernel.cu or frustum_cull.cu.
Can I modify the voxel size parameters without rebuilding?
Yes. The voxel_size parameter in vm.voxelize() is a runtime argument, so you can adjust grid resolution without recompiling the extension. However, if you modify the kernel logic itself—such as changing the Morton code bit depth or culling algorithms—you must rerun setup.py build_ext --inplace to regenerate the shared objects.
How do I verify the extensions built correctly?
After compilation, check that voxel_morton_ext.*.so and frustum_cull_ext.*.so exist in demo_render/render_cuda_ext/. Then verify Python can load them:
import voxel_morton_ext, frustum_cull_ext
print("Extensions loaded successfully")
If import errors occur, ensure your LD_LIBRARY_PATH includes the CUDA toolkit libraries and that PyTorch can locate the compiled modules.
Are these extensions compatible with torch.compile?
The extensions expose kernels through torch.autograd.Function wrappers, making them compatible with PyTorch's eager mode and torch.compile graphs. Since the CUDA implementations are opaque to the Python tracer, they are treated as black-box operators during compilation, preserving their performance characteristics within optimized execution graphs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →