How Modly Detects GPU Hardware and Manages CUDA Version Compatibility: A Technical Deep Dive

Modly uses PyTorch runtime APIs, environment variable inspection, and conditional setuptools logic in api/texture_baker/setup.py and api/routers/extensions.py to detect GPU capabilities and automatically compile CUDA or Metal extensions while falling back to CPU-only builds when hardware is unavailable.

The lightningpixel/modly repository implements a robust hardware abstraction layer that determines GPU availability at both runtime and build time. This article examines the exact mechanisms Modly uses to detect GPU hardware, validate CUDA toolkit installations, and manage cross-platform compatibility across NVIDIA, AMD, and Apple Silicon architectures.

Runtime GPU Detection with PyTorch APIs

Modly’s hardware detection begins with PyTorch’s CUDA runtime interfaces. In api/routers/extensions.py, the application calls torch.cuda.is_available() to determine whether a CUDA-capable GPU is present on the host system.

If a GPU is detected, Modly retrieves the compute capability of the first device using torch.cuda.get_device_capability(0). This function returns a tuple of major and minor version numbers (e.g., (7, 5) for an RTX 3080) that the extensions router uses to gate CUDA-specific features and select appropriate kernel implementations.


# api/routers/extensions.py

if torch.cuda.is_available():
    major, minor = torch.cuda.get_device_capability(0)  # e.g., (7, 5)

    # Logic gates CUDA-specific API routes based on compute capability

Environment Variable Configuration and CUDA Toolkit Validation

Before compiling native extensions, Modly validates the presence of the CUDA toolkit through environment variables. The build script in api/texture_baker/setup.py inspects CUDA_HOME to locate the CUDA installation directory. If this variable is unset, the system falls back to CPU-only compilation regardless of GPU presence.

The USE_CUDA flag provides manual override capability. By default, Modly enables CUDA when torch.cuda.is_available() returns True, but developers can force CPU-only builds by setting USE_CUDA=0.


# api/texture_baker/setup.py

use_cuda = os.getenv("USE_CUDA", "1" if torch.cuda.is_available() else "0") == "1"
use_cuda = use_cuda and CUDA_HOME is not None  # Ensure toolkit is present

Conditional Compilation: CUDA vs CPU Extensions

When both GPU hardware and the CUDA toolkit are confirmed, Modly switches the extension type from CppExtension to CUDAExtension. This change instructs setuptools to compile .cu source files and link against cudart and c10_cuda libraries.

The build script also injects the THRUST_IGNORE_CUB_VERSION_CHECK macro to suppress version mismatch warnings between Thrust and CUB components. CUDA source files are globbed from the csrc directory only when use_cuda evaluates to True.


# api/texture_baker/setup.py

extension = CUDAExtension if use_cuda else CppExtension

if use_cuda:
    define_macros += [("THRUST_IGNORE_CUB_VERSION_CHECK", None)]
    sources += glob.glob(os.path.join(this_dir, library_name, "csrc", "**", "*.cu"), recursive=True)
    libraries += ["cudart", "c10_cuda"]

Apple Silicon Support and Metal Fallback

On macOS systems, Modly detects Apple Silicon GPUs through torch.backends.mps.is_available(). When Metal Performance Shaders (MPS) are available and the USE_METAL environment variable is enabled, the build system compiles Metal-based extensions using .mm (Objective-C++) source files instead of CUDA kernels.

This logic coexists with CUDA checks in the same setup.py script, allowing the repository to maintain a unified build pipeline for heterogeneous hardware targets.


# Metal detection in setup.py

use_metal = os.getenv("USE_METAL", "1" if torch.backends.mps.is_available() else "0") == "1"
if use_metal:
    define_macros += [("WITH_MPS", None)]
    sources += glob.glob(os.path.join(this_dir, library_name, "csrc", "**", "*.mm"), recursive=True)

Graceful Degradation and Cache Management

If no GPU is detected or the CUDA toolkit is missing, Modly prints a warning message and skips CUDA extension compilation, ensuring the framework remains functional on CPU-only systems. After generation tasks complete, api/services/generators/base.py explicitly clears CUDA memory caches using torch.cuda.empty_cache() to prevent memory leaks on systems where CUDA is active.

Summary

  • PyTorch runtime checks in api/routers/extensions.py determine GPU availability and compute capability using torch.cuda.is_available() and torch.cuda.get_device_capability().
  • Environment variables (CUDA_HOME, USE_CUDA) control whether the build system attempts CUDA compilation or falls back to CPU-only extensions.
  • Conditional compilation switches between CppExtension and CUDAExtension, linking against cudart and c10_cuda when appropriate.
  • Apple Silicon detection via torch.backends.mps.is_available() enables Metal-based builds using .mm source files as an alternative to CUDA.
  • Graceful degradation ensures Modly installs and runs on CPU-only machines when GPU hardware or CUDA toolkits are absent.

Frequently Asked Questions

How does Modly detect which CUDA version is installed?

Modly does not explicitly parse the CUDA version string. Instead, it checks for the presence of the CUDA_HOME environment variable and validates that torch.cuda.is_available() returns True. The actual CUDA version compatibility is handled by PyTorch’s underlying CUDA runtime, which ensures that the compiled extensions match the installed toolkit version.

Can I force Modly to compile without CUDA even if I have a GPU?

Yes. Set the environment variable USE_CUDA=0 before running the build. This overrides the automatic GPU detection logic in api/texture_baker/setup.py and forces the use of CppExtension instead of CUDAExtension, resulting in a CPU-only build.

Does Modly support AMD GPUs or Intel Arc GPUs?

The current implementation focuses on NVIDIA CUDA and Apple Metal. AMD and Intel GPU support would require ROCm or SYCL backend implementations, which are not present in the current api/texture_baker/setup.py build logic. The CPU fallback path ensures functionality on these devices but without GPU acceleration.

What happens if CUDA_HOME is set but the GPU is not available?

Modly performs a logical AND operation between the USE_CUDA flag and the presence of CUDA_HOME. However, it also checks torch.cuda.is_available() at build time. If the GPU is not available (e.g., CUDA drivers missing), the build will default to CPU-only compilation even if CUDA_HOME points to a valid toolkit installation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →