How llmfit Detects Unified Memory Systems Like Apple Silicon

llmfit treats Apple Silicon as a unified memory system by mapping total system RAM to GPU VRAM, setting the unified_memory flag to true, and querying Metal's "recommendedMaxWorkingSetSize" to determine the GPU's available working set.

The llmfit repository implements specialized hardware detection for unified memory systems like Apple Silicon in llmfit-core/src/hardware.rs. Unlike discrete GPUs with dedicated VRAM, these systems require distinct logic to handle shared physical memory pools between CPU and GPU.

Detection Flow in SystemSpecs::detect()

The hardware detection process begins in SystemSpecs::detect() at line 272 of llmfit-core/src/hardware.rs. When the constructor runs, it first attempts to detect Apple Silicon hardware by calling Self::detect_apple_gpu(total_ram_gb).

If the machine is running on Apple Silicon, this function returns the total system RAM in GiB; otherwise, it returns None. This check determines whether the system uses a unified memory architecture.

The detect_apple_gpu Function Implementation

The detect_apple_gpu function, defined around line 1600, identifies Apple Silicon devices and returns the total system RAM as the available GPU memory. This approach recognizes that on unified memory systems, the GPU shares the same physical memory pool as the CPU.

When a value is returned, the detection logic constructs a GpuInfo struct and pushes it to the GPU list between lines 272-285:

gpus.push(GpuInfo {
    name,
    vram_gb: Some(vram),          // = total RAM
    backend: GpuBackend::Metal,   // Metal is the only backend on macOS
    count: 1,
    unified_memory: true,
});

The unified_memory field is set to true to signal that the GPU and CPU share physical memory, while the backend is forced to Metal because Apple Silicon devices expose their GPU exclusively through the Metal API.

Querying Available GPU Memory with Metal

After constructing the GPU list, SystemSpecs::detect() checks the unified_memory flag. When true, it calls detect_gpu_available_gb() at lines 334-338 to query the actual GPU-available memory using the Metal API.

This function reads "recommendedMaxWorkingSetSize", which represents the amount of the shared pool the GPU is allowed to wire up. This value is typically lower than total system RAM and represents the practical limit for GPU memory usage. The result is stored in gpu_available_gb within the final SystemSpecs struct.

Resulting SystemSpecs Structure

The final SystemSpecs struct contains the following fields for unified memory systems:

  • total_ram_gb – Physical RAM size
  • gpu_vram_gb – Identical to total_ram_gb for Apple Silicon
  • unified_memory: true – Signals shared memory architecture
  • backend: GpuBackend::Metal – Inference backend for model execution
  • gpu_available_gb – Metal's recommended working set size

This information enables the fit-scoring logic to treat the GPU as having VRAM-first memory while applying correct capacity limits when estimating model fit.

Practical Implementation Example

To detect unified memory systems in your own code using llmfit:

use llmfit_core::hardware::SystemSpecs;

fn main() {
    // Detect the current machine's hardware.
    let specs = SystemSpecs::detect();

    // Unified-memory detection (Apple Silicon, NVIDIA Grace, AMD APU, etc.)
    if specs.unified_memory {
        println!("Unified-memory system detected!");
        println!("Total RAM (and GPU VRAM): {:.2} GiB", specs.total_ram_gb);
        println!(
            "GPU-available (Metal's working-set limit): {:?} GiB",
            specs.gpu_available_gb
        );
        println!("Backend used for inference: {:?}", specs.backend);
    } else {
        println!("Discrete-GPU system – normal VRAM handling.");
    }
}

Running this on an Apple Silicon Mac outputs:


Unified-memory system detected!
Total RAM (and GPU VRAM): 32.00 GiB
GPU-available (Metal's working-set limit): Some(24.00) GiB
Backend used for inference: Metal

Summary

  • llmfit detects Apple Silicon by calling detect_apple_gpu() in llmfit-core/src/hardware.rs, which returns total system RAM when running on macOS ARM architecture.
  • The system sets unified_memory: true in the GpuInfo struct to indicate shared physical memory between CPU and GPU.
  • The Metal backend is forced for all Apple Silicon detection, as Metal is the exclusive GPU API on macOS.
  • Available GPU memory is determined by querying Metal's "recommendedMaxWorkingSetSize" rather than total RAM, providing accurate working set limits.
  • The resulting SystemSpecs enables the fit-scoring engine to correctly evaluate model compatibility on unified memory architectures.

Frequently Asked Questions

How does llmfit distinguish between Apple Silicon and discrete GPUs?

llmfit distinguishes Apple Silicon by calling the detect_apple_gpu function at line 1600 in llmfit-core/src/hardware.rs. This function checks for macOS ARM architecture and returns the total system RAM as VRAM if detected. For discrete GPUs, different detection paths populate separate VRAM values without setting the unified_memory flag.

Why does llmfit use Metal's "recommendedMaxWorkingSetSize" instead of total RAM?

The "recommendedMaxWorkingSetSize" query provides the amount of unified memory the GPU is actually allowed to wire up, which is typically lower than total system RAM. This value represents the practical GPU memory limit available for model inference, preventing overallocation that could starve the system or other processes.

Can llmfit detect other unified memory systems besides Apple Silicon?

While the current implementation in llmfit-core/src/hardware.rs specifically targets Apple Silicon through the detect_apple_gpu function, the architecture supports other unified memory systems. The unified_memory boolean flag in GpuInfo is designed to handle any unified memory architecture, though additional detection logic would be required for systems like NVIDIA Grace or AMD APUs.

What happens if recommendedMaxWorkingSetSize is not available?

If the Metal API cannot provide the "recommendedMaxWorkingSetSize" value, the detect_gpu_available_gb() function returns None for gpu_available_gb. In this case, the fit-scoring logic falls back to using the total system RAM value, though this represents the maximum theoretical memory rather than the GPU's practical working set limit.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →