How ModelFit Memory and VRAM Fields Feed Into Kubernetes DRA ResourceClaim Specs

The min_ram_gb and min_vram_gb fields from a ModelFit struct populate the memory and gpu.memory resource requests in Kubernetes Dynamic Resource Allocation (DRA) ResourceClaim specifications, with optional estimated_memory_gb and estimated_vram_gb values overriding these when more precise runtime measurements are available.

The llmfit open-source project automates hardware provisioning for large language model inference by translating model analysis data into Kubernetes-native objects. Understanding how memory and VRAM calculations flow from the internal ModelFit structure into DRA ResourceClaim specifications enables operators to optimize GPU scheduling and resource reservation for AI workloads.

ModelFit Memory and VRAM Fields

In llmfit-core/src/fit.rs, the ModelFit struct captures resource requirements through four key memory-related fields:

  • min_ram_gb (or ram_gb): Minimum system RAM required for CPU-only inference, measured in gigabytes.
  • min_vram_gb (or vram_gb): Minimum GPU VRAM required for GPU-accelerated inference, measured in gigabytes.
  • estimated_memory_gb: Optional field providing a tighter estimate of total memory consumption for the chosen runtime configuration.
  • estimated_vram_gb: Optional field providing a tighter estimate of GPU memory usage for the specific inference engine and quantization settings.

Converting ModelFit to DRA ResourceClaim Specifications

The conversion logic resides in llmfit-core/src/claim.rs, where the resource_claim_from_fit function transforms these Rust struct fields into a valid Kubernetes ResourceClaim manifest compatible with the DRA API.

Resource Key Mapping

According to the source code implementation, the function constructs a HashMap<String, Quantity> to populate the resources field in the ResourceClaim spec:

ModelFit Field DRA ResourceClaim Key Resource Class
min_ram_gb memory Standard system memory
min_vram_gb gpu.memory GPU memory pool
estimated_memory_gb memory Overrides minimum when available
estimated_vram_gb gpu.memory Overrides minimum when available

ResourceClaim YAML Structure

The resulting ResourceClaim manifest targets the gpu-memory-class resource class and formats values as Kubernetes quantities with Gi suffixes:

apiVersion: resource.k8s.io/v1alpha2
kind: ResourceClaim
metadata:
  name: <model-name>-claim
spec:
  resourceClassName: gpu-memory-class
  resources:
    memory: "<value>Gi"
    gpu.memory: "<value>Gi"

Implementation Details in claim.rs

Priority Logic for Memory Values

The implementation in llmfit-core/src/claim.rs applies a precedence check when building resource requests. If estimated values exist, they take priority over minimum requirements:

// Simplified representation of the mapping logic in claim.rs
let mut resources: HashMap<String, String> = HashMap::new();

// System RAM: use estimate if available, else fall back to minimum
let memory_val = fit.estimated_memory_gb
    .unwrap_or(fit.min_ram_gb);
resources.insert("memory".to_string(), format!("{}Gi", memory_val));

// GPU VRAM: use estimate if available, else fall back to minimum  
let vram_val = fit.estimated_vram_gb
    .unwrap_or(fit.min_vram_gb);
resources.insert("gpu.memory".to_string(), format!("{}Gi", vram_val));

Generating ResourceClaims from Model Fits

To generate a ResourceClaim from an analyzed model:

use llmfit_core::fit::ModelFit;
use llmfit_core::claim::resource_claim_from_fit;

// Assuming `fit` is a ModelFit instance from model analysis
let claim = resource_claim_from_fit(&fit);

// Output YAML for kubectl application
println!("{}", serde_yaml::to_string(&claim).unwrap());

ResourceClaim Output Example

For a 7B parameter model analysis that calculates 12GB system RAM and 8GB VRAM requirements, the generated ResourceClaim appears as:

apiVersion: resource.k8s.io/v1alpha2
kind: ResourceClaim
metadata:
  name: llama-7b-claim
spec:
  resourceClassName: gpu-memory-class
  resources:
    memory: "12Gi"
    gpu.memory: "8Gi"

Summary

  • min_ram_gb and min_vram_gb serve as the baseline fields from ModelFit that feed into Kubernetes DRA ResourceClaim specifications as memory and gpu.memory respectively.
  • estimated_memory_gb and estimated_vram_gb provide optional overrides when the runtime configuration allows for more precise resource constraints than the generic minimums.
  • The conversion logic in llmfit-core/src/claim.rs formats all values as Kubernetes quantities with Gi (gibibyte) suffixes within a HashMap<String, Quantity> structure.
  • Resulting ResourceClaims target the gpu-memory-class resource class for proper GPU scheduling through the DRA controller.

Frequently Asked Questions

What determines whether estimated or minimum memory values are used in the ResourceClaim?

The logic in llmfit-core/src/claim.rs checks for the presence of estimated_memory_gb and estimated_vram_gb fields first. If these optional values exist, they take precedence over min_ram_gb and min_vram_gb respectively, allowing the ResourceClaim to request resources that match the specific runtime configuration rather than generic minimum requirements.

How are memory values formatted in the Kubernetes ResourceClaim?

Raw gigabyte values from the ModelFit struct are converted to Kubernetes Quantity strings with a Gi suffix indicating gibibytes. For example, a min_vram_gb value of 8 becomes the string "8Gi" in the ResourceClaim's gpu.memory field.

Can these ResourceClaims be used with standard Kubernetes resource limits?

DRA ResourceClaims function alongside standard resource limits but are processed specifically by the Dynamic Resource Allocation controller. While the memory field corresponds to traditional RAM, the gpu.memory resource requires a DRA-enabled cluster with an appropriate resource driver installed to handle GPU memory scheduling.

Where are the ModelFit struct and conversion logic defined in the llmfit repository?

The ModelFit struct and its memory-related fields are defined in llmfit-core/src/fit.rs, while the conversion logic that maps these fields to Kubernetes DRA ResourceClaim specifications resides in llmfit-core/src/claim.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →