How Kubernetes DRA ResourceClaim Rendering Translates Hardware into Allocatable Resources
Kubernetes DRA ResourceClaim rendering converts LLM model specifications—parameter count, quantization, and throughput targets—into CEL-based selectors that match against node-level hardware attributes like memory capacity and bandwidth.
The llmfit project (available at AlexsJones/llmfit) implements a specialized Device Resource Allocation (DRA) driver that bridges abstract AI model requirements with concrete cluster resources. Understanding how Kubernetes DRA ResourceClaim rendering functions enables operators to schedule GPU and high-bandwidth memory workloads efficiently. The core logic resides in llmfit-core/src/claim.rs, where raw hardware specs transform into scheduler-compatible YAML manifests.
The Translation Pipeline: From Model Specs to CEL Selectors
The rendering process follows a deterministic seven-step pipeline that translates human-readable model requirements into machine-evaluable resource constraints.
Step 1: Deriving Model Weight Constants
The process begins in llmfit-core/src/models.rs by calculating the model’s on-disk footprint. The LlmModel::estimate_disk_gb method derives the weights size from the parameter count and selected quantization format (e.g., Q4_K_M). This establishes the baseline storage requirement that will later inform memory allocation.
Step 2: Calculating Memory Floors
In llmfit-core/src/claim.rs, the fit_bounds function computes the minimum VRAM requirement. When using the model’s default quantization, the system leverages the database-provided min_vram_gb value, which already incorporates KV cache and runtime headroom. For quantization overrides, the weight size is multiplied by the constant WEIGHTS_HEADROOM = 1.2 to ensure sufficient working memory. The result is always rounded up to the nearest whole GiB.
Step 3: Computing Bandwidth Requirements
To satisfy latency objectives, the renderer translates tokens-per-second targets into memory bandwidth minimums. The formula implemented in fit_bounds rearranges the throughput equation to:
bandwidth ≥ (min_tps × weights_gb × 100) / efficiency_pct
This calculation ensures that the selected device can feed the model weights to compute units quickly enough to achieve the requested inference speed. The value rounds up to whole GB/s.
Step 4: Sanitizing Claim Names
The claim_name function in llmfit-core/src/claim.rs transforms model identifiers into DNS-safe Kubernetes labels. This involves lower-casing, replacing non-alphanumeric characters with dashes, collapsing duplicate dashes, and truncating to 48 characters. The result serves as the metadata.name in the generated ResourceClaim.
Step 5: Emitting CEL Selector Expressions
The render function constructs a Common Expression Language (CEL) selector that evaluates three critical conditions against device attributes published by the DRA driver:
- The device belongs to the
llmfit.aidriver domain device.capacity['llmfit.ai'].memorymeets or exceeds the calculated memory floordevice.attributes['llmfit.ai'].memoryBandwidthGBssatisfies the bandwidth requirement
Guard clauses (e.g., 'memory' in device.capacity['llmfit.ai']) prevent runtime errors when nodes lack the expected attributes, causing graceful non-matches rather than crashes.
Step 6: Adding Auditability Headers
The generated YAML includes a human-readable comment block disclosing the model’s parameter count, quantization format, weight size, and the derivation formulas. This transparency supports GitOps workflows and troubleshooting when claims are stored in version control.
Step 7: Scheduler Integration
The final output—available via render for YAML or render_json for programmatic consumption—contains a complete ResourceClaim or ResourceClaimTemplate. The Kubernetes scheduler evaluates the embedded CEL expression against each node’s Device objects, selecting only nodes where llmfit-core/src/hardware.rs has published sufficient capacity and bandwidth.
Core Implementation in llmfit-core/src/claim.rs
The translation logic centers on three primary functions exported from the core library:
fit_bounds– Computesmemory_gbandmin_bwfrom model metadata and performance targetsrender– Generates a complete ResourceClaim YAML with CEL selectors and audit commentsrender_json– Produces a structured JSON representation for API consumption by downstream controllers
These functions consume data from llmfit-core/src/models.rs (model specifications) and target attributes detected by llmfit-core/src/hardware.rs (system RAM, GPU memory, and bandwidth capabilities).
Practical Rendering Examples
The following Rust code demonstrates how to generate a ResourceClaim for a 7B parameter model requiring 20 tokens per second at 55% efficiency:
use llmfit_core::claim::{ClaimTarget, render, render_json};
use llmfit_core::models::LlmModel;
// Load model from the embedded catalog
let model: LlmModel = /* ... */;
// Configure performance targets
let mut target = ClaimTarget::default();
target.min_tps = 20.0;
target.efficiency_pct = 55;
// Generate YAML ResourceClaim
let yaml = render(&model, &target).expect("render failed");
println!("{}", yaml);
The resulting YAML contains a CEL expression requiring at least 6Gi of memory and 148GB/s of bandwidth:
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: test-model-7b-fit
spec:
devices:
requests:
- name: model
exactly:
deviceClassName: llmfit.ai
selectors:
- cel:
expression: >-
'memory' in device.capacity['llmfit.ai'] &&
device.capacity['llmfit.ai'].memory.compareTo(quantity('6Gi')) >= 0 &&
'memoryBandwidthGBs' in device.attributes['llmfit.ai'] &&
device.attributes['llmfit.ai'].memoryBandwidthGBs >= 148 &&
'healthy' in device.attributes['llmfit.ai'] &&
device.attributes['llmfit.ai'].healthy
For programmatic consumption, the JSON endpoint exposed in llmfit-tui/src/serve_api.rs returns structured data:
let json = render_json(&model, &target, env!("CARGO_PKG_VERSION"))
.expect("JSON render failed");
Summary
- Kubernetes DRA ResourceClaim rendering translates abstract LLM requirements into concrete hardware selectors using CEL expressions evaluated by the scheduler.
- The
fit_boundsfunction inllmfit-core/src/claim.rscalculates memory floors using a 1.2x headroom multiplier and derives bandwidth requirements from tokens-per-second targets. - CEL selectors generated by
rendermatch againstdevice.capacityanddevice.attributespublished by thellmfit-dradriver, ensuring nodes can satisfy both capacity and throughput constraints. - DNS-safe claim names are automatically generated via
claim_nameto comply with Kubernetes labeling rules. - Hardware detection in
llmfit-core/src/hardware.rspopulates the attributes that the scheduler evaluates against these claims.
Frequently Asked Questions
How does the CEL expression prevent scheduling on incompatible nodes?
The expression includes guard clauses that check for attribute existence (e.g., 'memory' in device.capacity['llmfit.ai']) before evaluating magnitude comparisons. If a node lacks the llmfit.ai device class or the specific memory attributes, the CEL evaluation returns false, excluding that node from the scheduling decision without causing a runtime error.
Why does llmfit use a 1.2x headroom factor for quantization overrides?
The WEIGHTS_HEADROOM = 1.2 constant in llmfit-core/src/claim.rs accounts for KV cache memory, activation storage, and runtime overhead not captured in the raw weights calculation. When users override the default quantization, the database-provided min_vram_gb (which includes these factors) is unavailable, necessitating this conservative multiplier to prevent out-of-memory errors during inference.
Can the generated ResourceClaims target specific GPU models?
The current implementation in llmfit-core/src/claim.rs focuses on capacity and bandwidth constraints rather than device model names. However, the CEL expression architecture supports arbitrary attribute matching. Operators can extend the render function to include additional selectors based on attributes like device.attributes['llmfit.ai'].modelName if the DRA driver publishes such metadata.
What is the difference between render and render_json?
The render function produces a complete ResourceClaim YAML suitable for kubectl apply or Helm charts, including human-readable audit comments. The render_json function returns a structured JSON object consumed by the API exposed in llmfit-tui/src/serve_api.rs, enabling programmatic claim generation without YAML parsing overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →