How Kubernetes DRA ResourceClaim Rendering Translates Hardware into Allocatable Resources

Kubernetes DRA ResourceClaim rendering converts LLM model specifications—parameter count, quantization, and throughput targets—into CEL-based selectors that match against node-level hardware attributes like memory capacity and bandwidth.

The llmfit project (available at AlexsJones/llmfit) implements a specialized Device Resource Allocation (DRA) driver that bridges abstract AI model requirements with concrete cluster resources. Understanding how Kubernetes DRA ResourceClaim rendering functions enables operators to schedule GPU and high-bandwidth memory workloads efficiently. The core logic resides in llmfit-core/src/claim.rs, where raw hardware specs transform into scheduler-compatible YAML manifests.

The Translation Pipeline: From Model Specs to CEL Selectors

The rendering process follows a deterministic seven-step pipeline that translates human-readable model requirements into machine-evaluable resource constraints.

Step 1: Deriving Model Weight Constants

The process begins in llmfit-core/src/models.rs by calculating the model’s on-disk footprint. The LlmModel::estimate_disk_gb method derives the weights size from the parameter count and selected quantization format (e.g., Q4_K_M). This establishes the baseline storage requirement that will later inform memory allocation.

Step 2: Calculating Memory Floors

In llmfit-core/src/claim.rs, the fit_bounds function computes the minimum VRAM requirement. When using the model’s default quantization, the system leverages the database-provided min_vram_gb value, which already incorporates KV cache and runtime headroom. For quantization overrides, the weight size is multiplied by the constant WEIGHTS_HEADROOM = 1.2 to ensure sufficient working memory. The result is always rounded up to the nearest whole GiB.

Step 3: Computing Bandwidth Requirements

To satisfy latency objectives, the renderer translates tokens-per-second targets into memory bandwidth minimums. The formula implemented in fit_bounds rearranges the throughput equation to:


bandwidth ≥ (min_tps × weights_gb × 100) / efficiency_pct

This calculation ensures that the selected device can feed the model weights to compute units quickly enough to achieve the requested inference speed. The value rounds up to whole GB/s.

Step 4: Sanitizing Claim Names

The claim_name function in llmfit-core/src/claim.rs transforms model identifiers into DNS-safe Kubernetes labels. This involves lower-casing, replacing non-alphanumeric characters with dashes, collapsing duplicate dashes, and truncating to 48 characters. The result serves as the metadata.name in the generated ResourceClaim.

Step 5: Emitting CEL Selector Expressions

The render function constructs a Common Expression Language (CEL) selector that evaluates three critical conditions against device attributes published by the DRA driver:

  • The device belongs to the llmfit.ai driver domain
  • device.capacity['llmfit.ai'].memory meets or exceeds the calculated memory floor
  • device.attributes['llmfit.ai'].memoryBandwidthGBs satisfies the bandwidth requirement

Guard clauses (e.g., 'memory' in device.capacity['llmfit.ai']) prevent runtime errors when nodes lack the expected attributes, causing graceful non-matches rather than crashes.

Step 6: Adding Auditability Headers

The generated YAML includes a human-readable comment block disclosing the model’s parameter count, quantization format, weight size, and the derivation formulas. This transparency supports GitOps workflows and troubleshooting when claims are stored in version control.

Step 7: Scheduler Integration

The final output—available via render for YAML or render_json for programmatic consumption—contains a complete ResourceClaim or ResourceClaimTemplate. The Kubernetes scheduler evaluates the embedded CEL expression against each node’s Device objects, selecting only nodes where llmfit-core/src/hardware.rs has published sufficient capacity and bandwidth.

Core Implementation in llmfit-core/src/claim.rs

The translation logic centers on three primary functions exported from the core library:

  • fit_bounds – Computes memory_gb and min_bw from model metadata and performance targets
  • render – Generates a complete ResourceClaim YAML with CEL selectors and audit comments
  • render_json – Produces a structured JSON representation for API consumption by downstream controllers

These functions consume data from llmfit-core/src/models.rs (model specifications) and target attributes detected by llmfit-core/src/hardware.rs (system RAM, GPU memory, and bandwidth capabilities).

Practical Rendering Examples

The following Rust code demonstrates how to generate a ResourceClaim for a 7B parameter model requiring 20 tokens per second at 55% efficiency:

use llmfit_core::claim::{ClaimTarget, render, render_json};
use llmfit_core::models::LlmModel;

// Load model from the embedded catalog
let model: LlmModel = /* ... */;

// Configure performance targets
let mut target = ClaimTarget::default();
target.min_tps = 20.0;
target.efficiency_pct = 55;

// Generate YAML ResourceClaim
let yaml = render(&model, &target).expect("render failed");
println!("{}", yaml);

The resulting YAML contains a CEL expression requiring at least 6Gi of memory and 148GB/s of bandwidth:

apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
  name: test-model-7b-fit
spec:
  devices:
    requests:
      - name: model
        exactly:
          deviceClassName: llmfit.ai
          selectors:
            - cel:
                expression: >-
                  'memory' in device.capacity['llmfit.ai'] &&
                  device.capacity['llmfit.ai'].memory.compareTo(quantity('6Gi')) >= 0 &&
                  'memoryBandwidthGBs' in device.attributes['llmfit.ai'] &&
                  device.attributes['llmfit.ai'].memoryBandwidthGBs >= 148 &&
                  'healthy' in device.attributes['llmfit.ai'] &&
                  device.attributes['llmfit.ai'].healthy

For programmatic consumption, the JSON endpoint exposed in llmfit-tui/src/serve_api.rs returns structured data:

let json = render_json(&model, &target, env!("CARGO_PKG_VERSION"))
    .expect("JSON render failed");

Summary

  • Kubernetes DRA ResourceClaim rendering translates abstract LLM requirements into concrete hardware selectors using CEL expressions evaluated by the scheduler.
  • The fit_bounds function in llmfit-core/src/claim.rs calculates memory floors using a 1.2x headroom multiplier and derives bandwidth requirements from tokens-per-second targets.
  • CEL selectors generated by render match against device.capacity and device.attributes published by the llmfit-dra driver, ensuring nodes can satisfy both capacity and throughput constraints.
  • DNS-safe claim names are automatically generated via claim_name to comply with Kubernetes labeling rules.
  • Hardware detection in llmfit-core/src/hardware.rs populates the attributes that the scheduler evaluates against these claims.

Frequently Asked Questions

How does the CEL expression prevent scheduling on incompatible nodes?

The expression includes guard clauses that check for attribute existence (e.g., 'memory' in device.capacity['llmfit.ai']) before evaluating magnitude comparisons. If a node lacks the llmfit.ai device class or the specific memory attributes, the CEL evaluation returns false, excluding that node from the scheduling decision without causing a runtime error.

Why does llmfit use a 1.2x headroom factor for quantization overrides?

The WEIGHTS_HEADROOM = 1.2 constant in llmfit-core/src/claim.rs accounts for KV cache memory, activation storage, and runtime overhead not captured in the raw weights calculation. When users override the default quantization, the database-provided min_vram_gb (which includes these factors) is unavailable, necessitating this conservative multiplier to prevent out-of-memory errors during inference.

Can the generated ResourceClaims target specific GPU models?

The current implementation in llmfit-core/src/claim.rs focuses on capacity and bandwidth constraints rather than device model names. However, the CEL expression architecture supports arbitrary attribute matching. Operators can extend the render function to include additional selectors based on attributes like device.attributes['llmfit.ai'].modelName if the DRA driver publishes such metadata.

What is the difference between render and render_json?

The render function produces a complete ResourceClaim YAML suitable for kubectl apply or Helm charts, including human-readable audit comments. The render_json function returns a structured JSON object consumed by the API exposed in llmfit-tui/src/serve_api.rs, enabling programmatic claim generation without YAML parsing overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →