How llmfit Renders Kubernetes DRA ResourceClaim Manifests from Model Fits
llmfit-core/src/claim.rs transforms a model fit—comprising model weights, quantization level, and target throughput—into a Kubernetes Device Resource Allocation (DRA) ResourceClaim or ResourceClaimTemplate by calculating hardware bounds, generating DNS-safe identifiers, and serializing guarded CEL selectors as YAML or JSON.
The AlexsJones/llmfit repository automates the pairing of Large Language Model (LLM) inference workloads with Kubernetes hardware resources. Central to this capability is the rendering pipeline in llmfit-core/src/claim.rs, which translates abstract performance requirements into concrete Kubernetes manifests that DRA controllers consume to bind Pods to specific GPU or accelerator devices.
Overview of the Rendering Pipeline
The conversion process follows three logical stages defined in claim.rs: resolving hardware requirements, establishing resource identity, and emitting the final manifest format.
Resolving Fit Bounds with fit_bounds()
The fit_bounds() function (lines 60-94) computes the minimum viable hardware configuration to satisfy the requested tokens-per-second (min_tps). This calculation begins by selecting the effective quantization—either from the ClaimTarget override or the model’s default—and estimating the weight size via model.estimate_disk_gb(&quant).
When the requested quantization differs from the database entry, the function applies a WEIGHTS_HEADROOM multiplier of 1.2 to account for runtime memory overhead (lines 78-88). It then derives the minimum bandwidth using the relationship tok/s ≈ bandwidth × efficiency / weights, calculated as:
min_bandwidth_gbs = min_tps * weights_gb * 100 / efficiency_pct
The function returns a FitBounds struct containing memory_gi, min_bandwidth_gbs, weights_gb, and the selected quantization for downstream rendering.
Building Deterministic Claim Names
To ensure reproducible GitOps workflows, claim_name() (lines 99-115) generates a DNS-label-safe identifier derived from the model name, such as llama-2-7b-chat-fit. While callers may override this via ClaimTarget::name, the automatic generation guarantees uniqueness while respecting Kubernetes naming constraints of 63 characters and valid charset restrictions.
YAML and JSON Manifest Generation
The final stage produces consumable output through two primary entry points:
render()(lines 48-122): Constructs human-readable YAML representing either aResourceClaimorResourceClaimTemplate(whentemplate = true). The implementation builds a CEL selector that guards every optional device attribute lookup, ensuring missing capabilities result in a scheduling non-match rather than a runtime error.render_json()(lines 17-41): Serializes the same fit bounds usingserde_json::json!for programmatic consumption by the llmfit-dra controller, rounding weights to one decimal place for precision.
CEL Selector Construction for Safe Device Matching
The generated CEL expression represents the core matching logic that binds the claim to suitable hardware. Rather than directly accessing device attributes, the selector uses existence guards to prevent evaluation errors on heterogeneous clusters. The expression validates memory capacity, bandwidth availability, and device health status:
'memory' in device.capacity['llmfit.ai'] &&
device.capacity['llmfit.ai'].memory.compareTo(quantity('6Gi')) >= 0 &&
'memoryBandwidthGBs' in device.attributes['llmfit.ai'] &&
device.attributes['llmfit.ai'].memoryBandwidthGBs >= 152 &&
'healthy' in device.attributes['llmfit.ai'] &&
device.attributes['llmfit.ai'].healthy
This defensive pattern ensures the Kubernetes DRA ResourceClaim remains valid even when cluster nodes lack the extended llmfit.ai device class attributes.
Practical Code Examples
The following Rust code demonstrates generating both YAML and JSON manifests from a model fit:
use llmfit_core::claim::{render, render_json, ClaimTarget};
use llmfit_core::models::LlmModel;
// Load a model from the embedded catalog
let model: LlmModel = /* ... */;
// Configure target throughput and efficiency
let mut target = ClaimTarget::default();
target.min_tps = 30.0;
target.efficiency_pct = 60;
// Generate YAML for GitOps workflows
let yaml = render(&model, &target).expect("failed to render claim");
println!("{}", yaml);
The resulting YAML includes a provenance header commenting the model parameters and calculated bounds:
# Generated by llmfit claim — do not compute these constants by hand.
# model: Llama‑2‑7B‑Chat (7.0B params, Q4_K_M ≈ 3.9 GB weights)
# fit: tok/s ≈ bandwidth × 60% / 3.9 GB ⇒ bandwidth ≥ 152 GB/s for ≥ 30 tok/s
# memory: ≥ 6 Gi (weights + KV/runtime headroom)
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: llama-2-7b-chat-fit
spec:
devices:
requests:
- name: model
exactly:
deviceClassName: llmfit.ai
selectors:
- cel:
expression: >-
'memory' in device.capacity['llmfit.ai'] &&
device.capacity['llmfit.ai'].memory.compareTo(quantity('6Gi')) >= 0 &&
'memoryBandwidthGBs' in device.attributes['llmfit.ai'] &&
device.attributes['llmfit.ai'].memoryBandwidthGBs >= 152 &&
'healthy' in device.attributes['llmfit.ai'] &&
device.attributes['llmfit.ai'].healthy
For controller integration, use JSON output:
let json = render_json(&model, &target, env!("CARGO_PKG_VERSION"))
.expect("failed to render JSON");
println!("{}", json);
This produces a structured representation suitable for automated pipelines:
{
"model": "Llama-2-7B-Chat",
"claimName": "llama-2-7b-chat-fit",
"quant": "Q4_K_M",
"weightsGb": 3.9,
"memoryGi": 6,
"minBandwidthGBs": 152,
"minTps": 30.0,
"efficiencyPct": 60,
"deviceClass": "llmfit.ai",
"resolverVersion": "0.3.1"
}
Summary
fit_bounds()inllmfit-core/src/claim.rscalculates minimum memory and bandwidth requirements using the formulamin_tps * weights_gb * 100 / efficiency_pct, applying a 1.2x headroom multiplier for non-default quantization selections.- Claim identifiers are auto-generated from model names but can be overridden via
ClaimTarget::nameto support specific organizational naming conventions. - CEL selectors employ guarded existence checks (
'attribute' in device...) to safely match devices without causing scheduler runtime errors on missing attributes. render()produces human-readable YAML with embedded provenance comments, whilerender_json()generates machine-readable output for the llmfit-dra controller.- The pipeline supports both standard
ResourceClaimandResourceClaimTemplateKubernetes kinds via thetemplateboolean flag inClaimTarget.
Frequently Asked Questions
What constitutes a ModelFit in llmfit?
A ModelFit represents the combination of a specific LLM architecture (including parameter count), a chosen quantization level (such as Q4_K_M), and a desired throughput measured in tokens-per-second. According to the implementation in llmfit-core/src/claim.rs, the fit captures the hardware requirements necessary to achieve that performance level, which the rendering pipeline then converts into concrete Kubernetes resource requests.
How does the CEL selector prevent runtime errors when device attributes are missing?
The selector uses existence guards like 'memory' in device.capacity['llmfit.ai'] before attempting to evaluate attribute values. As implemented in lines 61-68 of claim.rs, this pattern ensures that if a node lacks the llmfit.ai device class or specific extended attributes, the expression evaluates to false (resulting in a non-match) rather than throwing a runtime exception that could crash the Kubernetes scheduler.
When should I use render_json() instead of render()?
Use render_json() when integrating with automation pipelines, CI/CD systems, or the llmfit-dra controller, as it outputs structured data without YAML comments or template wrappers. Use render() when generating GitOps manifests or human-readable configuration files intended for direct application via kubectl, as it produces standard Kubernetes YAML with embedded comments explaining the model parameters and calculated hardware bounds.
How does overriding quantization affect ResourceClaim memory requirements?
When ClaimTarget.quant differs from the model's default quantization, fit_bounds() applies a WEIGHTS_HEADROOM constant of 1.2 to the weight size calculation (lines 78-88 of claim.rs). This multiplier accounts for increased memory fragmentation and KV-cache overhead that occurs when running non-optimized quantization formats, resulting in a higher memory_gi value in the final ResourceClaim to ensure safe allocation and prevent out-of-memory errors during inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →