How llmfit-core `plan.rs` Estimates Memory, Throughput, and Hardware Upgrade Needs
The plan.rs module in llmfit-core analyzes a PlanRequest to calculate VRAM and RAM requirements, predict tokens-per-second (TPS) across GPU and CPU execution paths, and generate specific hardware upgrade deltas by comparing current system specs against minimum and recommended thresholds.
The llmfit repository by AlexsJones provides a Rust-based LLM operations toolkit where the planning subsystem helps users determine if their hardware can run specific models. By processing a PlanRequest containing context length, quantization preferences, and target throughput, the system produces a PlanEstimate that maps resource requirements to actionable upgrade recommendations.
Architecture of the Planning Workflow
The core planning logic resides in llmfit-core/src/plan.rs, specifically within the estimate_model_plan_with_config function (around lines 670-794). This orchestrates a nine-step pipeline that transforms high-level request parameters into concrete hardware assessments:
- Request validation – Rejects zero-context or non-positive TPS values (lines 638-645).
- Quantization resolution –
normalize_quantmaps user strings to canonical forms, defaulting to the model's native quantization if none specified (lines 36-77). - KV-cache quantization – Defaults to
Fp16when not explicitly provided (line 53). - Path estimation – Builds GPU, CPU-offload, and CPU-only execution profiles via
build_path_estimate(lines 885-1015). - Current system evaluation –
evaluate_currentassesses real-world performance on the existing machine (lines 631-702). - Path selection – Prioritizes feasible GPU paths, then CPU-offload, then CPU-only (lines 713-726).
- Upgrade delta calculation – Computes VRAM, RAM, and CPU core gaps (lines 738-795).
- KV-cache alternatives –
compute_kv_alternativesenumerates memory impacts of different cache quantization options (lines 1120-1180). - Result consolidation – Returns the populated
PlanEstimatestruct.
Memory Estimation Logic
Memory calculations rely on the model's estimate_memory_gb_with_kv method (invoked at line 885), which aggregates three components: weight size based on the selected quantization, KV-cache size determined by the cache quantization level, and a small operational overhead.
For GPU execution paths, build_path_estimate sets the minimum VRAM requirement to exactly the model_mem value (line 951). The recommended VRAM adds a 20% headroom buffer to accommodate runtime fluctuations (line 952).
CPU-offload and CPU-only paths shift these requirements to system RAM rather than VRAM (lines 1043-1055). The planner evaluates these against available system memory using fit_level_for to determine if the configuration is feasible, tight, or insufficient (lines 1000-1023).
Throughput (TPS) Prediction
The system estimates tokens-per-second through estimate_tps_with_gpu (lines 79-101) or the fallback estimate_tps for CPU-only scenarios. The algorithm follows a hierarchical approach:
- Known GPU bandwidth – When
resolve_gpu_bandwidthidentifies the specific GPU, the function delegates tofit::estimate_tps(lines 124-131), ensuring consistency with the broaderllmfit-corefitting heuristics. - Backend K-factors – For unknown hardware, it applies a constant performance factor per backend (lines 158-166) modified by
quant_speed_multiplier(lines 170-176). - CPU scaling – CPU-only paths use CPU-specific constants (lines 178-185) and apply core-count multipliers (lines 190-194).
- Run-mode efficiency – Applies user-tuned efficiency values from
CalcConfig(lines 191-193) to alignplanandfitsubcommand outputs.
For each path, minimum_cores_for_target (lines 889-904) iterates to find the smallest core count that satisfies the optional target_tps parameter, ensuring the recommendation meets user performance expectations.
Hardware Upgrade Delta Calculation
The upgrade_deltas struct (defined at lines 84-92) quantifies the gap between current hardware and execution requirements through three resource-specific calculations:
- VRAM upgrades – Calculated as
max(min_vram - current_vram, 0)for "Good" fit (lines 743-749) andmax(rec_vram - current_vram, 0)for "Perfect" fit (lines 754-760). - RAM upgrades – Triggered when
minimum.ram_gbexceeds available system RAM (lines 770-777). - CPU core requirements – Suggested when
minimum.cpu_coressurpassessystem.total_cpu_cores(lines 783-794).
These deltas feed directly into CLI and TUI outputs, displaying precise values like "+2.0 GB VRAM → Good" to guide hardware procurement decisions.
Code Examples
Rust API Integration
use llmfit_core::{
plan::{estimate_model_plan, PlanRequest},
hardware::SystemSpecs,
};
fn main() -> Result<(), String> {
// Detect current hardware configuration
let system = SystemSpecs::detect().map_err(|e| e.to_string())?;
// Define planning constraints: 8192-token context targeting 10 TPS
let request = PlanRequest {
context: 8192,
quant: None, // Use model default
target_tps: Some(10.0),
kv_quant: None, // Defaults to Fp16
};
// Load model from internal database
let model = llmfit_core::models::ModelDatabase::new()
.map_err(|e| e.to_string())?
.find_by_name("Qwen-7B")
.ok_or("model not found")?;
// Generate estimate
let plan = estimate_model_plan(&model, &request, &system)?;
let best_path = plan.preferred_path.unwrap();
println!("Recommended execution: {}", best_path.path.label());
println!("Estimated TPS: {:.1}", best_path.estimated_tps);
Ok(())
}
Command-Line Interface
# Plan for Qwen-7B with 4096-context and 8 TPS target
cargo run -- plan "Qwen-7B" --context 4096 --target-tps 8
# Output includes:
# • Minimum/recommended VRAM and RAM requirements
# • TPS estimates for GPU, CPU-offload, and CPU-only paths
# • Upgrade suggestions (e.g., "+4.2 GB VRAM → Good")
Summary
- Memory calculation combines model weights, KV-cache size, and overhead in
estimate_memory_gb_with_kv, with GPU paths requiring exact VRAM matches and recommended 20% headroom. - Throughput prediction leverages
estimate_tps_with_gpuwith GPU bandwidth detection, backend-specific K-factors, and quantization multipliers to project tokens-per-second across execution paths. - Upgrade recommendations derive from the
upgrade_deltasstruct, calculating specific VRAM, RAM, and CPU core deficits against currentSystemSpecs. - Path flexibility evaluates GPU, CPU-offload, and CPU-only configurations, automatically selecting the most performant feasible option while documenting alternatives.
Frequently Asked Questions
How does plan.rs determine if my GPU has sufficient VRAM?
The system calls fit_level_for (lines 1000-1023) to compare the model's memory requirements against detected VRAM in SystemSpecs. For GPU paths, "Good" fit requires at least the minimum VRAM (exact model size), while "Perfect" fit requires the recommended VRAM (120% of model size). If current VRAM exceeds both thresholds, the path is marked feasible with room to spare.
What happens if I specify a target TPS that my hardware cannot achieve?
minimum_cores_for_target (lines 889-904) attempts to find a core count and execution path that meets your target. If no configuration satisfies the requirement, the planner still returns estimates for all paths but marks them as sub-target, allowing you to see the maximum achievable TPS and the specific hardware upgrades (additional VRAM or CPU cores) needed to reach your goal.
How does KV-cache quantization affect the memory estimate?
When compute_kv_alternatives (lines 1120-1180) processes the request, it evaluates each KvQuant variant (FP16, Q8, Q4, etc.) against the model's kv_cache_gb calculation. Lower-precision KV-cache formats reduce memory footprint proportionally, and the planner displays these as alternative scenarios showing exact memory savings and any backend-specific compatibility notes.
Can the planner detect my GPU automatically?
Yes, through SystemSpecs::detect() in llmfit-core/src/hardware.rs, the system identifies installed GPUs and queries resolve_gpu_bandwidth to retrieve known memory bandwidth values. If your GPU is not in the internal database, the planner falls back to backend-specific K-factors (lines 158-166), which provide conservative throughput estimates based on quantization level rather than hardware-specific optimization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →