How llmfit-core `plan.rs` Estimates Memory, Throughput, and Hardware Upgrade Needs

The plan.rs module in llmfit-core analyzes a PlanRequest to calculate VRAM and RAM requirements, predict tokens-per-second (TPS) across GPU and CPU execution paths, and generate specific hardware upgrade deltas by comparing current system specs against minimum and recommended thresholds.

The llmfit repository by AlexsJones provides a Rust-based LLM operations toolkit where the planning subsystem helps users determine if their hardware can run specific models. By processing a PlanRequest containing context length, quantization preferences, and target throughput, the system produces a PlanEstimate that maps resource requirements to actionable upgrade recommendations.

Architecture of the Planning Workflow

The core planning logic resides in llmfit-core/src/plan.rs, specifically within the estimate_model_plan_with_config function (around lines 670-794). This orchestrates a nine-step pipeline that transforms high-level request parameters into concrete hardware assessments:

  1. Request validation – Rejects zero-context or non-positive TPS values (lines 638-645).
  2. Quantization resolution – normalize_quant maps user strings to canonical forms, defaulting to the model's native quantization if none specified (lines 36-77).
  3. KV-cache quantization – Defaults to Fp16 when not explicitly provided (line 53).
  4. Path estimation – Builds GPU, CPU-offload, and CPU-only execution profiles via build_path_estimate (lines 885-1015).
  5. Current system evaluation – evaluate_current assesses real-world performance on the existing machine (lines 631-702).
  6. Path selection – Prioritizes feasible GPU paths, then CPU-offload, then CPU-only (lines 713-726).
  7. Upgrade delta calculation – Computes VRAM, RAM, and CPU core gaps (lines 738-795).
  8. KV-cache alternatives – compute_kv_alternatives enumerates memory impacts of different cache quantization options (lines 1120-1180).
  9. Result consolidation – Returns the populated PlanEstimate struct.

Memory Estimation Logic

Memory calculations rely on the model's estimate_memory_gb_with_kv method (invoked at line 885), which aggregates three components: weight size based on the selected quantization, KV-cache size determined by the cache quantization level, and a small operational overhead.

For GPU execution paths, build_path_estimate sets the minimum VRAM requirement to exactly the model_mem value (line 951). The recommended VRAM adds a 20% headroom buffer to accommodate runtime fluctuations (line 952).

CPU-offload and CPU-only paths shift these requirements to system RAM rather than VRAM (lines 1043-1055). The planner evaluates these against available system memory using fit_level_for to determine if the configuration is feasible, tight, or insufficient (lines 1000-1023).

Throughput (TPS) Prediction

The system estimates tokens-per-second through estimate_tps_with_gpu (lines 79-101) or the fallback estimate_tps for CPU-only scenarios. The algorithm follows a hierarchical approach:

  • Known GPU bandwidth – When resolve_gpu_bandwidth identifies the specific GPU, the function delegates to fit::estimate_tps (lines 124-131), ensuring consistency with the broader llmfit-core fitting heuristics.
  • Backend K-factors – For unknown hardware, it applies a constant performance factor per backend (lines 158-166) modified by quant_speed_multiplier (lines 170-176).
  • CPU scaling – CPU-only paths use CPU-specific constants (lines 178-185) and apply core-count multipliers (lines 190-194).
  • Run-mode efficiency – Applies user-tuned efficiency values from CalcConfig (lines 191-193) to align plan and fit subcommand outputs.

For each path, minimum_cores_for_target (lines 889-904) iterates to find the smallest core count that satisfies the optional target_tps parameter, ensuring the recommendation meets user performance expectations.

Hardware Upgrade Delta Calculation

The upgrade_deltas struct (defined at lines 84-92) quantifies the gap between current hardware and execution requirements through three resource-specific calculations:

  • VRAM upgrades – Calculated as max(min_vram - current_vram, 0) for "Good" fit (lines 743-749) and max(rec_vram - current_vram, 0) for "Perfect" fit (lines 754-760).
  • RAM upgrades – Triggered when minimum.ram_gb exceeds available system RAM (lines 770-777).
  • CPU core requirements – Suggested when minimum.cpu_cores surpasses system.total_cpu_cores (lines 783-794).

These deltas feed directly into CLI and TUI outputs, displaying precise values like "+2.0 GB VRAM → Good" to guide hardware procurement decisions.

Code Examples

Rust API Integration

use llmfit_core::{
    plan::{estimate_model_plan, PlanRequest},
    hardware::SystemSpecs,
};

fn main() -> Result<(), String> {
    // Detect current hardware configuration
    let system = SystemSpecs::detect().map_err(|e| e.to_string())?;
    
    // Define planning constraints: 8192-token context targeting 10 TPS
    let request = PlanRequest {
        context: 8192,
        quant: None,           // Use model default
        target_tps: Some(10.0),
        kv_quant: None,        // Defaults to Fp16
    };
    
    // Load model from internal database
    let model = llmfit_core::models::ModelDatabase::new()
        .map_err(|e| e.to_string())?
        .find_by_name("Qwen-7B")
        .ok_or("model not found")?;
    
    // Generate estimate
    let plan = estimate_model_plan(&model, &request, &system)?;
    let best_path = plan.preferred_path.unwrap();
    
    println!("Recommended execution: {}", best_path.path.label());
    println!("Estimated TPS: {:.1}", best_path.estimated_tps);
    
    Ok(())
}

Command-Line Interface


# Plan for Qwen-7B with 4096-context and 8 TPS target

cargo run -- plan "Qwen-7B" --context 4096 --target-tps 8

# Output includes:

# • Minimum/recommended VRAM and RAM requirements

# • TPS estimates for GPU, CPU-offload, and CPU-only paths

# • Upgrade suggestions (e.g., "+4.2 GB VRAM → Good")

Summary

  • Memory calculation combines model weights, KV-cache size, and overhead in estimate_memory_gb_with_kv, with GPU paths requiring exact VRAM matches and recommended 20% headroom.
  • Throughput prediction leverages estimate_tps_with_gpu with GPU bandwidth detection, backend-specific K-factors, and quantization multipliers to project tokens-per-second across execution paths.
  • Upgrade recommendations derive from the upgrade_deltas struct, calculating specific VRAM, RAM, and CPU core deficits against current SystemSpecs.
  • Path flexibility evaluates GPU, CPU-offload, and CPU-only configurations, automatically selecting the most performant feasible option while documenting alternatives.

Frequently Asked Questions

How does plan.rs determine if my GPU has sufficient VRAM?

The system calls fit_level_for (lines 1000-1023) to compare the model's memory requirements against detected VRAM in SystemSpecs. For GPU paths, "Good" fit requires at least the minimum VRAM (exact model size), while "Perfect" fit requires the recommended VRAM (120% of model size). If current VRAM exceeds both thresholds, the path is marked feasible with room to spare.

What happens if I specify a target TPS that my hardware cannot achieve?

minimum_cores_for_target (lines 889-904) attempts to find a core count and execution path that meets your target. If no configuration satisfies the requirement, the planner still returns estimates for all paths but marks them as sub-target, allowing you to see the maximum achievable TPS and the specific hardware upgrades (additional VRAM or CPU cores) needed to reach your goal.

How does KV-cache quantization affect the memory estimate?

When compute_kv_alternatives (lines 1120-1180) processes the request, it evaluates each KvQuant variant (FP16, Q8, Q4, etc.) against the model's kv_cache_gb calculation. Lower-precision KV-cache formats reduce memory footprint proportionally, and the planner displays these as alternative scenarios showing exact memory savings and any backend-specific compatibility notes.

Can the planner detect my GPU automatically?

Yes, through SystemSpecs::detect() in llmfit-core/src/hardware.rs, the system identifies installed GPUs and queries resolve_gpu_bandwidth to retrieve known memory bandwidth values. If your GPU is not in the internal database, the planner falls back to backend-specific K-factors (lines 158-166), which provide conservative throughput estimates based on quantization level rather than hardware-specific optimization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →