What Is the llmfit Plan Command? Complete Hardware Feasibility Guide
The llmfit plan command is a high-level estimator that generates a detailed feasibility report showing how a specific LLM model will run on your current hardware and exactly what upgrades are needed to improve performance.
The plan command in the AlexsJones/llmfit repository transforms raw hardware specs and model requirements into an actionable deployment roadmap. Unlike simple compatibility checks, it produces a PlanEstimate structure that quantifies throughput in tokens-per-second (TPS), identifies bottlenecks, and calculates precise resource deltas needed to reach better fit levels. This functionality is implemented in llmfit-core/src/plan.rs and employs the same speed heuristics as the fit command to ensure consistent performance predictions across the tool.
Core Components of the PlanEstimate
At the heart of the plan command lies the PlanEstimate data structure, built by the estimate_model_plan function in llmfit-core/src/plan.rs. This estimate contains several critical analysis layers:
- Model Configuration: The selected model name, provider, context size, and quantization scheme.
- Hardware Specifications: Both minimum requirements (VRAM, RAM, CPU cores) to run the model at all, and recommended specifications that provide comfortable headroom for stable operation.
- Run-Path Feasibility: Per-path analysis for GPU-only, CPU-offload, and CPU-only execution modes, each annotated with estimated TPS and a fit level.
- Fit Level Classification: A graded assessment ranging from Perfect → Good → Marginal → TooTight, indicating how well the current hardware matches the model's demands.
- Upgrade Deltas: Specific resource increments (e.g., "+30.0 GB VRAM") required to advance to the next fit tier, eliminating guesswork when budgeting hardware upgrades.
- KV-Cache Alternatives: "What-if" scenarios showing memory usage and quality impact of different KV-cache quantization strategies.
Shared Speed Heuristics and Configuration
The plan command maintains consistency with llmfit's broader analysis engine by delegating TPS calculations to the same underlying estimators used by the fit command. Specifically, the estimate_tps_with_gpu function in llmfit-core/src/plan.rs calls crate::fit::estimate_tps to compute throughput predictions.
This shared logic means that any configuration tweaks—such as efficiency adjustments or DDR bandwidth overrides specified via estimate_model_plan_with_config—are honored by both commands. Users see identical TPS projections whether they are validating current hardware with fit or exploring future scenarios with plan.
Hardware Profiles and "What-If" Modeling
The plan command supports hardware profiles that allow you to simulate deployment on machines other than your current one. When a profile is supplied via the --profile flag, the estimator uses the profile's GPU bandwidth specifications instead of the detected machine's bandwidth.
This capability is verified in llmfit-tui/tests/hardware_profiles_cli.rs by the test plan_estimates_from_the_profile_bandwidth_not_the_default_config, which confirms that users can accurately model performance impacts of faster memory or different GPU architectures before purchasing new hardware.
Practical Usage Examples
Generate a JSON feasibility report for a 120B parameter model with 8K context:
llmfit plan openai/gpt-oss-120b --context 8192 --json
Evaluate the same model with specific quantization and a target throughput threshold:
llmfit plan openai/gpt-oss-120b --context 8192 \
--quant Q4_K_M --target-tps 20 --json
Simulate performance on faster hardware using a predefined profile:
llmfit --profile test-unified-fast plan openai/gpt-oss-120b \
--context 8192 --json
Understanding the PlanEstimate Output
The JSON output from the plan command provides structured data for automation and decision-making. A typical response includes:
{
"model_name":"gpt-oss-120b",
"context":8192,
"quantization":"Q4_K_M",
"minimum":{"vram_gb":30.0,"ram_gb":45.0,"cpu_cores":8},
"recommended":{"vram_gb":60.0,"ram_gb":90.0,"cpu_cores":16},
"run_paths":[
{"path":"GPU","feasible":true,"estimated_tps":12.3,"fit_level":"Good"},
{"path":"CPU offload","feasible":true,"estimated_tps":5.1,"fit_level":"Marginal"},
{"path":"CPU-only","feasible":true,"estimated_tps":2.7,"fit_level":"Marginal"}
],
"upgrade_deltas":[
{"resource":"vram_gb","add_gb":30.0,"target_fit":"Good","description":"+30.0 GB VRAM -> Good"}
]
}
Run paths detail three execution strategies: full GPU acceleration, partial CPU offloading for memory-constrained systems, and CPU-only inference for compatibility testing. Each path reports feasibility, estimated TPS, and fit level. Upgrade deltas provide concrete procurement guidance, specifying exactly how much additional VRAM, RAM, or CPU cores are necessary to achieve a "Good" or "Perfect" fit level.
TUI Integration and User Interface
In interactive mode, the plan command leverages components from llmfit-tui/src/tui_app.rs to store user-entered parameters including context window size, quantization settings, KV-cache quantization preferences, and target TPS values. The rendering logic in llmfit-tui/src/tui_ui.rs then visualizes the PlanEstimate data, presenting hardware minima, recommendations, and upgrade suggestions in a navigable interface.
Summary
- The
llmfit plancommand generates a PlanEstimate inllmfit-core/src/plan.rsthat predicts LLM performance on specific hardware. - It calculates minimum and recommended specs, grades hardware fit as Perfect/Good/Marginal/TooTight, and provides exact upgrade deltas.
- Throughput estimates use shared heuristics from
llmfit-core/src/fit.rsviaestimate_tps, ensuring consistency with thefitcommand. - Hardware profiles enable "what-if" analysis using alternate GPU bandwidth specifications.
- Output includes per-path feasibility for GPU, CPU-offload, and CPU-only execution modes with associated TPS benchmarks.
Frequently Asked Questions
What does the llmfit plan command do?
The plan command analyzes whether a specific LLM model can run on given hardware and produces a detailed feasibility report. It estimates throughput in tokens-per-second, identifies whether the current configuration is sufficient, and calculates exact hardware upgrades needed to improve performance.
How does plan differ from fit in llmfit?
While both commands use the same speed estimation logic from crate::fit::estimate_tps, fit evaluates how well your current hardware runs a model, whereas plan acts as a forward-looking estimator that can simulate different hardware profiles and provides structured upgrade recommendations.
What hardware specifications does the plan command analyze?
The command evaluates VRAM requirements for GPU execution, RAM requirements for CPU-offload scenarios, and CPU core counts for pure CPU inference. It generates minimum thresholds for basic operation and recommended specs for optimal performance.
Can I simulate different hardware with the plan command?
Yes. By specifying a hardware profile with the --profile flag, you can override the detected system's GPU bandwidth and other characteristics. This allows you to model performance on faster or different hardware before making purchasing decisions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →