How the llmfit Analysis Pipeline Workflow Ranks LLMs for Your Hardware
The llmfit analysis pipeline workflow transforms raw hardware specifications into a ranked list of compatible language models by detecting system capabilities, indexing installed providers, filtering the model catalog, and scoring candidates through memory analysis, runtime prediction, and benchmark-calibrated throughput estimation.
The llmfit analysis pipeline workflow powers the AlexsJones/llmfit repository, enabling developers to determine which large language models will run efficiently on their specific hardware configuration. Written in Rust, this pipeline bridges the gap between static model metadata and dynamic hardware constraints to deliver precise fit recommendations for CPUs, GPUs, and unified memory systems.
Stage-by-Stage Pipeline Execution
The workflow consists of nine distinct stages, each implemented in specific modules of the llmfit-core crate.
1. Hardware Detection with SystemSpecs
The pipeline begins by calling SystemSpecs::detect() in llmfit-core/src/hardware.rs (lines 72–155). This function probes the host machine to gather total RAM, CPU core count, and GPU specifications including VRAM capacity, device type, and unified memory flags. It interrogates multiple vendor tools—nvidia-smi, rocm-smi, lspci, Windows WMI, and macOS system-profiler—to construct a comprehensive hardware profile.
2. Indexing Installed Models
Next, InstalledIndex::detect_all() spawns a dedicated thread for each supported provider (Ollama, MLX, llama.cpp, Docker-MR, LM Studio, vLLM, and RamaLama) as implemented in llmfit-core/src/analysis.rs (lines 59–92). This concurrent scan aggregates locally available model identifiers, creating an index used later to mark catalog entries as already present on disk.
3. Loading the Model Catalog
The ModelDatabase::new() constructor, located in llmfit-core/src/models.rs, loads the embedded hf_models.json catalog containing Hugging Face and ONNX model metadata. It merges these entries with any custom or cached specifications, exposing parameter counts, quantization options, minimum memory requirements, and use-case tags.
4. Filtering Rank-able Models
The rankable_models() function in llmfit-core/src/analysis.rs (lines 74–99) filters the master catalog to exclude models incompatible with the current backend or "sanitization-demoted" entries featuring draft heads and size mismatches. The result is a curated iterator containing only hardware-compatible candidates.
5. Computing Model Fits
The core scoring logic resides in build_model_fits() (lines 41–62 of llmfit-core/src/analysis.rs). For each candidate model, the pipeline invokes ModelFit::analyze_with_forced_runtime() or ModelFit::analyze_with_config() from llmfit-core/src/fit.rs to calculate:
- Required memory (RAM and VRAM)
- Inference runtime (GPU-accelerated, CPU-offload, or mixed)
- Estimated throughput (tokens per second)
- Fit level classification (Perfect, Good, Marginal, or TooTight)
6. Benchmark Enrichment
For each ModelFit object, the pipeline queries measured throughput data from three hierarchical sources in llmfit-core/src/analysis.rs (lines 98–115): user-local benchmarks, community-submitted llmfit results, and pre-computed measured-TPS indexes. When real measurements exist, they override formulaic estimates and update the confidence flag.
7. Local Calibration
The apply_local_calibration() function (lines 26–39 of llmfit-core/src/analysis.rs) computes a hardware-specific scaling factor from the ratio of measured versus estimated throughput on dense, large models. This factor, clamped to the range [0.05, 3.0], scales all estimated TPS values to improve prediction accuracy for the specific host machine.
8. Marking Installed Status
The pipeline consults the InstalledIndex to set the installed boolean flag on each ModelFit object (lines 108–112 of llmfit-core/src/analysis.rs). This allows downstream consumers to highlight models already available for immediate inference.
9. Returning Results
Finally, build_model_fits() returns a Vec<ModelFit> (lines 41–45 of llmfit-core/src/analysis.rs) containing unsorted fit objects. The calling interface—whether CLI, TUI, HTTP API, or MCP—receives this vector and handles domain-specific sorting, filtering, and presentation.
Core Implementation Architecture
The pipeline relies on five primary source files in the llmfit-core crate:
llmfit-core/src/hardware.rs— ImplementsSystemSpecs::detect()for cross-platform hardware probing.llmfit-core/src/analysis.rs— Contains the orchestration logic includingrankable_models(),build_model_fits(), and calibration routines.llmfit-core/src/models.rs— DefinesModelDatabaseand catalog parsing forhf_models.json.llmfit-core/src/providers.rs— Defines provider backends and theInstalledIndexaggregation logic.llmfit-core/src/fit.rs— ImplementsModelFitanalysis methods for runtime selection and throughput estimation.
Entry points in llmfit-tui/src/main.rs and llmfit-tui/src/serve_api.rs wire these core functions into command-line and HTTP interfaces respectively.
Programmatic Usage Examples
Running a Full Analysis
To execute the complete pipeline from a Rust binary:
use llmfit_core::{
analysis::{build_model_fits, rankable_models},
hardware::SystemSpecs,
models::ModelDatabase,
providers::InstalledIndex,
};
fn main() {
// Detect hardware
let specs = SystemSpecs::detect();
// Find installed models
let installed = InstalledIndex::detect_all();
// Load the model catalog
let db = ModelDatabase::new();
// Build fit results without forced runtime or context cap
let fits = build_model_fits(&db, &specs, &installed, None, None);
// Sort by fit level then estimated throughput
let mut sorted = fits;
sorted.sort_by(|a, b| a.compare(b));
// Display top results
for fit in sorted.iter().take(10) {
println!(
"{} – {:?} – {:.2} tps – installed: {}",
fit.model.name,
fit.run_mode,
fit.estimated_tps,
fit.installed
);
}
}
API Integration
For HTTP API handlers using Axum:
let specs = SystemSpecs::detect();
let installed = InstalledIndex::detect_all();
let db = ModelDatabase::new();
let fits = build_model_fits(&db, &specs, &installed, Some(8192), None);
serde_json::to_string(&fits).unwrap()
Custom Hardware Profiles
Apply bandwidth and efficiency overrides using CalcConfig:
use llmfit_core::fit::CalcConfig;
let custom_cfg = CalcConfig {
// Configure custom bandwidth, efficiency metrics, etc.
..Default::default()
};
let fits = build_model_fits_with_config(&db, &specs, &installed, None, custom_cfg);
Summary
- The llmfit analysis pipeline workflow executes nine sequential stages from hardware detection to result generation.
- SystemSpecs::detect() in
hardware.rsprovides the foundation by probing RAM, CPU, and GPU resources across multiple vendor tools. - InstalledIndex::detect_all() concurrently scans all supported providers to identify locally available models.
- The ModelDatabase loads embedded and custom catalogs, while rankable_models() filters for hardware compatibility.
- build_model_fits() orchestrates scoring through
ModelFitanalysis methods, enriched with real benchmark data and local calibration factors. - Results are returned as a vector of ModelFit objects containing fit levels, runtime modes, and throughput estimates ready for consumer presentation.
Frequently Asked Questions
How does llmfit detect GPU specifications across different vendors?
The pipeline invokes SystemSpecs::detect() in llmfit-core/src/hardware.rs (lines 72–155), which executes platform-specific discovery commands including nvidia-smi for NVIDIA GPUs, rocm-smi for AMD cards, and lspci for generic PCI devices on Linux. On Windows, it queries WMI classes, while macOS uses system-profiler to detect Metal-capable GPUs and unified memory configurations.
What distinguishes estimated from measured throughput in the analysis?
Estimated throughput derives from analytical formulas based on parameter count, quantization, and hardware bandwidth, while measured throughput originates from actual benchmark runs stored in local indexes, community databases, or pre-computed TPS tables. When measured data exists in llmfit-core/src/analysis.rs (lines 98–115), it overrides the estimate and updates the result's confidence flag to indicate empirical validation.
Can the pipeline run if some LLM providers are not installed?
Yes. InstalledIndex::detect_all() in llmfit-core/src/analysis.rs (lines 59–92) spawns independent threads for each provider. Missing providers simply return empty result sets without failing the overall analysis. The pipeline functions correctly with zero providers detected, though no models will carry the installed flag.
How does local calibration improve throughput predictions?
apply_local_calibration() in llmfit-core/src/analysis.rs (lines 26–39) calculates a scaling factor by comparing measured versus estimated throughput on large, dense models present in your benchmark history. This factor, constrained between 0.05 and 3.0, adjusts all subsequent estimates to account for your specific hardware's actual inference efficiency, correcting for thermal throttling, driver variations, or memory bandwidth limitations not captured by generic formulas.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →