ModelFit Analysis Methods in llmfit: 5 Ways to Evaluate LLM Hardware Fit

The ModelFit struct in AlexsJones/llmfit provides five analysis methods—analyze, analyze_with_context_limit, analyze_with_forced_runtime, analyze_with_config, and refresh_estimate_confidence—that evaluate how well an LLM fits your hardware by calculating memory requirements, selecting optimal runtimes, and scoring compatibility levels.

The llmfit crate determines whether Large Language Models can run efficiently on specific hardware through its core evaluation engine. Located in llmfit-core/src/fit.rs, the ModelFit implementation combines system detection, quantization strategies, and runtime selection to generate detailed compatibility reports. Mastering these ModelFit analysis methods enables precise control over model deployment across diverse environments from consumer GPUs to CPU-only servers.

The Five Public Analysis Entry Points

The ModelFit API exposes five distinct entry points that funnel into a shared internal pipeline. Each method offers different levels of control over the evaluation process.

analyze: Default Hardware Evaluation

The analyze method serves as the primary entry point for standard evaluations. It automatically detects the optimal runtime, quantization level, and execution path without manual configuration.

According to lines 81-85 in llmfit-core/src/fit.rs, this method delegates to analyze_with_context_limit with context_limit set to None, allowing the engine to maximize the available context window based on detected hardware capabilities.

use llmfit_core::{ModelFit, SystemSpecs, LlmModel};

let system = SystemSpecs::detect()?;
let model = LlmModel::load("meta-llama/Meta-Llama-3-8B")?;
let fit = ModelFit::analyze(&model, &system);
println!("Fit score: {:.1}", fit.score);

analyze_with_context_limit: Constrained Context Windows

Use this method when you need to enforce a maximum token limit regardless of hardware capacity. This proves essential for memory-constrained deployments or specific workload requirements.

As implemented in lines 86-90, this function passes your explicit context_limit through to analyze_inner, overriding the automatic context detection while preserving runtime selection logic.

// Force analysis with maximum 4096 tokens
let fit = ModelFit::analyze_with_context_limit(&model, &system, Some(4096));
println!("Usable context: {} tokens", fit.usable_context);

analyze_with_forced_runtime: Backend-Specific Benchmarking

This method overrides the auto-detected inference runtime, enabling benchmarking of specific backends like LlamaCpp versus MLX on Apple Silicon. Lines 94-100 show this function accepting an Option<InferenceRuntime> that bypasses the default platform detection.

use llmfit_core::types::InferenceRuntime;

// Force LlamaCpp runtime instead of default MLX on Metal
let fit = ModelFit::analyze_with_forced_runtime(
    &model,
    &system,
    None,
    Some(InferenceRuntime::LlamaCpp),
);

analyze_with_config: Custom Calculation Parameters

For fine-grained performance tuning, analyze_with_config accepts a CalcConfig struct that modifies efficiency factors, scoring weights, and run-mode multipliers. This addresses specific calibration needs referenced in issue #449 when default token-per-second estimates require adjustment.

The implementation (lines 108-115) passes your custom configuration directly into analyze_inner, affecting how the engine calculates memory estimates and fit scores.

use llmfit_core::config::{CalcConfig, ScoringWeights};

let mut cfg = CalcConfig::default();
cfg.efficiency = 0.65; // More optimistic bandwidth assumption

let fit = ModelFit::analyze_with_config(&model, &system, cfg);

refresh_estimate_confidence: Recalibrating with Measured Data

After obtaining actual benchmark results, call this method to recompute the confidence label based on measured throughput rather than estimates. Lines 118-119 invoke effective_estimate_confidence to update the FitLevel classification.

let mut fit = ModelFit::analyze(&model, &system);
fit.measured_tps = Some(actual_benchmark_result);
fit.refresh_estimate_confidence();
println!("Updated confidence: {}", fit.estimate_confidence.label());

The Internal Analysis Pipeline

All public methods converge on the private analyze_inner function (lines 124-129). This unified workflow executes nine distinct calculation phases:

  1. Context window handling (lines 130-142) — Respects explicit limits and applies CalcConfig::context_cap constraints
  2. Memory requirement calculation (lines 144-148) — Invokes model.estimate_memory_gb with selected quantization
  3. Runtime selection (lines 150-168) — Evaluates force_runtime, cluster mode, pre-quantized models, and Apple Silicon (Metal) before falling back to LlamaCpp
  4. Execution-path decision (lines 170-240) — Selects from five RunMode variants: Gpu, MoeOffload, CpuOffload, CpuOnly, or TensorParallel
  5. Fit scoring (lines 267-270) — Calls score_fit to map memory ratios to FitLevel classifications (Perfect, Good, Marginal, TooTight)
  6. Quantization hierarchy selection (lines 310-339) — Chooses optimal quantization from QUANT_HIERARCHY, MLX_QUANT_HIERARCHY, or ONNX_QUANT_HIERARCHY
  7. Throughput estimation (lines 341-368) — Calculates tokens-per-second via estimate_tps and records inputs in EstimateBasis
  8. Weighted scoring (lines 374-380) — Combines ScoreComponents using configurable weighted_score logic
  9. Usable context derivation (lines 442-468) — Computes actual token capacity within available memory pools

Key Supporting Functions

The analysis pipeline relies on specialized helper functions defined in llmfit-core/src/fit.rs:

  • score_fit (lines 336-343) — Converts pure memory ratios into capped FitLevel determinations
  • pure_ratio_verdict (lines 405-416) — Maps numerical memory ratios to categorical fit assessments
  • cap_for_run_mode (lines 418-426) — Downgrades Perfect scores to Good for non-GPU execution paths
  • best_quant_for_runtime_budget (lines 718-754) — Selects the smallest viable quantization fitting within memory constraints

Summary

  • ModelFit::analyze provides fully automatic evaluation using detected hardware capabilities
  • analyze_with_context_limit enforces specific token limits while maintaining automatic runtime selection
  • analyze_with_forced_runtime enables benchmarking of specific inference backends like LlamaCpp or MLX
  • analyze_with_config exposes calculation parameters for environments requiring custom efficiency factors or scoring weights
  • refresh_estimate_confidence updates fit assessments when real benchmark data becomes available
  • All methods share the analyze_inner pipeline in llmfit-core/src/fit.rs, which coordinates memory estimation, quantization selection, and run-mode decisions

Frequently Asked Questions

What is the difference between analyze and analyze_with_context_limit?

The analyze method maximizes the context window based on available hardware memory, while analyze_with_context_limit enforces a specific token ceiling you provide. Use the latter when deploying models in memory-constrained environments or when your application requires fixed-size context windows regardless of hardware capacity.

When should I use analyze_with_forced_runtime?

Use this method when benchmarking specific backends or when deployment constraints require a particular runtime. For example, you might force InferenceRuntime::LlamaCpp on Apple Silicon to compare performance against the default MLX backend, or lock to TensorParallel for multi-GPU cluster deployments.

How does ModelFit choose between different quantization levels?

The engine selects quantization using hierarchy-specific rules defined in best_quant_for_runtime_budget (lines 718-754). It evaluates QUANT_HIERARCHY for standard paths, MLX_QUANT_HIERARCHY for Apple Silicon, or ONNX_QUANT_HIERARCHY for ONNX runtimes, choosing the smallest quantization that satisfies the memory budget while maximizing model quality.

Can I update a ModelFit analysis after getting real benchmark data?

Yes. After obtaining measured throughput data from llmfit-core/src/bench.rs, attach the MeasuredTps to your ModelFit instance and call refresh_estimate_confidence. This recalculates the confidence label and fit level using actual performance metrics rather than theoretical estimates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →