llmfit ModelFit Analysis Methods: Context Limits, Runtime Overrides, and Custom Config
All three entry points are thin wrappers around the internal analyze_inner function that respectively control the maximum token context window, force a specific inference backend, or provide full calculation customization via a configuration struct.
The llmfit library from the AlexsJones/llmfit repository provides a full-stack "fit" analysis to determine how well a large language model runs on your hardware. While the base ModelFit::analyze method uses sensible defaults, three specialized APIs give you granular control over specific analysis parameters without modifying the core calculation engine.
analyze_with_context_limit: Token Window Constraints
Located in llmfit-core/src/fit.rs at lines 85‑90, analyze_with_context_limit accepts an optional context_limit: Option<u32> parameter that caps the KV-cache memory estimate.
When you provide a token count, the function forwards this limit to analyze_inner, where it is applied before the model's native context_length and the library's DEFAULT_ESTIMATION_CTX. This ensures the memory calculation never exceeds your specified token budget, which is useful when you know your inference workloads will use shorter prompts than the model's theoretical maximum.
use llmfit_core::{ModelFit, LlmModel, SystemSpecs};
// Limit analysis to 4096 tokens regardless of model capabilities
let fit = ModelFit::analyze_with_context_limit(&model, &system, Some(4096));
analyze_with_forced_runtime: Backend Override
The analyze_with_forced_runtime function (lines 98‑105 in llmfit-core/src/fit.rs) bypasses the automatic runtime-selection logic by accepting an optional InferenceRuntime enum variant.
When force_runtime is Some, the internal logic skips the default decision tree (cluster mode → VLLM, Metal + unified memory → MLX, etc.) and uses your specified backend instead. This is particularly useful on Apple Silicon where you might want to force LlamaCpp rather than accept the automatic selection, though pre-quantized models still default to vllm regardless of the override.
use llmfit_core::InferenceRuntime;
// Force LlamaCpp runtime while keeping default context handling
let fit = ModelFit::analyze_with_forced_runtime(
&model,
&system,
None, // no context limit override
Some(InferenceRuntime::LlamaCpp),
);
analyze_with_config: Full Parameter Control
Defined at lines 108‑115 of llmfit-core/src/fit.rs, analyze_with_config accepts a complete CalcConfig struct, enabling you to override TPS efficiency factors, run-mode multipliers, scoring weights, and the context cap simultaneously.
The function extracts any context_cap defined in the config and forwards the entire configuration object to analyze_inner. This is the most flexible entry point, allowing you to tune the fit calculation's internal knobs without forking the library.
use llmfit_core::CalcConfig;
let mut cfg = CalcConfig::default();
cfg.context_cap = Some(8000); // cap context at 8000 tokens
cfg.tps_efficiency = 0.85; // adjust tokens-per-second efficiency
let fit = ModelFit::analyze_with_config(&model, &system, cfg);
The Common Core: analyze_inner
All three public methods delegate to analyze_inner(model, system, context_limit, force_runtime, config) defined in llmfit-core/src/fit.rs.
Inside this private function (lines 31‑40 for context handling, lines 95‑107 for runtime selection), the analysis proceeds through five stages:
- Context resolution – Derives the effective estimation context from the supplied limit, model's
context_length, andDEFAULT_ESTIMATION_CTX - Runtime selection – Uses
force_runtimeif provided, otherwise applies the automatic backend selection logic - Quantization fitting – Evaluates which quantization fits the selected runtime's memory budget using the effective context size
- Execution path selection – Chooses
RunMode(e.g.,TensorParallel,Gpu,CpuOffload) based on system topology - Scoring – Calculates fit level, throughput estimates, and final scores using any custom weights from
CalcConfig
Summary
analyze_with_context_limitis context-centric: It only affects the token-window estimate passed to the memory calculator, making it ideal for constraining KV-cache projections to known workload sizes.analyze_with_forced_runtimeis runtime-centric: It forces a specific inference backend (likeLlamaCpporMLX) while leaving all other parameters at their defaults, useful for testing specific execution paths.analyze_with_configis full-config: It accepts aCalcConfigstruct that can simultaneously override context caps, TPS efficiency factors, and scoring weights for completely custom fit calculations.
Frequently Asked Questions
When should I use analyze_with_context_limit versus analyze_with_config?
Use analyze_with_context_limit when you only need to cap the token count for memory estimation, as it provides a simpler API for that single purpose. Use analyze_with_config when you need to adjust multiple parameters like TPS efficiency or scoring weights in addition to the context limit, since it accepts the full CalcConfig struct.
Does analyze_with_forced_runtime affect quantization selection?
Yes, indirectly. While the parameter forces a specific backend runtime, the internal analyze_inner function still runs quantization fitting logic against that runtime's memory budget. However, pre-quantized models bypass some of this logic and default to vllm regardless of your forced runtime choice.
What happens if I pass None to analyze_with_context_limit?
Passing None is equivalent to calling the base ModelFit::analyze method, which internally invokes analyze_with_context_limit(..., None). This tells the analyzer to use the model's advertised context_length capped by the library's DEFAULT_ESTIMATION_CTX constant without any user-specified override.
Can I combine runtime forcing with context limits?
Not in a single method call. The API design provides separate entry points for clarity. To apply both constraints, use analyze_with_config, which allows you to set context_cap in the CalcConfig while also specifying a forced runtime through config.force_runtime if you extend the configuration, or manually chain the logic by calling the methods sequentially and comparing results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →