analyze_with_forced_runtime vs analyze_with_config in llmfit: Runtime Control vs Configuration Tuning

analyze_with_forced_runtime overrides automatic runtime detection to force a specific inference engine, directly changing which run-mode (GPU, CPU, etc.) is evaluated, while analyze_with_config applies custom calculation parameters through CalcConfig without altering the underlying runtime selection.

The ModelFit struct in the llmfit library provides two distinct entry points in llmfit-core/src/fit.rs for analyzing model fitness. While both methods ultimately delegate to the private analyze_inner function, they serve fundamentally different purposes in the fitting pipeline, particularly regarding how the run-mode is determined and how scoring calculations are performed.

Core Functional Differences

Runtime Override with analyze_with_forced_runtime

The analyze_with_forced_runtime method allows you to bypass the automatic runtime detection logic and specify a particular InferenceRuntime variant. According to the source at lines 398-405 in llmfit-core/src/fit.rs, this method accepts an optional force_runtime parameter that, when provided, sets the runtime first-hand inside analyze_inner before any evaluation occurs.

This forced selection directly influences the subsequent run-mode decision—whether the model fits in RunMode::Gpu, RunMode::CpuOnly, or RunMode::TensorParallel. For example, forcing InferenceRuntime::LlamaCpp on an Apple Silicon machine will evaluate the GPU path using LlamaCpp quantization rules instead of the default MLX path.

Calculation Tuning with analyze_with_config

In contrast, analyze_with_config lets you supply a custom CalcConfig structure to adjust calculation parameters without changing the runtime. As implemented at lines 107-115 in fit.rs, this method accepts a mandatory CalcConfig that can specify context_cap, tps_efficiency, and run-mode weighting.

When using this method, the runtime selection follows the default detection path because force_runtime is always None. The configuration modifies how memory requirements are calculated and how the final score is computed, but the underlying runtime variable remains automatically selected based on system specs in llmfit-core/src/hardware.rs.

How Runtime Detection Shapes Run-Mode Selection

Inside the private analyze_inner function at lines 97-107 of llmfit-core/src/fit.rs, the runtime selection follows a strict cascade defined in llmfit-core/src/providers.rs when force_runtime is None:

  • Cluster environments trigger InferenceRuntime::Vllm
  • Pre-quantized models select InferenceRuntime::Vllm
  • Apple Silicon systems default to InferenceRuntime::Mlx
  • Fallback routes to InferenceRuntime::LlamaCpp

When analyze_with_forced_runtime provides Some(runtime), this cascade is bypassed, and the forced value is used immediately. The selected runtime—whether forced or auto-detected—determines which run-mode evaluation path executes (e.g., GPU memory calculations for Mlx differ from LlamaCpp quantization rules). This directly influences whether the output contains RunMode::Gpu, RunMode::CpuOnly, or RunMode::TensorParallel.

Method Signatures and Source Locations

The public API differs significantly in their parameters:

// Lines 398-405 in llmfit-core/src/fit.rs
pub fn analyze_with_forced_runtime(
    model: &LlmModel, 
    system: &SystemSpecs, 
    context_limit: Option<u32>, 
    force_runtime: Option<InferenceRuntime>
) -> Self
// Lines 107-115 in llmfit-core/src/fit.rs  
pub fn analyze_with_config(
    model: &LlmModel, 
    system: &SystemSpecs, 
    config: CalcConfig
) -> Self

Note that analyze_with_forced_runtime mirrors the standard analyze signature but adds the force_runtime parameter, while analyze_with_config replaces the optional parameters with a mandatory CalcConfig object defined in llmfit-core/src/calc.rs.

Practical Usage Examples

Use analyze_with_forced_runtime when you need to benchmark or deploy with a specific inference engine regardless of automatic detection:

use llmfit_core::fit::ModelFit;
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;
use llmfit_core::providers::InferenceRuntime;

// Force LlamaCpp on Apple Silicon to compare quantization approaches
let forced_fit = ModelFit::analyze_with_forced_runtime(
    &my_model,
    &my_system,
    None,                              // No explicit context limit
    Some(InferenceRuntime::LlamaCpp), // Override automatic MLX selection
);

Use analyze_with_config when you need to adjust calculation assumptions for your specific deployment constraints:

use llmfit_core::calc::CalcConfig;

let mut cfg = CalcConfig::default();
cfg.context_cap = Some(8_000);      // Cap estimation at 8k tokens
cfg.tps_efficiency = 0.85;          // Adjust for slower storage subsystem

let config_fit = ModelFit::analyze_with_config(
    &my_model, 
    &my_system, 
    cfg
);

Summary

  • analyze_with_forced_runtime controls which inference engine evaluates the model fit, directly affecting whether ModelFit.runtime is set to LlamaCpp, Mlx, or Vllm, and consequently which run-mode (GPU, CPU, TensorParallel) is selected.
  • analyze_with_config adjusts how the fit is calculated through CalcConfig parameters like context_cap and tps_efficiency, potentially changing memory_required_gb and scoring components while keeping the runtime selection automatic.
  • Both methods ultimately call analyze_inner in llmfit-core/src/fit.rs, but they populate different internal flags that alter the analysis pipeline at distinct stages.

Frequently Asked Questions

Can I force a runtime and apply custom configuration simultaneously?

No, the current API in llmfit-core/src/fit.rs provides separate entry points that cannot be combined in a single call. analyze_with_forced_runtime does not accept a CalcConfig parameter, and analyze_with_config always passes None for the forced runtime. To achieve both effects, you would need to modify the private analyze_inner function directly or fork the validation logic.

How does forcing InferenceRuntime::LlamaCpp change the run-mode on Apple Silicon?

When you force LlamaCpp on Apple Silicon, the code skips the automatic Mlx selection at lines 97-107 of fit.rs. This causes the run-mode evaluation to use LlamaCpp's quantization rules and memory estimation instead of MLX's, potentially resulting in RunMode::Gpu with different memory requirements than the native MLX path would calculate.

Does the context_cap in CalcConfig affect which run-mode is selected?

Indirectly, yes. While analyze_with_config does not change the runtime variable, the context_cap reduces the estimated KV-cache size in the memory calculation. This can shift the boundary between Gpu and CpuOffload modes if the reduced memory footprint allows the model to fit in GPU memory where the full context would not.

Why would I use analyze_with_forced_runtime instead of the standard analyze method?

Use the forced variant when automatic detection selects a suboptimal runtime for your specific hardware or when benchmarking. For example, forcing Vllm for multi-node cluster tests or LlamaCpp to compare quantization performance against the default MLX implementation on Mac systems allows explicit control over the inference engine without changing system detection logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →