# analyze_with_forced_runtime vs analyze_with_config in llmfit: Runtime Control vs Configuration Tuning

> Understand analyze_with_forced_runtime vs analyze_with_config in llmfit. Force inference engines or tune calculation parameters to control LLM fitting efficiently. Choose the right run mode for your needs.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-09-12

---

**`analyze_with_forced_runtime` overrides automatic runtime detection to force a specific inference engine, directly changing which run-mode (GPU, CPU, etc.) is evaluated, while `analyze_with_config` applies custom calculation parameters through `CalcConfig` without altering the underlying runtime selection.**

The `ModelFit` struct in the `llmfit` library provides two distinct entry points in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) for analyzing model fitness. While both methods ultimately delegate to the private `analyze_inner` function, they serve fundamentally different purposes in the fitting pipeline, particularly regarding how the **run-mode** is determined and how scoring calculations are performed.

## Core Functional Differences

### Runtime Override with analyze_with_forced_runtime

The `analyze_with_forced_runtime` method allows you to **bypass the automatic runtime detection logic** and specify a particular `InferenceRuntime` variant. According to the source at lines 398-405 in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), this method accepts an optional `force_runtime` parameter that, when provided, sets the runtime first-hand inside `analyze_inner` before any evaluation occurs.

This forced selection directly influences the subsequent **run-mode decision**—whether the model fits in `RunMode::Gpu`, `RunMode::CpuOnly`, or `RunMode::TensorParallel`. For example, forcing `InferenceRuntime::LlamaCpp` on an Apple Silicon machine will evaluate the GPU path using LlamaCpp quantization rules instead of the default MLX path.

### Calculation Tuning with analyze_with_config

In contrast, `analyze_with_config` lets you supply a custom **`CalcConfig`** structure to adjust calculation parameters without changing the runtime. As implemented at lines 107-115 in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs), this method accepts a mandatory `CalcConfig` that can specify `context_cap`, `tps_efficiency`, and **run-mode weighting**.

When using this method, the runtime selection follows the default detection path because `force_runtime` is always `None`. The configuration modifies how **memory requirements** are calculated and how the final **score** is computed, but the underlying runtime variable remains automatically selected based on system specs in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs).

## How Runtime Detection Shapes Run-Mode Selection

Inside the private `analyze_inner` function at lines 97-107 of [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the runtime selection follows a strict cascade defined in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) when `force_runtime` is `None`:

- **Cluster environments** trigger `InferenceRuntime::Vllm`
- **Pre-quantized models** select `InferenceRuntime::Vllm`  
- **Apple Silicon systems** default to `InferenceRuntime::Mlx`
- **Fallback** routes to `InferenceRuntime::LlamaCpp`

When `analyze_with_forced_runtime` provides `Some(runtime)`, this cascade is bypassed, and the forced value is used immediately. The selected runtime—whether forced or auto-detected—determines which **run-mode** evaluation path executes (e.g., GPU memory calculations for `Mlx` differ from `LlamaCpp` quantization rules). This directly influences whether the output contains `RunMode::Gpu`, `RunMode::CpuOnly`, or `RunMode::TensorParallel`.

## Method Signatures and Source Locations

The public API differs significantly in their parameters:

```rust
// Lines 398-405 in llmfit-core/src/fit.rs
pub fn analyze_with_forced_runtime(
    model: &LlmModel, 
    system: &SystemSpecs, 
    context_limit: Option<u32>, 
    force_runtime: Option<InferenceRuntime>
) -> Self

```

```rust
// Lines 107-115 in llmfit-core/src/fit.rs  
pub fn analyze_with_config(
    model: &LlmModel, 
    system: &SystemSpecs, 
    config: CalcConfig
) -> Self

```

Note that `analyze_with_forced_runtime` mirrors the standard `analyze` signature but adds the `force_runtime` parameter, while `analyze_with_config` replaces the optional parameters with a mandatory `CalcConfig` object defined in [`llmfit-core/src/calc.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/calc.rs).

## Practical Usage Examples

Use `analyze_with_forced_runtime` when you need to benchmark or deploy with a specific inference engine regardless of automatic detection:

```rust
use llmfit_core::fit::ModelFit;
use llmfit_core::hardware::SystemSpecs;
use llmfit_core::models::LlmModel;
use llmfit_core::providers::InferenceRuntime;

// Force LlamaCpp on Apple Silicon to compare quantization approaches
let forced_fit = ModelFit::analyze_with_forced_runtime(
    &my_model,
    &my_system,
    None,                              // No explicit context limit
    Some(InferenceRuntime::LlamaCpp), // Override automatic MLX selection
);

```

Use `analyze_with_config` when you need to adjust calculation assumptions for your specific deployment constraints:

```rust
use llmfit_core::calc::CalcConfig;

let mut cfg = CalcConfig::default();
cfg.context_cap = Some(8_000);      // Cap estimation at 8k tokens
cfg.tps_efficiency = 0.85;          // Adjust for slower storage subsystem

let config_fit = ModelFit::analyze_with_config(
    &my_model, 
    &my_system, 
    cfg
);

```

## Summary

- **`analyze_with_forced_runtime`** controls *which inference engine* evaluates the model fit, directly affecting whether `ModelFit.runtime` is set to `LlamaCpp`, `Mlx`, or `Vllm`, and consequently which **run-mode** (GPU, CPU, TensorParallel) is selected.
- **`analyze_with_config`** adjusts *how the fit is calculated* through `CalcConfig` parameters like `context_cap` and `tps_efficiency`, potentially changing `memory_required_gb` and scoring components while keeping the runtime selection automatic.
- Both methods ultimately call `analyze_inner` in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), but they populate different internal flags that alter the analysis pipeline at distinct stages.

## Frequently Asked Questions

### Can I force a runtime and apply custom configuration simultaneously?

No, the current API in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) provides separate entry points that cannot be combined in a single call. `analyze_with_forced_runtime` does not accept a `CalcConfig` parameter, and `analyze_with_config` always passes `None` for the forced runtime. To achieve both effects, you would need to modify the private `analyze_inner` function directly or fork the validation logic.

### How does forcing `InferenceRuntime::LlamaCpp` change the run-mode on Apple Silicon?

When you force `LlamaCpp` on Apple Silicon, the code skips the automatic `Mlx` selection at lines 97-107 of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs). This causes the run-mode evaluation to use LlamaCpp's quantization rules and memory estimation instead of MLX's, potentially resulting in `RunMode::Gpu` with different memory requirements than the native MLX path would calculate.

### Does the `context_cap` in `CalcConfig` affect which run-mode is selected?

Indirectly, yes. While `analyze_with_config` does not change the runtime variable, the `context_cap` reduces the estimated KV-cache size in the memory calculation. This can shift the boundary between `Gpu` and `CpuOffload` modes if the reduced memory footprint allows the model to fit in GPU memory where the full context would not.

### Why would I use `analyze_with_forced_runtime` instead of the standard `analyze` method?

Use the forced variant when automatic detection selects a suboptimal runtime for your specific hardware or when benchmarking. For example, forcing `Vllm` for multi-node cluster tests or `LlamaCpp` to compare quantization performance against the default MLX implementation on Mac systems allows explicit control over the inference engine without changing system detection logic.