# llmfit ModelFit Analysis Methods: Context Limits, Runtime Overrides, and Custom Config

> Explore llmfit ModelFit analysis methods. Understand context limits, runtime overrides, and custom config for your LLM analysis with AlexsJones/llmfit. Optimize your models effectively.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: api-reference
- Published: 2026-09-11

---

**All three entry points are thin wrappers around the internal `analyze_inner` function that respectively control the maximum token context window, force a specific inference backend, or provide full calculation customization via a configuration struct.**

The `llmfit` library from the AlexsJones/llmfit repository provides a full-stack "fit" analysis to determine how well a large language model runs on your hardware. While the base `ModelFit::analyze` method uses sensible defaults, three specialized APIs give you granular control over specific analysis parameters without modifying the core calculation engine.

## `analyze_with_context_limit`: Token Window Constraints

Located in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) at lines 85‑90, `analyze_with_context_limit` accepts an optional `context_limit: Option<u32>` parameter that caps the KV-cache memory estimate.

When you provide a token count, the function forwards this limit to `analyze_inner`, where it is applied **before** the model's native `context_length` and the library's `DEFAULT_ESTIMATION_CTX`. This ensures the memory calculation never exceeds your specified token budget, which is useful when you know your inference workloads will use shorter prompts than the model's theoretical maximum.

```rust
use llmfit_core::{ModelFit, LlmModel, SystemSpecs};

// Limit analysis to 4096 tokens regardless of model capabilities
let fit = ModelFit::analyze_with_context_limit(&model, &system, Some(4096));

```

## `analyze_with_forced_runtime`: Backend Override

The `analyze_with_forced_runtime` function (lines 98‑105 in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)) bypasses the automatic runtime-selection logic by accepting an optional `InferenceRuntime` enum variant.

When `force_runtime` is `Some`, the internal logic skips the default decision tree (cluster mode → VLLM, Metal + unified memory → MLX, etc.) and uses your specified backend instead. This is particularly useful on Apple Silicon where you might want to force `LlamaCpp` rather than accept the automatic selection, though pre-quantized models still default to `vllm` regardless of the override.

```rust
use llmfit_core::InferenceRuntime;

// Force LlamaCpp runtime while keeping default context handling
let fit = ModelFit::analyze_with_forced_runtime(
    &model,
    &system,
    None,  // no context limit override
    Some(InferenceRuntime::LlamaCpp),
);

```

## `analyze_with_config`: Full Parameter Control

Defined at lines 108‑115 of [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), `analyze_with_config` accepts a complete `CalcConfig` struct, enabling you to override TPS efficiency factors, run-mode multipliers, scoring weights, and the context cap simultaneously.

The function extracts any `context_cap` defined in the config and forwards the entire configuration object to `analyze_inner`. This is the most flexible entry point, allowing you to tune the fit calculation's internal knobs without forking the library.

```rust
use llmfit_core::CalcConfig;

let mut cfg = CalcConfig::default();
cfg.context_cap = Some(8000);      // cap context at 8000 tokens
cfg.tps_efficiency = 0.85;          // adjust tokens-per-second efficiency

let fit = ModelFit::analyze_with_config(&model, &system, cfg);

```

## The Common Core: `analyze_inner`

All three public methods delegate to `analyze_inner(model, system, context_limit, force_runtime, config)` defined in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs).

Inside this private function (lines 31‑40 for context handling, lines 95‑107 for runtime selection), the analysis proceeds through five stages:

1. **Context resolution** – Derives the effective estimation context from the supplied limit, model's `context_length`, and `DEFAULT_ESTIMATION_CTX`
2. **Runtime selection** – Uses `force_runtime` if provided, otherwise applies the automatic backend selection logic
3. **Quantization fitting** – Evaluates which quantization fits the selected runtime's memory budget using the effective context size
4. **Execution path selection** – Chooses `RunMode` (e.g., `TensorParallel`, `Gpu`, `CpuOffload`) based on system topology
5. **Scoring** – Calculates fit level, throughput estimates, and final scores using any custom weights from `CalcConfig`

## Summary

- **`analyze_with_context_limit`** is **context-centric**: It only affects the token-window estimate passed to the memory calculator, making it ideal for constraining KV-cache projections to known workload sizes.
- **`analyze_with_forced_runtime`** is **runtime-centric**: It forces a specific inference backend (like `LlamaCpp` or `MLX`) while leaving all other parameters at their defaults, useful for testing specific execution paths.
- **`analyze_with_config`** is **full-config**: It accepts a `CalcConfig` struct that can simultaneously override context caps, TPS efficiency factors, and scoring weights for completely custom fit calculations.

## Frequently Asked Questions

### When should I use `analyze_with_context_limit` versus `analyze_with_config`?

Use `analyze_with_context_limit` when you only need to cap the token count for memory estimation, as it provides a simpler API for that single purpose. Use `analyze_with_config` when you need to adjust multiple parameters like TPS efficiency or scoring weights in addition to the context limit, since it accepts the full `CalcConfig` struct.

### Does `analyze_with_forced_runtime` affect quantization selection?

Yes, indirectly. While the parameter forces a specific backend runtime, the internal `analyze_inner` function still runs quantization fitting logic against that runtime's memory budget. However, pre-quantized models bypass some of this logic and default to `vllm` regardless of your forced runtime choice.

### What happens if I pass `None` to `analyze_with_context_limit`?

Passing `None` is equivalent to calling the base `ModelFit::analyze` method, which internally invokes `analyze_with_context_limit(..., None)`. This tells the analyzer to use the model's advertised `context_length` capped by the library's `DEFAULT_ESTIMATION_CTX` constant without any user-specified override.

### Can I combine runtime forcing with context limits?

Not in a single method call. The API design provides separate entry points for clarity. To apply both constraints, use `analyze_with_config`, which allows you to set `context_cap` in the `CalcConfig` while also specifying a forced runtime through `config.force_runtime` if you extend the configuration, or manually chain the logic by calling the methods sequentially and comparing results.