What Does a "TooTight" Fit Level Signify in llmfit?

A "TooTight" fit level in llmfit indicates that a language model does not fit into available GPU VRAM or system RAM, making it impossible to load or run on the current hardware.

In the llmfit repository, fit levels are a core abstraction that describe how comfortably a language model can execute on detected hardware. The levels form a severity-ordered enum: Perfect → Good → Marginal → TooTight. According to the source code in llmfit-core/src/fit.rs, the TooTight variant explicitly means "Does not fit in available memory"【https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L182】.

How llmfit Determines a "TooTight" Classification

The classification occurs during hardware-model analysis in the core planning module.

In llmfit-core/src/plan.rs, the planning logic compares a model's memory requirements against detected GPU VRAM or system RAM. When available memory is insufficient, the function returns FitLevel::TooTight. This prevents downstream components from attempting to load models that would immediately fail with out-of-memory errors.

The enum definition in llmfit-core/src/fit.rs at line 182 establishes this as the most severe fit level, reserved exclusively for memory-exhaustion scenarios rather than performance degradation.

Where "TooTight" Appears in the Codebase

CLI Filtering with --tight and --runnable

The command-line interface exposes explicit controls for handling "TooTight" models.

In llmfit-tui/src/main.rs at line 1152, the CLI parses flags into a FitArg enum. The --tight flag filters results to only TooTight models, while --runnable explicitly excludes them. This dual-mode design supports both troubleshooting ("what's blocking me?") and operational workflows ("what can I actually run?").


# Display only models that exceed current memory capacity

cargo run -- fit --tight

# Display only models that can actually execute

cargo run -- fit --runnable

Backend Filtering in Rust

Applications built on llmfit-core can filter programmatically using the FitLevel enum:

use llmfit_core::fit::FitLevel;
use llmfit_core::analysis::ModelFit;

// `fits` is a Vec<ModelFit> from build_model_fits()
let too_tight_fits: Vec<&ModelFit> = fits
    .iter()
    .filter(|fit| fit.fit_level == FitLevel::TooTight)
    .collect();

// runnable_fits excludes memory-exhausted models
let runnable_fits: Vec<&ModelFit> = fits
    .iter()
    .filter(|fit| fit.fit_level != FitLevel::TooTight)
    .collect();

TUI Visual Indicators

The terminal UI renders "TooTight" entries with distinctive red coloring at llmfit-tui/src/display.rs line 255 and llmfit-tui/src/tui_ui.rs lines 696-705. This includes a red circle icon and badge styling that makes memory-blocked models immediately obvious to users scanning results.

Resolving "TooTight" Conditions

When llmfit reports TooTight, three resolution paths exist:

  • Quantization: Reduce model precision (e.g., FP16 → INT8) to shrink memory footprint
  • Memory liberation: Close other GPU/CPU processes to free VRAM or system RAM
  • Hardware upgrade: Deploy on systems with larger memory capacity

The "TooTight" designation prevents wasted time attempting to load models that cannot physically fit, distinguishing memory exhaustion from performance concerns handled by Marginal or Good levels.

Summary

Frequently Asked Questions

How does "TooTight" differ from "Marginal" in llmfit?

"Marginal" indicates a model fits in memory but runs with severely degraded performance—uncomfortably slow or near resource limits. "TooTight" means the model cannot load at all due to memory exhaustion. The distinction prevents users from attempting runs guaranteed to fail versus runs that merely perform poorly.

Can a "TooTight" model ever run successfully?

No. By definition in llmfit-core/src/fit.rs, "TooTight" models exceed available memory. The only resolution paths are reducing model size through quantization, freeing memory from other processes, or upgrading hardware. The designation is binary: either sufficient memory exists, or it does not.

Why would I use the --tight CLI flag?

The --tight flag serves diagnostic and planning purposes. It reveals which models your hardware cannot run, helping identify upgrade requirements or quantization targets. For operational deployments, --runnable is the inverse—filtering to models that will actually execute.

Does "TooTight" apply to both GPU and CPU execution?

Yes. The FitLevel::TooTight check in llmfit-core/src/fit.rs applies to VRAM for GPU-accelerated inference and system RAM for CPU-only execution. The same enum variant covers both execution paths, unified under the "insufficient memory" semantics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →