How llmfit's `best_quant_for_budget` Function Works: Automatic Quantization Selection with Context Window Adjustments

The best_quant_for_budget function automatically selects the highest-quality quantization level that fits within a specified memory budget by iterating through quantization hierarchies and automatically halving the context window when necessary to meet hardware constraints.

The best_quant_for_budget method in the llmfit repository serves as the core decision engine that determines which model quantization can run within available RAM or VRAM limits. Implemented in llmfit-core/src/models.rs (lines 555-586), this function balances model quality against memory constraints by testing quantization options from highest to lowest fidelity, employing an intelligent fallback mechanism that reduces context length to preserve quantization quality when possible.

Core Implementation in llmfit-core

The primary logic resides within the LlmModel struct in llmfit-core/src/models.rs, exposing two related public methods that handle quantization selection with varying degrees of customization.

Public API and Wrapper Methods

The entry point best_quant_for_budget(budget_gb, ctx) serves as a convenience wrapper that delegates immediately to best_quant_for_budget_with using the default QUANT_HIERARCHY constant (lines 557-559). This design allows standard usage with minimal parameters while maintaining flexibility for advanced scenarios.

For deployments requiring runtime-specific quantization preferences—such as MLX-optimized hierarchies for Apple Silicon—the best_quant_for_budget_with(budget_gb, ctx, hierarchy) method accepts a custom ordered slice of quantization names. This enables fine-tuned optimization while preserving the same budget-checking logic and fallback behavior.

The Quantization Hierarchy

The hierarchy parameter represents an ordered list of quantization identifiers (e.g., ["Q8_0", "Q6_K", "Q5_K_M", "Q4_K_M", "Q2_K"]) ranked from best to worst quality. The function traverses this sequence sequentially, testing each option against the memory budget until identifying a suitable candidate that fits within the specified budget_gb limit.

Two-Pass Selection Algorithm

The function employs a deterministic two-pass strategy that prioritizes maintaining the full context window before resorting to context reduction as a fallback.

First Pass: Full Context Evaluation

During the initial iteration, the method scans the quantization hierarchy from highest to lowest quality. For each quantization level q, it invokes self.estimate_memory_gb(q, ctx) to calculate the precise memory footprint required for that quantization paired with the specified context length. If the estimated memory is less than or equal to budget_gb, the function immediately returns Some((q, mem)), pairing the quantization name with its memory requirement.

This approach ensures that when hardware permits, the model runs with both optimal quantization quality and maximum context length, preserving inference capabilities.

Second Pass: Context Window Reduction

If no quantization level fits within the budget at the full context length, the function automatically triggers a fallback mechanism. It calculates half_ctx = ctx / 2 and verifies that this value meets the minimum threshold of 1024 tokens. If the halved context remains viable, the function re-scans the entire quantization hierarchy using half_ctx instead of the original context length.

The first quantization that satisfies the budget under the reduced context is returned, allowing models to execute on constrained hardware by trading context length for quantization quality rather than immediately defaulting to the lowest-quality quantization format.

Handling Unsatisfiable Constraints

When neither the full nor halved context yields a suitable quantization, the function returns None, explicitly signaling that the model cannot execute within the specified memory constraints. This failure mode enables higher-level planning logic in llmfit-core/src/plan.rs to make informed decisions about resource allocation or model selection.

Integration with Runtime Budget Helpers

The best_quant_for_runtime_budget function in llmfit-core/src/fit.rs (lines 22-46) serves as the higher-level orchestrator that bridges runtime detection with model-specific selection logic. This helper determines the appropriate quantization hierarchy based on the execution target—whether MLX, ONNX, or generic inference engines—and retrieves the estimation context (often the model's default context length) before invoking best_quant_for_budget_with.

This architecture separates runtime-specific configuration from the core budget mathematics, enabling llmfit to support multiple inference backends while maintaining consistent memory budgeting semantics. The selected quantization and context parameters then flow into llmfit-core/src/plan.rs, which constructs the concrete deployment plan and may apply additional context trimming if necessary.

Code Example: Selecting Quantization for a Memory Budget

The following Rust example demonstrates how to use the budget selection functions with both default and custom hierarchies:

use llmfit_core::models::{LlmModel, QUANT_HIERARCHY, MLX_QUANT_HIERARCHY};

// Assume `model` is an initialized LlmModel instance
let budget_gb = 5.0;          // 5 GB RAM/VRAM budget
let ctx = 4096;               // Desired context length in tokens

// Method 1: Use the default quantization hierarchy
if let Some((quant, mem)) = model.best_quant_for_budget(budget_gb, ctx) {
    println!("Best quant: {quant}, estimated memory: {mem:.2} GB");
} else {
    println!("No quantization fits within the {budget_gb} GB budget");
}

// Method 2: Supply a custom MLX-specific hierarchy
if let Some((quant, mem)) = model.best_quant_for_budget_with(
    budget_gb, 
    ctx, 
    &MLX_QUANT_HIERARCHY
) {
    println!("MLX-optimal quant: {quant} ({mem:.2} GB)");
}

For a typical 7-billion-parameter model with a 5 GB budget and 4096-token context, this code outputs:

Best quant: Q4_K_M, estimated memory: 4.23 GB

If the budget were reduced to 2 GB, the function automatically halves the context to 2048 tokens and potentially returns a lower-quality quantization such as Q2_K to meet the constraint.

Summary

  • best_quant_for_budget in llmfit-core/src/models.rs implements the core logic for selecting quantization levels based on memory constraints, iterating from highest to lowest quality.
  • The function uses a two-pass algorithm that first attempts to fit the model using the full context window, then automatically falls back to halving the context (minimum 1024 tokens) if necessary.
  • Custom hierarchies are supported via best_quant_for_budget_with, enabling runtime-specific quantization preferences such as MLX_QUANT_HIERARCHY for Apple Silicon deployments.
  • Runtime integration occurs through best_quant_for_runtime_budget in llmfit-core/src/fit.rs, which selects appropriate hierarchies based on the inference backend before delegating to the core selection method.
  • The function returns None when no quantization configuration fits the budget, providing clear feedback for resource planning and deployment decisions.

Frequently Asked Questions

What happens if no quantization fits the budget even with context reduction?

If neither the full context nor the halved context (down to a minimum of 1024 tokens) yields a quantization level that fits within the specified budget, best_quant_for_budget returns None. This signals to the calling code in llmfit-core/src/plan.rs that the model cannot run on the available hardware, allowing the system to either select a different model or report the resource constraint to the user.

How does the context window adjustment work?

The context window adjustment serves as an automatic fallback mechanism when memory constraints prevent using the desired context length. When no quantization fits at the original ctx size, the function calculates half_ctx = ctx / 2 and re-evaluates the entire quantization hierarchy using this reduced context. This allows the system to prioritize maintaining higher quantization quality at the expense of context length, ensuring the best possible model performance within strict memory limits. The adjustment occurs only once per function call; further reductions are handled by the planning logic in plan.rs.

Can I use a custom quantization hierarchy?

Yes. While best_quant_for_budget uses the default QUANT_HIERARCHY, the best_quant_for_budget_with method accepts any slice of quantization strings as a hierarchy parameter. This enables runtime-specific optimizations, such as passing MLX_QUANT_HIERARCHY for Apple Silicon deployments or custom hierarchies prioritizing specific quantization formats for ONNX Runtime, without modifying the core budget-checking algorithm.

Where does the memory estimation come from?

The memory estimation is provided by the estimate_memory_gb method on the LlmModel struct, which calculates the exact RAM or VRAM required for a specific quantization level and context length. This calculation considers the model's parameter count, the bits per weight for the specified quantization format, and the KV cache requirements determined by the context window size. Higher-level functions like best_quant_for_runtime_budget determine which context length to use for these estimations based on runtime configuration and model defaults.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →