# llmfit | Alex Jones | Knowledge Base | Instagit

Hundreds of models & providers. One command to find what runs on your hardware.

GitHub Stars: 30.5k

Repository: https://github.com/AlexsJones/llmfit

---

## Articles

### [What Is the llmfit Plan Command? Complete Hardware Feasibility Guide](/AlexsJones/llmfit/purpose-of-llmfit-plan-command)

Understand the llmfit plan command. This guide details how to estimate LLM hardware feasibility and identify necessary upgrades for optimal performance on your system.

- Tags: how-to-guide
- Published: 2026-09-13

### [How the llmfit Benchmark Command Works: Automated LLM Inference Testing](/AlexsJones/llmfit/how-llmfit-benchmark-command-works)

Learn how the llmfit bench command automates LLM inference testing. Measure throughput, detect servers, and submit results to the community repository.

- Tags: internals
- Published: 2026-09-13

### [Which LLM Runtimes Does llmfit Support? A Complete Guide to Local AI Backends](/AlexsJones/llmfit/llmfit-llm-runtime-integrations)

Discover which LLM runtimes llmfit supports including Ollama MLX llama cpp Docker LM Studio vLLM and RamaLama with its unified ModelProvider trait for seamless local AI backend integration.

- Tags: tutorial
- Published: 2026-09-13

### [How to Override Memory Bandwidth or Efficiency Settings in llmfit](/AlexsJones/llmfit/override-llmfit-memory-bandwidth-efficiency)

Override llmfit memory bandwidth and efficiency settings using hardware profiles, TUI, environment variables, or CLI flags for custom TPS estimates on your hardware.

- Tags: how-to-guide
- Published: 2026-09-13

### [How the llmfit Fit Algorithm Chooses a Runtime: Hardware-Aware Execution in Rust](/AlexsJones/llmfit/llmfit-fit-algorithm-runtime-choice)

Discover how the llmfit fit algorithm chooses a runtime. It ranks hardware against model needs, filters by backend, and picks the first GPU-first option for efficient execution.

- Tags: internals
- Published: 2026-09-13

### [min_vram_gb vs min_ram_gb in llmfit: GPU VRAM vs CPU RAM Requirements Explained](/AlexsJones/llmfit/llmfit-min_vram_gb-vs-min_ram_gb)

Understand min_vram_gb vs min_ram_gb in llmfit. Learn the difference between GPU VRAM and CPU RAM requirements for LLM inference based on model size. Optimize your setup now.

- Tags: deep-dive
- Published: 2026-09-13

### [How llmfit Handles Unified Memory on Apple Silicon for LLM Inference](/AlexsJones/llmfit/llmfit-unified-memory-apple-silicon)

Discover how llmfit leverages unified memory on Apple Silicon for faster LLM inference. Optimize performance by treating system RAM as a shared pool.

- Tags: internals
- Published: 2026-09-13

### [How llmfit Determines the FitLevel of a Model: Memory-Aware Analysis in Rust](/AlexsJones/llmfit/how-llmfit-determines-fitlevel)

Discover how llmfit determines model FitLevel by analyzing memory requirements against available hardware and runtime constraints. Learn more about this Rust library.

- Tags: deep-dive
- Published: 2026-09-13

### [RunMode Options in llmfit: 5 Execution Strategies for LLM Inference](/AlexsJones/llmfit/llmfit-runmode-options)

Explore llmfit RunMode options: Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly. Optimize LLM inference by distributing model weights across GPU VRAM, system RAM, and CPU.

- Tags: deep-dive
- Published: 2026-09-13

### [ModelFit Analysis Methods in llmfit: 5 Ways to Evaluate LLM Hardware Fit](/AlexsJones/llmfit/llmfit-model-fit-analysis-methods)

Explore the 5 analysis methods in llmfit ModelFit to evaluate LLM hardware fit. Calculate memory, optimize runtime, and score compatibility for your LLM.

- Tags: deep-dive
- Published: 2026-09-13

### [How llmfit Estimates LLM Throughput: Roofline Modeling and Efficiency Factors](/AlexsJones/llmfit/how-llmfit-estimates-llm-throughput)

Discover how llmfit estimates LLM throughput using roofline modeling and efficiency factors. Learn about TPS prediction, GPU memory bandwidth, and runtime mode adjustments for optimal performance.

- Tags: deep-dive
- Published: 2026-09-13

### [How the llmfit Analysis Pipeline Workflow Ranks LLMs for Your Hardware](/AlexsJones/llmfit/llmfit-analysis-pipeline-workflow)

Discover the llmfit analysis pipeline workflow. It ranks language models for your hardware by analyzing system capabilities, filtering models, and estimating performance. Get the best LLMs for your setup.

- Tags: internals
- Published: 2026-09-13

### [How to Add a New Model to llmfit's Catalog: A Step-by-Step Guide](/AlexsJones/llmfit/how-to-add-new-model-to-llmfit)

Learn how to add a new model to llmfit's catalog. Follow this step by step guide to update the Hugging Face model ID, run the scraper, and rebuild the Rust project easily.

- Tags: how-to-guide
- Published: 2026-09-13

### [What Is llmfit-core? The Rust Engine Behind Hardware-Aware LLM Recommendations](/AlexsJones/llmfit/purpose-of-llmfit-core-library)

Discover llmfit-core, the Rust engine that analyzes your hardware, assesses LLM compatibility, and provides fit scores to recommend the best models for your unique CPU, RAM, and GPU.

- Tags: internals
- Published: 2026-09-13

### [How llmfit Handles Custom User Models: A Complete Guide](/AlexsJones/llmfit/how-llmfit-handles-custom-user-models)

Discover how llmfit handles custom user models with a JSON overlay file. Extend the built-in model catalog easily without code changes. Get the complete guide.

- Tags: how-to-guide
- Published: 2026-09-13

### [What Are the Four Dimensions llmfit Uses to Score Models?](/AlexsJones/llmfit/llmfit-model-scoring-dimensions)

llmfit scores LLMs across four dimensions: Quality, Speed, Fit, and Context. Learn how these combine for a comprehensive model evaluation.

- Tags: deep-dive
- Published: 2026-09-13

### [How Does llmfit Analyze and Score LLM Models? A Deep Dive into the Core Pipeline](/AlexsJones/llmfit/how-llmfit-analyzes-llm-models)

llmfit analyzes and scores LLM models by evaluating hardware specs, generating memory-fit, throughput, and composite scores through its 13-step pipeline.

- Tags: deep-dive
- Published: 2026-09-13

### [Core Components of the llmfit Rust Workspace: A Deep Dive into the Cargo Architecture](/AlexsJones/llmfit/llmfit-rust-workspace-components)

Explore the core components of the llmfit Rust workspace. Understand the Cargo architecture including llmfit-core, llmfit-tui, and llmfit-desktop for your development.

- Tags: deep-dive
- Published: 2026-09-13

### [What Programming Language Is llmfit Written In?](/AlexsJones/llmfit/what-programming-language-is-llmfit-written-in)

Discover what programming language powers llmfit. This powerful tool is built exclusively in Rust, utilizing its 2024 edition for CLI, TUI, GUI, and web dashboard.

- Tags: getting-started
- Published: 2026-09-13

### [How llmfit Detects System CPU, RAM, and GPU Hardware: A Deep Dive into the Detection Logic](/AlexsJones/llmfit/how-does-llmfit-detect-system-hardware)

Discover how llmfit detects system CPU, RAM, and GPU hardware. Explore its layered detection logic using sysinfo and vendor-specific tools for comprehensive multi-platform support.

- Tags: deep-dive
- Published: 2026-09-13

### [How llmfit Serializes and Reloads TUI Filter State Across Sessions](/AlexsJones/llmfit/llmfit-tui-filter_config.rs-tui-filter-state-serialization-and-reloading)

Discover how llmfit serializes and reloads TUI filter state across sessions. Learn about JSON persistence, struct handling, and state translation for seamless user experience.

- Tags: internals
- Published: 2026-09-12

### [How llmfit-core Assembles a PR-Ready Payload from Locally Cached Benchmark Results](/AlexsJones/llmfit/llmfit-core-share.rs-pr-payload-assembly-from-cached-results)

Discover how llmfit-core assembles PR-ready payloads from cached benchmark results. Learn how it sanitizes paths, adds metadata, and creates pull requests.

- Tags: internals
- Published: 2026-09-12

### [How the Doctor Command Aggregates System Diagnostics in llmfit](/AlexsJones/llmfit/llmfit-core-doctor.rs-doctor-command-diagnostic-aggregation)

Learn how the llmfit doctor command aggregates system diagnostics in llmfit-core/src/doctor.rs. It compiles a markdown report with truncated output from external tools.

- Tags: internals
- Published: 2026-09-12

### [How llmfit Discovers Ollama, MLX, llama.cpp, Docker Model Runner, LM Studio, vLLM, and RamaLama via the `which` Crate](/AlexsJones/llmfit/llmfit-core-providers.rs-runtime-discovery-via-which)

Discover how llmfit uses the Rust which crate to find Ollama MLX llama.cpp Docker Model Runner LM Studio vLLM and RamaLama across Linux macOS and Windows by checking the system PATH efficiently.

- Tags: internals
- Published: 2026-09-12

### [Where Is TableState Mutably Borrowed in llmfit-tui? A Rust Borrow Checker Analysis](/AlexsJones/llmfit/llmfit-tui-tui_events.rs-table-state-mutable-borrowing)

Analyze where TableState is mutably borrowed in llmfit-tui. Discover how update_model_viewport manages &mut borrows in Rust's borrow checker for safe TUI event handling.

- Tags: deep-dive
- Published: 2026-09-12

### [How llmfit Enforces Stateless TUI Rendering While Handling Events in Rust](/AlexsJones/llmfit/llmfit-tui-tui_ui.rs-stateless-rendering-vs-tui_events.rs-app-mutation)

Discover how llmfit enforces stateless TUI rendering in Rust by separating UI drawing from event handling. Learn to manage state effectively for cleaner applications.

- Tags: internals
- Published: 2026-09-12

### [How llmfit Applies Bandwidth and Efficiency Overrides from Hardware Profiles to CalcConfig](/AlexsJones/llmfit/llmfit-core-hwprofile.rs-calcconfig-bandwidth-efficiency-overrides)

Learn how llmfit-core/src/hwprofile.rs applies bandwidth and efficiency overrides from hardware profiles to CalcConfig. Optimize your LLM inference performance.

- Tags: internals
- Published: 2026-09-12

### [How llmfit-core Loads Schema-v1 Hardware Profiles and Applies Capacity Overrides to SystemSpecs](/AlexsJones/llmfit/llmfit-core-hwprofile.rs-schema-v1-hardware-profile-loading-and-systemspecs-overrides)

Learn how llmfit-core loads schema-v1 hardware profiles and applies capacity overrides to SystemSpecs by merging bundled and user-supplied configurations.

- Tags: internals
- Published: 2026-09-12

### [Why llmfit Multiplies RAM by 1.2 and VRAM by 1.1 in scripts/scrape_hf_models.py](/AlexsJones/llmfit/scripts-scrape_hf_models.py-ram-vram-formula-multipliers)

Discover why llmfit uses 1.2x RAM and 1.1x VRAM multipliers in scrape_hf_models.py to accurately estimate resource needs for CPU and GPU inference.

- Tags: internals
- Published: 2026-09-12

### [How the llmfit Python Scraper Derives RAM and VRAM Requirements from Parameter Count](/AlexsJones/llmfit/scripts-scrape_hf_models.py-model-parameter-to-ram-vram-derivation)

Discover how the llmfit Python scraper calculates RAM and VRAM needs from model parameter counts. Learn about quantization formats, byte-per-parameter values, and overhead multipliers for accurate estimations.

- Tags: internals
- Published: 2026-09-12

### [Why llmfit MCP Server Results Are Converted Through serve_shared.rs](/AlexsJones/llmfit/llmfit-tui-mcp-server-result-conversion-serve_shared.rs)

Discover why llmfit MCP server results are converted through serve_shared.rs. Ensure consistent JSON serialization and eliminate code duplication across all interfaces.

- Tags: internals
- Published: 2026-09-12

### [How llmfit Embeds the React dist/ Folder into the Binary via include_str! in the Build Script](/AlexsJones/llmfit/llmfit-tui-build-script-react-dist-bundling-include_str)

Discover how llmfit embeds the React dist folder into its binary using include_str! in the build script. Learn about static byte slices and compiled Rust modules for your executable.

- Tags: internals
- Published: 2026-09-12

### [How ModelFit Memory and VRAM Fields Feed Into Kubernetes DRA ResourceClaim Specs](/AlexsJones/llmfit/llmfit-core-claim.rs-kubernetes-dra-resourceclaim-spec-memory-vram-fields)

Learn how ModelFit min_ram_gb and min_vram_gb fields populate Kubernetes DRA ResourceClaim specs for memory and GPU. Optimize resource allocation with this guide.

- Tags: internals
- Published: 2026-09-12

### [How llmfit Renders Kubernetes DRA ResourceClaim Manifests from Model Fits](/AlexsJones/llmfit/llmfit-core-claim.rs-kubernetes-dra-resourceclaim-rendering-from-modelfit)

Discover how llmfit Renders Kubernetes DRA ResourceClaim manifests from ModelFits. Learn about transforming model fits into DRA claims with hardware bounds and CEL selectors.

- Tags: internals
- Published: 2026-09-12

### [Where Local, Community, and Measured Presets Take Precedence in llmfit Benchmarks](/AlexsJones/llmfit/llmfit-core-fit.rs-benchmark-preset-precedence)

Discover how llmfit prioritizes local measured benchmarks over community presets for faster, more accurate results. Learn the fitting logic in llmfit-core/src/fit.rs.

- Tags: benchmarks
- Published: 2026-09-12

### [How benchmarks.rs Creates a Measured Throughput Index Overriding the Memory-Bandwidth Formula in llmfit](/AlexsJones/llmfit/llmfit-core-benchmarks.rs-measured-throughput-index-vs-memory-bandwidth-formula)

Discover how benchmarks.rs crafts a measured throughput index, replacing memory bandwidth calculations in llmfit. Access real-world TPS benchmarks for accurate performance insights.

- Tags: deep-dive
- Published: 2026-09-12

### [How SystemSpecs::detect Distinguishes Apple Silicon Unified Memory from Discrete VRAM in llmfit](/AlexsJones/llmfit/llmfit-core-systemspecs-detect-apple-silicon-vs-discrete-vram-cpudoffload-impact)

Discover how SystemSpecs::detect identifies Apple Silicon unified memory, bypassing discrete VRAM and the CpuOffload path for optimal performance in llmfit.

- Tags: deep-dive
- Published: 2026-09-12

### [analyze_with_forced_runtime vs analyze_with_config in llmfit: Runtime Control vs Configuration Tuning](/AlexsJones/llmfit/llmfit-core-fit.rs-analyze_with_forced_runtime-vs-analyze_with_config-run-modes)

Understand analyze_with_forced_runtime vs analyze_with_config in llmfit. Force inference engines or tune calculation parameters to control LLM fitting efficiently. Choose the right run mode for your needs.

- Tags: deep-dive
- Published: 2026-09-12

### [How Themes Are Persisted and Applied Across TUI Sessions in llmfit-tui](/AlexsJones/llmfit/how-are-themes-persisted-tui-sessions)

Learn how llmfit-tui persists and applies themes across TUI sessions by storing theme variants in config files and reloading them on startup.

- Tags: internals
- Published: 2026-09-11

### [How Process Replacement Differs Between Unix and Windows in llmfit-python](/AlexsJones/llmfit/process-replacement-unix-vs-windows-llmfit-python)

Understand process replacement differences in llmfit-python between Unix and Windows. Learn how Unix uses os.execv for PID preservation while Windows uses subprocess.run.

- Tags: internals
- Published: 2026-09-11

### [How the Tauri Desktop Crate Calls llmfit‑core and Manages Ollama Pull State](/AlexsJones/llmfit/tauri-desktop-calling-llmfit-core-ollama-state)

Discover how the Tauri desktop crate integrates with llmfit-core to manage Ollama pull state. Learn about its command facade, PullHandle, and non-blocking poll commands for seamless UI updates.

- Tags: internals
- Published: 2026-09-11

### [How llmfit-core bench.rs Measures Tokens-per-Second Without Blocking the TUI](/AlexsJones/llmfit/how-llmfit-core-bench-rs-measures-tok-s-non-blocking)

Discover how llmfit-core bench.rs measures tokens-per-second without blocking the TUI. Learn how synchronous HTTP benchmarks achieve responsiveness through background threads and callbacks.

- Tags: internals
- Published: 2026-09-11

### [Silent-Failure Paths for GPU VRAM Detection in llmfit: A Complete Analysis](/AlexsJones/llmfit/silent-failure-paths-gpu-vram-detection)

Discover silent-failure paths for GPU VRAM detection in llmfit. Learn how the library handles missing or failing detection tools without logging errors.

- Tags: deep-dive
- Published: 2026-09-11

### [How llmfit Detects GPU VRAM Using nvidia-smi, rocm-smi, system_profiler, and objc2-metal](/AlexsJones/llmfit/how-gpu-vram-detected-nvidia-rocm-system-profiler)

Discover how llmfit detects GPU VRAM using nvidia-smi, rocm-smi, system_profiler, and objc2-metal. Learn about cross-platform GPU memory detection and analysis.

- Tags: internals
- Published: 2026-09-11

### [How llmfit Authenticates Benchmark Submissions to GitHub Without the gh CLI](/AlexsJones/llmfit/benchmark-submission-github-auth-no-gh-cli)

Discover how llmfit authenticates benchmark submissions to GitHub using OAuth Device Flow and environment variables, bypassing the gh CLI. Learn more about this streamlined process.

- Tags: how-to-guide
- Published: 2026-09-11

### [How Crossterm Keybindings Drive Filter, Sort, and Navigation State Changes in llmfit](/AlexsJones/llmfit/crossterm-keybindings-tui-app-state)

Learn how crossterm keybindings in llmfit TUI control filter, sort, and navigation state changes. Discover event handling and app struct mutation for seamless updates.

- Tags: internals
- Published: 2026-09-11

### [How llmfit Maintains Stateless TUI Rendering While Event Handlers Mutate App State](/AlexsJones/llmfit/tui-stateless-rendering-event-handlers)

Discover how llmfit achieves stateless TUI rendering. Learn how event handlers mutate app state while keeping UI rendering functions read-only and efficient.

- Tags: internals
- Published: 2026-09-11

### [How llmfit-core `plan.rs` Estimates Memory, Throughput, and Hardware Upgrade Needs](/AlexsJones/llmfit/how-llmfit-core-plan-rs-estimates-needs)

Discover how llmfit-core plan.rs estimates VRAM, RAM, and TPS. Learn to predict performance and identify hardware upgrade needs for your LLM setup.

- Tags: internals
- Published: 2026-09-11

### [How llmfit-core Discovers Ollama, MLX, llama.cpp, and Other Model Runners](/AlexsJones/llmfit/how-llmfit-providers-discovers-model-runners)

Discover how llmfit-core finds Ollama, MLX, and llama.cpp. Learn how the ModelProvider trait unifies local LLM discovery by probing binaries, endpoints, and caches.

- Tags: internals
- Published: 2026-09-11

### [How the llmfit Axum Server Routes Between the React Dashboard and `/api/v1/` JSON Endpoints](/AlexsJones/llmfit/axum-server-routing-react-api)

Discover how the llmfit Axum server routes between the React dashboard and api v1 JSON endpoints. Learn about nesting routes and client-side routing support.

- Tags: architecture
- Published: 2026-09-11

### [How Hardware Profiles Override Detected SystemSpecs Capacity and CalcConfig Bandwidth in llmfit](/AlexsJones/llmfit/how-hardware-profiles-override-systemspecs-calcconfig)

Discover how llmfit hardware profiles override detected SystemSpecs capacity and CalcConfig bandwidth by replacing raw values and mutating estimator parameters via apply methods.

- Tags: internals
- Published: 2026-09-11

### [How the llmfit-tui MCP Server Exposes Hardware, Model, Runtime, and Planning Tools Over stdio](/AlexsJones/llmfit/how-llmfit-tui-mcp-server-exposes-tools)

Discover how llmfit-tui mcp_server uses stdio to expose hardware, model, runtime, and planning tools. Access powerful features from any language.

- Tags: internals
- Published: 2026-09-11

### [How Kubernetes DRA ResourceClaim Rendering Translates Hardware into Allocatable Resources](/AlexsJones/llmfit/kubernetes-dra-resourceclaim-hardware-translation)

Learn how Kubernetes DRA ResourceClaim rendering translates LLM hardware specs like parameter count and throughput into CEL selectors for efficient resource allocation.

- Tags: how-to-guide
- Published: 2026-09-11

### [How Local, Community, and Measured Benchmark Results Override Throughput Estimates in llmfit](/AlexsJones/llmfit/how-benchmark-results-override-throughput)

Discover how llmfit overrides throughput estimates by prioritizing local, community, and measured benchmarks for accurate, real-world AI performance data.

- Tags: deep-dive
- Published: 2026-09-11

### [The Role of scripts/scrape_hf_models.py in Maintaining the 33-Model Schema](/AlexsJones/llmfit/role-of-scrape_hf_models-py)

Discover how scripts scrape_hf_models.py maintains the 33-model schema for llmfit. This script is the central source for discovering, validating, and refreshing models for hardware-fit calculations.

- Tags: internals
- Published: 2026-09-11

### [How to Regenerate the Embedded hf_models.json Catalog in llmfit](/AlexsJones/llmfit/how-is-hf_models-json-regenerated)

Learn how to regenerate the embedded hf_models.json catalog in llmfit using the scrape_hf_models.py script. Understand the process of fetching and embedding Hugging Face model metadata.

- Tags: how-to-guide
- Published: 2026-09-11

### [llmfit ModelFit Analysis Methods: Context Limits, Runtime Overrides, and Custom Config](/AlexsJones/llmfit/difference-between-model-fit-analyze-methods)

Explore llmfit ModelFit analysis methods. Understand context limits, runtime overrides, and custom config for your LLM analysis with AlexsJones/llmfit. Optimize your models effectively.

- Tags: api-reference
- Published: 2026-09-11

### [Why Apple Silicon Unified Memory Skips CpuOffload in llmfit](/AlexsJones/llmfit/why-apple-silicon-skips-cpuoffload)

Discover why Apple Silicon unified memory skips CpuOffload in llmfit. Learn how direct GPU access to system RAM boosts LLM performance and eliminates data copying.

- Tags: internals
- Published: 2026-09-11

### [How llmfit Chooses Between GPU, MoE-Offload, CPU-Offload, CPU-Only, and Tensor-Parallel Run Paths](/AlexsJones/llmfit/how-does-llmfit-choose-run-paths)

Discover how llmfit selects the best run path GPU, MoE-offload, CPU-offload, CPU-only, or tensor-parallel. Understand its hardware, memory, and speed evaluation.

- Tags: internals
- Published: 2026-09-11

### [How to Use llmfit with a Specific Runtime or Quantization](/AlexsJones/llmfit/how-to-use-llmfit-with-a-specific-runtime-or-quantization)

Master llmfit by learning to specify inference runtimes and quantization levels. Control your AI model's performance and resource usage with simple CLI flags and API parameters. Optimize your llmfit experience today.

- Tags: how-to-guide
- Published: 2026-08-23

### [How llmfit Handles Mixture of Experts (MoE) Models: A Bandwidth-Aware Approach](/AlexsJones/llmfit/how-does-llmfit-handle-mixture-of-experts-moe-models)

Discover how llmfit manages Mixture of Experts MoE models with its bandwidth-aware approach, keeping active experts in VRAM and streaming inactive ones from system RAM.

- Tags: deep-dive
- Published: 2026-08-23

### [Can llmfit Analyze Models with Different Quantization Levels?](/AlexsJones/llmfit/can-llmfit-analyze-models-with-different-quantization-levels)

Discover if llmfit analyzes models with different quantization levels. Evaluate memory, speed, and quality degradation from full-precision down to 2-bit formats.

- Tags: how-to-guide
- Published: 2026-08-23

### [How llmfit Estimates Model Speed (Throughput): The Physics-Based TPS Calculator Explained](/AlexsJones/llmfit/how-does-llmfit-estimate-model-speed-throughput)

Discover how llmfit estimates model speed using its physics-based TPS calculator. Learn the formula and factors influencing token per second calculations for LLMs.

- Tags: internals
- Published: 2026-08-23

### [Understanding llmfit Fit Analysis Dimensions: Quality, Speed, Fit, and Context](/AlexsJones/llmfit/what-are-llmfit-fit-analysis-dimensions)

Discover llmfit fit analysis dimensions: Quality, Speed, Fit, and Context. Understand how these factors rank hardware-model compatibility with a composite score.

- Tags: deep-dive
- Published: 2026-08-23

### [How to Add New Models to the llmfit Catalog: A Step-by-Step Guide](/AlexsJones/llmfit/how-to-add-new-models-to-llmfit-catalog)

Learn how to add new models to the llmfit catalog with this step-by-step guide. Update the scraper script and regenerate the JSON to include your model's Hugging Face repo ID.

- Tags: how-to-guide
- Published: 2026-08-23

### [What Is the llmfit HTTP API? A Complete Endpoint Reference](/AlexsJones/llmfit/what-is-the-llmfit-http-api)

Explore the llmfit HTTP API to query hardware, find LLM models, download them, and calculate execution plans. This Axum-based REST interface empowers local clients with full control over your LLM environment.

- Tags: api-reference
- Published: 2026-08-23

### [llmfit TUI Interface: Interactive Terminal Architecture](/AlexsJones/llmfit/what-is-the-llmfit-tui-interface)

Explore the llmfit TUI interface a ratatui terminal app for AI model hardware compatibility. Discover keyboard driven interaction stateless rendering and real time filtering.

- Tags: architecture
- Published: 2026-08-23

### [How llmfit's Tauri Desktop Application Integrates with the Core Rust Library](/AlexsJones/llmfit/llmfit-tauri-desktop-integration)

Discover how llmfit's Tauri desktop app integrates with the core Rust library. The app exposes llmfit-core functionality via serializable commands for seamless JavaScript to Rust interaction.

- Tags: internals
- Published: 2026-08-22

### [Can llmfit Run on Apple Silicon Macs Using the MLX Framework?](/AlexsJones/llmfit/llmfit-apple-silicon-mlx-support)

Yes llmfit runs natively on Apple Silicon Macs using MLX for fast inference. Discover how to leverage this powerful framework for your machine learning tasks.

- Tags: how-to-guide
- Published: 2026-08-22

### [How llmfit Uses TurboQuant for KV Cache Optimization: Implementation Guide](/AlexsJones/llmfit/llmfit-turboguant-kv-cache-optimization)

Discover how llmfit leverages TurboQuant for efficient KV cache optimization. Learn about 3-bit and 2-bit quantization for memory reduction on vLLM systems.

- Tags: implementation-guide
- Published: 2026-08-22

### [KV Cache Quantization Options in llmfit: Byte Sizes and Memory Optimization Guide](/AlexsJones/llmfit/llmfit-kv-cache-quantization-options)

Explore KV cache quantization in llmfit. Discover options from FP16 (2.0 bytes) to TurboQuant (~0.34 bytes) and achieve up to 83% VRAM reduction for efficient LLM inference.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit's `best_quant_for_budget` Function Works: Automatic Quantization Selection with Context Window Adjustments](/AlexsJones/llmfit/llmfit-best-quant-for-budget-function)

Discover how llmfit's best_quant_for_budget function auto-selects optimal quantization levels within your memory budget, adjusting context windows to meet hardware needs.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Performs Dynamic Quantization Selection for LLM Deployment](/AlexsJones/llmfit/llmfit-dynamic-quantization-selection)

llmfit automatically selects optimal LLM quantization at runtime by scanning compression formats and choosing the first fit for your hardware memory budget.

- Tags: internals
- Published: 2026-08-22

### [llmfit Run Modes: Speed Multipliers and Performance Factors Explained](/AlexsJones/llmfit/llmfit-speed-multipliers-run-modes)

Discover llmfit run modes and their speed multipliers from 1.0 on GPU to 0.3 on CPU. Understand performance factors impacting token generation.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Estimates Inference Speed for Mixture-of-Experts (MoE) Models](/AlexsJones/llmfit/llmfit-moe-model-speed-estimation)

llmfit estimates MoE model inference speed by analyzing active expert weights and memory bandwidth, providing physics-based formulas for GPU and offloaded configurations.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Estimates Speed on Unrecognized GPUs: 7-Step Fallback Method Explained](/AlexsJones/llmfit/llmfit-fallback-speed-estimation-unrecognized-gpu)

Discover how llmfit estimates speed on unrecognized GPUs using a 7-step fallback method. Learn about its deterministic path and fallback techniques.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Calculates Theoretical Maximum Tokens Per Second from GPU Memory Bandwidth](/AlexsJones/llmfit/llmfit-calculate-theoretical-tps)

Discover how llmfit calculates theoretical maximum tokens per second using GPU memory bandwidth and model size. Learn the formula and efficiency factor for precise TPS estimation.

- Tags: internals
- Published: 2026-08-22

### [The Memory-Bandwidth Roof-Line Model Used by llmfit for Estimating LLM Inference Speed](/AlexsJones/llmfit/llmfit-llm-inference-speed-estimation-model)

Discover how llmfit uses a memory-bandwidth roof-line model to estimate LLM inference speed. Learn about its calculation of theoretical throughput and efficiency factors for real-world performance.

- Tags: deep-dive
- Published: 2026-08-22

### [How Default Scoring Weights Vary Across Use Cases in llmfit](/AlexsJones/llmfit/llmfit-scoring-weights-by-use-case)

Discover how llmfit's default scoring weights adapt for diverse use cases. Explore the shifts from quality-focused Reasoning to speed-optimized Embedding.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Calculates Scoring Components: Quality, Speed, Fit, and Context Explained](/AlexsJones/llmfit/llmfit-score-components-calculation)

Understand how llmfit calculates quality, speed, fit, and context scoring components. Learn how each metric is normalized and combined for a final composite score.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Determines the Best Fit Level: Perfect, Good, Marginal, or Too Tight](/AlexsJones/llmfit/llmfit-determine-model-fit-level)

Discover how llmfit determines the best fit level for your model by comparing memory needs with resources. Learn about its two-stage hardware selection and scoring process.

- Tags: internals
- Published: 2026-08-22

### [Understanding llmfit Run Modes: 5 Execution Strategies for LLM Inference](/AlexsJones/llmfit/llmfit-supported-llm-run-modes)

Explore the 5 llmfit Run Modes: Gpu, MoeOffload, CpuOffload, CpuOnly, and TensorParallel. Learn how to efficiently distribute LLM weights across GPU VRAM, system RAM, and clusters for optimal inference.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Calculates Recommended RAM for a Model: The Complete Formula](/AlexsJones/llmfit/llmfit-recommended-ram-formula)

Discover the exact formula llmfit uses to calculate recommended RAM for your model. Learn how it determines minimum memory and applies a safety multiplier for optimal performance.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Estimates RAM and VRAM Requirements for LLM Models](/AlexsJones/llmfit/llmfit-estimate-model-memory-requirements)

Learn how llmfit estimates RAM and VRAM for LLM models by converting parameters to GiB using Q4_K_M quantization, applying safety margins, and accounting for KV-cache overhead.

- Tags: how-to-guide
- Published: 2026-08-22

### [What Is the Schema for Models in llmfit's JSON Catalogs?](/AlexsJones/llmfit/llmfit-model-catalog-schema)

Understand the JSON schema for models in llmfit catalogs like hf_models.json and onnx_models.json. Ensure consistent metadata for Hugging Face models with llmfit.

- Tags: api-reference
- Published: 2026-08-22

### [How to Find the List of Supported LLM Models in llmfit: JSON Database and CLI Access](/AlexsJones/llmfit/llmfit-supported-llm-models-list)

Easily find supported LLM models in llmfit. Access the complete list via the JSON database or the simple llmfit list CLI command.

- Tags: how-to-guide
- Published: 2026-08-22

### [How llmfit Detects Unified Memory Systems Like Apple Silicon](/AlexsJones/llmfit/llmfit-unified-memory-handling)

Discover how llmfit detects unified memory systems like Apple Silicon. Learn about RAM mapping, unified memory flags, and Metal's working set size for optimal GPU utilization.

- Tags: internals
- Published: 2026-08-22

### [GPU Detection Methods in llmfit: How It Identifies NVIDIA, AMD, and Apple Silicon](/AlexsJones/llmfit/llmfit-gpu-detection-methods)

Discover how llmfit detects NVIDIA AMD and Apple Silicon GPUs using nvidia-smi rocm-smi and system_profiler commands. Learn about its hardware detection methods.

- Tags: deep-dive
- Published: 2026-08-22

### [How llmfit Detects Local Hardware Specifications: RAM, CPU, and GPU VRAM](/AlexsJones/llmfit/how-llmfit-detects-hardware-specs)

Discover how llmfit detects local hardware specs. Learn how it uses sysinfo and command-line probes for RAM, CPU, and GPU VRAM detection on your system.

- Tags: internals
- Published: 2026-08-22

### [How LLMFIT plan.rs Calculates Minimum and Recommended VRAM and RAM Requirements](/AlexsJones/llmfit/hardware-estimation-plan.rs-calculate-minimum-recommended-vram-ram-requirements-llmfit)

Discover how LLMFIT plan.rs calculates VRAM and RAM needs for GPU, CPU offload, and CPU-only execution. Learn memory thresholds based on model size and context.

- Tags: internals
- Published: 2026-08-21

### [How LLMFIT's Benchmarking Subsystem Measures Real tok/s and Submits Community Data](/AlexsJones/llmfit/benchmarking-subsystem-measure-real-toks-running-providers-submit-community-data-llmfit)

Discover how LLMFIT measures real tok/s by testing live inference requests to providers like Ollama and vLLM. Learn how it submits community data to enrich the public leaderboard.

- Tags: internals
- Published: 2026-08-21

### [How the LLMFIT Run Mode Selection Algorithm Ranks Fit Levels (Perfect, Good, Marginal, Too Tight)](/AlexsJones/llmfit/run-mode-selection-algorithm-rank-fit-levels-perfect-good-marginal-tootight-llmfit)

Understand LLMFIT run mode selection algorithm's fit level ranking Perfect Good Marginal TooTight. Discover how memory and execution modes determine fit.

- Tags: internals
- Published: 2026-08-21

### [LLMFIT Scoring Weights Per Use Case: How Quality, Speed, Fit, and Context Are Weighted](/AlexsJones/llmfit/scoring-weights-per-use-case-quality-speed-fit-context-llmfit)

Discover LLMFIT scoring weights for quality, speed, fit, and context. Learn how these factors are weighted per use case like Chat, Coding, and Vision to optimize LLM performance.

- Tags: deep-dive
- Published: 2026-08-21

### [How LLMFIT Filters MLX Models to Metal‑Only Systems Using Backend Compatibility Checks](/AlexsJones/llmfit/backend-compatibility-check-filter-mlx-models-metal-only-systems-llmfit)

Discover how LLMFIT ensures Metal-only system compatibility for MLX models using backend checks. Learn how it verifies GpuBackend::Metal and unified memory to prevent GPU execution errors.

- Tags: internals
- Published: 2026-08-21

### [How LLMFIT Reports Full DIMM Capacity on AMD APU Systems with BIOS UMA Carveouts](/AlexsJones/llmfit/amd-apu-bios-uma-carveout-override-report-full-dimm-capacity-llmfit)

LLMFIT overrides AMD APU BIOS UMA carveouts by querying SMBIOS directly through WMI, reporting full DIMM capacity. Learn how this tool reveals true memory size.

- Tags: internals
- Published: 2026-08-21

### [How LLMFIT Estimates MoE Model Speed Using Active vs. Full Parameters](/AlexsJones/llmfit/speed-estimation-moe-models-active-parameters-vs-full-parameters-llmfit)

LLMFIT estimates MoE model speed using active parameters for DDR offload and full parameters for GPU execution, providing accurate tokens per second predictions for sparse architectures.

- Tags: deep-dive
- Published: 2026-08-21

### [How LLMFIT Detects Ollama, llama.cpp, MLX, and LM Studio Runtimes](/AlexsJones/llmfit/provider-detection-system-discover-ollama-llama.cpp-mlx-lm-studio-runtimes-llmfit)

Discover how LLMFIT efficiently detects Ollama, llama.cpp, MLX, and LM Studio runtimes. Our system uses single-pass startup probes for network, binary, and cache checks, building a unified model catalog.

- Tags: internals
- Published: 2026-08-21

### [How Cluster Mode with vLLM Tensor Parallelism Distributes Work Across DGX Spark Nodes in LLMFIT](/AlexsJones/llmfit/cluster-mode-vllm-tensor-parallelism-distribute-work-dgx-spark-nodes-llmfit)

Learn how LLMFIT uses vLLM tensor parallelism in cluster mode to distribute inference workloads across DGX Spark nodes. Discover how model layers are partitioned for efficient distributed computing.

- Tags: deep-dive
- Published: 2026-08-21

### [How LLMFIT Uses Multi-GPU Tensor Splitting to Aggregate VRAM for Model Fit Scoring](/AlexsJones/llmfit/multi-gpu-tensor-splitting-aggregate-vram-same-model-cards-fit-scoring-llmfit)

Discover how LLMFIT aggregates VRAM across identical GPUs for model fit scoring. Learn to run large models across multiple cards with multi-GPU tensor splitting for efficient tensor-parallel inference.

- Tags: internals
- Published: 2026-08-21

### [How TurboQuant KV Compression Enables TooTight Models to Fit on CUDA Systems in LLMFIT](/AlexsJones/llmfit/turboquant-kv-compression-tootight-models-cuda-systems-llmfit)

Discover how TurboQuant KV compression slashes key-value cache size, allowing TooTight LLM models to fit and run on CUDA systems with LLMFIT and vLLM.

- Tags: deep-dive
- Published: 2026-08-21

### [How LLMFIT_DDR_BANDWIDTH Configures MoE Expert Streaming Estimates in llmfit](/AlexsJones/llmfit/llmfit_ddr_bandwidth-environment-variable-moe-expert-streaming-estimates)

Discover how LLMFIT_DDR_BANDWIDTH configures MoE expert streaming estimates in llmfit. Override system RAM calculations for accurate throughput predictions in MoE off-load inference.

- Tags: internals
- Published: 2026-08-21

### [Why MLX Runtime Produces Faster Estimates Than llama.cpp on Apple Silicon in LLMFIT](/AlexsJones/llmfit/mlx-runtime-faster-estimates-llama.cpp-apple-silicon-llmfit)

Discover why MLX Runtime outperforms llama.cpp on Apple Silicon for faster LLM estimates. Learn how direct Metal GPU access boosts performance in LLMFIT.

- Tags: performance
- Published: 2026-08-21

### [How the LLMFIT Advanced Configuration Panel (CalcConfig) Tunes Efficiency and Scoring Weights](/AlexsJones/llmfit/tui-advanced-configuration-calcconfig-tune-efficiency-scoring-weights-llmfit)

Discover how LLMFIT's CalcConfig panel tunes efficiency and scoring weights for accurate tokens-per-second estimates and model fit rankings. Customize GPU, CPU, and MoE multipliers.

- Tags: deep-dive
- Published: 2026-08-21

### [How Context Cap Estimation Prevents KV-Cache Memory Overestimation in LLMFIT](/AlexsJones/llmfit/context-cap-estimation-kv-cache-memory-overestimation-llmfit)

Context cap estimation prevents KV-cache memory overestimation by setting realistic context windows, avoiding false "TooTight" fits and ensuring accurate memory projections for LLMFIT.

- Tags: internals
- Published: 2026-08-21

### [How LLMFIT Validates CUDA Compute Capability for Pre-Quantized Models (AWQ/GPTQ)](/AlexsJones/llmfit/cuda-compute-capability-check-pretrained-model-compatibility-awq-gptq-llmfit)

LLMFIT validates pre-quantized AWQ/GPTQ model compatibility by checking CUDA compute capability, ensuring seamless inference. Discover how LLMFIT ensures GPU readiness.

- Tags: how-to-guide
- Published: 2026-08-21

### [How the VRAM Cache‑Pressure Penalty Affects MoE Throughput Estimation in LLMFIT](/AlexsJones/llmfit/vram-cache-pressure-penalty-moe-throughput-estimation-llmfit)

Discover how VRAM cache-pressure penalty impacts MoE throughput estimation in LLMFIT. Learn how utilization limits affect tokens-per-second calculations.

- Tags: deep-dive
- Published: 2026-08-21

### [LLMFIT Quantization Hierarchies for GGUF, MLX, and ONNX Models](/AlexsJones/llmfit/quantization-hierarchies-supported-gguf-mlx-onnx-models-llmfit)

Discover LLMFIT quantization hierarchies for GGUF, MLX, and ONNX. Explore native GGUF, MLX 4bit/8bit mappings, and ONNX KV-quant levels with TurboQuant for efficient model compression.

- Tags: deep-dive
- Published: 2026-08-21

### [Unified Memory Detection Mechanism in LLMFIT: Apple Silicon vs NVIDIA Grace Blackwell](/AlexsJones/llmfit/unified-memory-detection-apple-silicon-nvidia-grace-blackwell-llmfit)

Discover LLMFIT's unified memory detection for Apple Silicon vs NVIDIA Grace Blackwell. Our mechanism uses system APIs to identify shared RAM and optimize VRAM calculations.

- Tags: deep-dive
- Published: 2026-08-21

### [How the MoE Expert Offloading Path Functions in LLMFIT](/AlexsJones/llmfit/moe-expert-offloading-path-function-when-used-llmfit)

Learn how LLMFITs MoE expert offloading path streams inactive experts from RAM to GPU to enable inference on memory-constrained hardware. Optimize your LLM usage.

- Tags: internals
- Published: 2026-08-21

### [Dynamic Quantization Selection in LLMFIT: How It Automatically Chooses Optimal Model Precision](/AlexsJones/llmfit/dynamic-quantization-selection-optimal-quantization-llmfit)

Discover dynamic quantization selection in LLMFIT. Learn how it automatically finds the best model precision for your hardware, optimizing performance and memory usage.

- Tags: deep-dive
- Published: 2026-08-21

### [How the Memory-Bandwidth Roofline Model Estimates Throughput in LLMFIT](/AlexsJones/llmfit/memory-bandwidth-roofline-model-throughput-estimation-llmfit)

Understand how the memory-bandwidth roofline model estimates LLMFIT throughput. Discover how it calculates theoretical ceilings and applies efficiency factors for accurate token-generation speed predictions.

- Tags: deep-dive
- Published: 2026-08-21

### [GGUF Quantization Impact on Memory Usage and Speed in llmfit](/AlexsJones/llmfit/impact-gguf-quantization-memory-speed-llmfit)

Discover how GGUF quantization in llmfit slashes memory usage by 50% and boosts speed by 1.15x. Learn about bytes-per-parameter compression and quantized multipliers for efficient LLM inference.

- Tags: performance
- Published: 2026-08-20

### [How Quantization Levels Are Defined and Ordered in llmfit: A Complete Developer Guide](/AlexsJones/llmfit/how-quantization-levels-defined-ordered-llmfit)

Discover how llmfit defines and orders quantization levels through three hierarchies. Learn about fidelity, compression, and automatic selection for optimal performance.

- Tags: deep-dive
- Published: 2026-08-20

### [How llmfit Detects Models from Docker Containers: Architecture and Implementation](/AlexsJones/llmfit/how-llmfit-detect-models-docker-containers)

Discover how llmfit detects models from Docker containers. Learn about its architecture and implementation, including verifying Docker, configuring the DMR client, and probing the /v1/models endpoint.

- Tags: architecture
- Published: 2026-08-20

### [How llmfit Handles LM Studio as a Runtime Provider: Complete Integration Guide](/AlexsJones/llmfit/how-llmfit-handle-lm-studio-provider)

Discover how llmfit seamlessly integrates LM Studio as a runtime provider. Learn about installation detection, model discovery, and asynchronous downloads through its REST API.

- Tags: how-to-guide
- Published: 2026-08-20

### [Apple Silicon Specific Runtime Options (MLX) in llmfit: Complete Configuration Guide](/AlexsJones/llmfit/apple-silicon-specific-runtime-options-llmfit-mlx)

Unlock Apple Silicon performance with llmfit's MLX runtime options. Discover how to force or auto-detect MLX for faster language model execution on your Mac.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit Integrates with Ollama for LLM Management: Complete Technical Guide](/AlexsJones/llmfit/how-llmfit-integrate-ollama)

Learn how llmfit integrates with Ollama for seamless LLM management. This technical guide covers model detection, lifecycle, and inference using the OllamaProvider.

- Tags: how-to-guide
- Published: 2026-08-20

### [What Is the Purpose of the `--perfect` Flag in llmfit?](/AlexsJones/llmfit/purpose-perfect-flag-llmfit)

Understand the purpose of the --perfect flag in llmfit. Filter for models with perfect hardware fit, ensuring exact or ample RAM/VRAM match for your system specs.

- Tags: how-to-guide
- Published: 2026-08-20

### [How to Filter llmfit Results by Fit Level Using the CLI](/AlexsJones/llmfit/filter-llmfit-results-by-fit-level-cli)

Filter llmfit results by fit level using the CLI. Learn to use the --min-fit flag with perfect good marginal or tootight values for precise recommendations.

- Tags: how-to-guide
- Published: 2026-08-20

### [How Apple Silicon Unified Memory Affects llmfit Run Modes: Detection & Execution Logic](/AlexsJones/llmfit/how-apple-silicon-unified-memory-affect-llmfit-run-modes)

Discover how Apple Silicon unified memory impacts llmfit run modes, forcing GPU-only or CPU-only execution and bypassing CPU-offload.

- Tags: deep-dive
- Published: 2026-08-20

### [How llmfit Prioritizes VRAM Over System RAM for GPU Systems: A Deep Dive into the Architecture](/AlexsJones/llmfit/how-llmfit-prioritize-vram-over-system-ram)

Discover how llmfit prioritizes VRAM over system RAM for GPU systems. Learn about its architecture and execution modes for optimal performance.

- Tags: architecture
- Published: 2026-08-20

### [What Does a "TooTight" Fit Level Signify in llmfit?](/AlexsJones/llmfit/what-does-tootight-fit-level-signify-llmfit)

Learn what a TooTight fit level means in llmfit. Discover why your model won't load on current hardware due to insufficient GPU or system RAM.

- Tags: deep-dive
- Published: 2026-08-20

### [What Is a Marginal Fit in llmfit? Understanding Memory Constraints and Performance Risks](/AlexsJones/llmfit/implications-marginal-fit-llmfit)

Discover what a marginal fit in llmfit means. Understand memory constraints and performance risks of running models with minimal headroom.

- Tags: deep-dive
- Published: 2026-08-20

### [When Does llmfit Report a "Good" Fit Level?](/AlexsJones/llmfit/when-expect-good-fit-level-llmfit)

Discover when llmfit reports a 'Good' fit level. Learn about the 20% memory headroom requirement and how it compares to the 'Perfect' tier for optimal model performance.

- Tags: how-to-guide
- Published: 2026-08-20

### [What Does a "Perfect" Fit Level Mean in llmfit? Understanding GPU Memory Requirements](/AlexsJones/llmfit/what-does-perfect-fit-level-mean-llmfit)

Discover what a perfect fit level means in llmfit. Understand GPU memory needs and ensure efficient model execution with optimal VRAM utilization. Requires GPU acceleration.

- Tags: deep-dive
- Published: 2026-08-20

### [How to Override Hardware Detection in llmfit Using CLI Flags](/AlexsJones/llmfit/override-hardware-detection-llmfit-cli-flags)

Override llmfit hardware detection with CLI flags like --memory, --ram, and --cpu-cores. Take control of your model-fit analysis settings.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit Handles Multi-GPU VRAM Detection and Aggregation for Model Fitting](/AlexsJones/llmfit/how-llmfit-handle-multi-gpu-setups)

Learn how llmfit detects and aggregates multi-GPU VRAM for efficient model fitting. Discover its intelligent handling of identical GPU models for enhanced performance.

- Tags: internals
- Published: 2026-08-20

### [What Are the Five Execution Paths for LLMs in llmfit? A Deep Dive into RunMode](/AlexsJones/llmfit/what-are-the-five-execution-paths-for-llms)

Explore the five LLM execution paths in llmfit: Gpu, TensorParallel, MoeOffload, CpuOffload, and CpuOnly. Understand the RunMode enum and optimize your LLM performance with this deep dive.

- Tags: deep-dive
- Published: 2026-08-20

### [How GpuBackends Affect LLM Fit Scoring in llmfit: A Technical Deep-Dive](/AlexsJones/llmfit/how-different-gpubackends-affect-llm-fit-scoring)

Explore how GpuBackends impact llmfit scoring. Understand throughput multipliers, feature eligibility, and runtime paths for optimal LLM fitting.

- Tags: deep-dive
- Published: 2026-08-20

### [What GPU Backends Does llmfit Support? A Complete Hardware Abstraction Guide](/AlexsJones/llmfit/what-gpubackends-does-llmfit-support)

Discover the extensive GPU backends llmfit supports including CUDA Metal ROCm Vulkan SYCL Ascend and CPU variants Learn about hardware abstraction in this comprehensive guide

- Tags: api-reference
- Published: 2026-08-20

### [How to Troubleshoot GPU Detection Issues with llmfit doctor](/AlexsJones/llmfit/troubleshoot-gpu-detection-issues-llmfit-doctor)

Troubleshoot GPU detection issues with llmfit doctor. Run llmfit doctor --verbose to check drivers, permissions, and environment variables, common causes of GPU failures.

- Tags: how-to-guide
- Published: 2026-08-20

### [Common Reasons for GPU Detection Failure in llmfit: A Technical Deep-Dive](/AlexsJones/llmfit/common-reasons-gpu-detection-failure-llmfit)

Troubleshoot GPU detection failure in llmfit. Discover common causes like missing CLI tools, restricted sysfs access, and permission issues. Resolve your hardware detection problems now.

- Tags: deep-dive
- Published: 2026-08-20

### [How llmfit Detects System RAM and CPU Cores: Complete Implementation Guide](/AlexsJones/llmfit/how-llmfit-detect-system-ram-and-cpu-cores)

Learn how llmfit detects system RAM and CPU cores using its SystemSpecs::detect() method. Discover the implementation details and platform-specific fallbacks in this guide.

- Tags: how-to-guide
- Published: 2026-08-20

### [Privacy Implications of llmfit's Network Operations: A Technical Analysis](/AlexsJones/llmfit/llmfit-privacy-network-operations)

Analyze llmfit network operations privacy. Discover what data llmfit sends, how it protects user privacy, and its security measures in this technical deep dive.

- Tags: deep-dive
- Published: 2026-08-20

### [How to Integrate llmfit with LM Studio as a Local Model Provider](/AlexsJones/llmfit/llmfit-lm-studio-integration)

Easily integrate llmfit with LM Studio as your local model provider. Run LM Studio on port 1234 and llmfit automatically detects it with zero configuration needed. Streamline your AI development today.

- Tags: how-to-guide
- Published: 2026-08-20

### [Understanding llmfit's Context Cap Behavior: Why 8192 Tokens Is the Default](/AlexsJones/llmfit/llmfit-context-cap-behavior)

Discover llmfit's context cap behavior and why 8192 tokens is the default. Learn how llmfit limits memory estimation and when it falls back to this default value without explicit flags.

- Tags: internals
- Published: 2026-08-20

### [How to Use `llmfit doctor` for Hardware Detection Reports in LLMFit](/AlexsJones/llmfit/llmfit-doctor-command-usage)

Learn how to use llmfit doctor to generate detailed hardware detection reports. Get system info, GPU details, and runtime data formatted for GitHub issues.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit Detects VRAM Across CUDA, ROCm, Vulkan, and SYCL GPU Backends](/AlexsJones/llmfit/llmfit-vram-detection-gpu-backends)

Discover how llmfit detects VRAM across CUDA, ROCm, Vulkan, and SYCL GPU backends by probing vendor utilities and OS interfaces for accurate memory reporting.

- Tags: internals
- Published: 2026-08-20

### [How to Configure llmfit Efficiency and Run Mode Factors for TPS Calibration](/AlexsJones/llmfit/llmfit-configure-tps-calibration)

Calibrate llmfit token per second TPS predictions using efficiency and run mode factors in CalcConfig. Optimize your model's performance against real hardware.

- Tags: how-to-guide
- Published: 2026-08-20

### [LLMFIT Default Scoring Weights Explained: General, Coding, Reasoning, Chat, Multimodal](/AlexsJones/llmfit/llmfit-default-scoring-weights)

Discover LLMFIT default scoring weights for General, Coding, Reasoning, Chat, and Multimodal. Understand how quality, speed, fit, and context influence scores.

- Tags: deep-dive
- Published: 2026-08-20

### [How llmfit Community Benchmark Submission Works Without CLI or GitHub Accounts](/AlexsJones/llmfit/llmfit-community-benchmark-submission)

Discover how to submit llmfit community benchmarks without CLI or GitHub accounts. Learn about local JSON storage, OAuth device flow, and direct API calls for anonymous contributions.

- Tags: how-to-guide
- Published: 2026-08-20

### [How to Run llmfit TUI in Docker and Access JSON Output](/AlexsJones/llmfit/llmfit-docker-tui-json-output)

Easily run llmfit TUI in Docker and access JSON output. Override entrypoint with --tui and use --json for machine-readable results. Get started now.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit's Cluster Mode and TensorParallel Enable Multi-Node Inference](/AlexsJones/llmfit/llmfit-cluster-mode-tensor-parallel)

Discover how llmfit's cluster mode and TensorParallel orchestrate multi-node inference using vLLM for efficient distributed LLM deployment and human readable cluster descriptions.

- Tags: internals
- Published: 2026-08-20

### [How llmfit Integrates with Ollama, llama.cpp, MLX, and vLLM Runtimes](/AlexsJones/llmfit/llmfit-provider-integration)

Discover how llmfit unifies Ollama, llama.cpp, MLX, and vLLM runtimes with a common ModelProvider trait. Seamlessly integrate local LLM inference across multiple backends.

- Tags: deep-dive
- Published: 2026-08-20

### [How to Use llmfit's JSON API for Scripting and Automation](/AlexsJones/llmfit/llmfit-json-api-usage)

Automate model selection and deployment with llmfit's JSON API. Leverage its REST API for scripting and seamless integration.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit Detects Windows APU RAM With BIOS GPU Carveout: A Deep Dive into the WMI-Based Correction](/AlexsJones/llmfit/llmfit-windows-apu-ram-detection)

Learn how llmfit corrects Windows APU RAM detection using WMI to bypass BIOS GPU carveout. Discover the solution for accurate memory reporting on AMD Ryzen AI systems.

- Tags: deep-dive
- Published: 2026-08-20

### [llmfit TUI vs CLI Mode: What Is the Difference and When to Use Each](/AlexsJones/llmfit/llmfit-tui-vs-cli-modes)

Explore llmfit TUI vs CLI modes. Discover when to use the interactive terminal interface for exploration or command line for scripting and one-shot commands.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit's Benchmark and Share Feature Generates tok/s Measurements](/AlexsJones/llmfit/llmfit-benchmark-share-feature)

llmfit measures tok/s by combining local benchmarking with community data for accurate hardware performance insights. Discover how real-world results are generated.

- Tags: internals
- Published: 2026-08-20

### [How to Add Custom Models to llmfit's Catalog Without Recompiling](/AlexsJones/llmfit/llmfit-add-custom-models)

Easily add custom llmfit models without recompiling. Create a JSON overlay file in your user data directory for automatic integration. Enhance your llmfit experience today.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit Detects and Manages Apple Silicon Unified Memory: A Deep Dive into the Rust Implementation](/AlexsJones/llmfit/llmfit-apple-silicon-unified-memory)

Discover how llmfit detects and manages Apple Silicon unified memory by parsing system profiler output and leveraging Metal memory queries. Optimize your AI workloads.

- Tags: deep-dive
- Published: 2026-08-20

### [LLMFit Quantization Options: Supported Formats and Dynamic Selection Methods](/AlexsJones/llmfit/llmfit-quantization-options)

Explore LLMFit quantization options including GGUF MLX ONNX AWQ and GPTQ. Discover dynamic selection methods driven by memory budget for efficient model loading.

- Tags: deep-dive
- Published: 2026-08-20

### [How to Configure Custom Scoring Weights for Different llmfit Use Cases](/AlexsJones/llmfit/llmfit-custom-scoring-weights)

Learn to configure custom scoring weights for llmfit use cases. Adjust Quality, Speed, Fit, and Context importance using the ScoringWeights struct in fit.rs for tailored LLM performance.

- Tags: how-to-guide
- Published: 2026-08-20

### [How llmfit's Memory-Bandwidth TPS Estimation Model Works: A Technical Deep Dive](/AlexsJones/llmfit/llmfit-tps-estimation-model)

Explore llmfit's memory-bandwidth TPS estimation model. Learn how it treats inference as a memory-bound problem to calculate token throughput by dividing bandwidth by per-token memory traffic.

- Tags: deep-dive
- Published: 2026-08-20

### [llmfit Run Modes Explained: GPU, MoE Offload, CPU Offload, CPU Only, and Tensor Parallel](/AlexsJones/llmfit/llmfit-run-modes-comparison)

Explore llmfit run modes: GPU, MoE Offload, CPU Offload, CPU Only, and Tensor Parallel. Optimize LLM execution across your hardware for maximum performance and efficiency.

- Tags: deep-dive
- Published: 2026-08-20

### [How llmfit Handles Multi-GPU Configurations Across Different Vendors](/AlexsJones/llmfit/llmfit-multi-gpu-different-vendors)

llmfit seamlessly manages multi-GPU setups from various vendors, unifying VRAM and enabling efficient tensor-parallel inference on mixed hardware.

- Tags: how-to-guide
- Published: 2026-08-20

### [Where Is the llmfit Model Database Stored? A Complete Guide to the Embedded JSON Catalog](/AlexsJones/llmfit/where-is-the-llmfit-model-database-stored)

Discover where the llmfit model database is stored. This guide explains the embedded JSON catalog within the llmfit Rust crate and how it's compiled into the binary.

- Tags: how-to-guide
- Published: 2026-08-19

### [How to Benchmark LLM Models with llmfit: A Complete Technical Guide](/AlexsJones/llmfit/how-to-benchmark-llm-models-with-llmfit)

Benchmark LLM models using llmfit. Learn the three-stage pipeline for measuring performance including tokens per second and time to first token. A complete technical guide.

- Tags: how-to-guide
- Published: 2026-08-19

### [How to Disable the llmfit Web Dashboard](/AlexsJones/llmfit/how-to-disable-the-llmfit-web-dashboard)

Learn how to disable the llmfit Web dashboard by passing the --no-dashboard flag to the CLI binary. Stop the automatic background web server startup with this simple command.

- Tags: how-to-guide
- Published: 2026-08-19

### [How to Access the llmfit Web Dashboard: Complete Setup Guide](/AlexsJones/llmfit/how-to-access-the-llmfit-web-dashboard)

Learn how to access the llmfit Web dashboard with this complete setup guide. Run the binary with the --serve flag and open the provided URL to get started.

- Tags: how-to-guide
- Published: 2026-08-19

### [Complete Guide to the llmfit API Endpoints](/AlexsJones/llmfit/what-are-the-llmfit-api-endpoints)

Explore the llmfit API endpoints. Query hardware, manage model fits, download weights, and generate execution plans using our Axum-powered server. Get started today!

- Tags: api-reference
- Published: 2026-08-19

### [How to Serve the llmfit REST API: Complete Setup and Endpoint Guide](/AlexsJones/llmfit/how-to-serve-the-llmfit-rest-api)

Learn to serve the llmfit REST API using the llmfit-tui binary. This guide covers Axum server setup and JSON endpoints for model discovery, system specs, and deployment planning.

- Tags: how-to-guide
- Published: 2026-08-19

### [Complete Guide to llmfit CLI Subcommands: 10 Essential Commands for LLM Deployment](/AlexsJones/llmfit/what-are-the-available-llmfit-cli-subcommands)

Explore 10 essential llmfit CLI subcommands like fit, serve, and bench for efficient LLM deployment. Discover hardware detection, model catalog queries, and performance benchmarking.

- Tags: how-to-guide
- Published: 2026-08-19

### [How to Use llmfit in CLI Mode: Command-Line Interface Guide](/AlexsJones/llmfit/how-do-i-use-llmfit-in-cli-mode)

Master llmfit CLI mode. Run analysis directly in your terminal with fit, recommend, or info subcommands for efficient data processing. Get table or JSON output now.

- Tags: how-to-guide
- Published: 2026-08-19

### [Complete llmfit TUI Keyboard Shortcuts Reference: Modal Navigation Guide](/AlexsJones/llmfit/what-are-the-tui-keyboard-shortcuts-in-llmfit)

Master llmfit TUI keyboard shortcuts with this comprehensive guide. Learn modal navigation, list scrolling, visual selection, and essential commands for efficient interaction.

- Tags: api-reference
- Published: 2026-08-19

### [How llmfit Integrates with LM Studio: A Deep Dive into the Provider Implementation](/AlexsJones/llmfit/how-does-llmfit-integrate-with-lm-studio)

Discover how llmfit integrates with LM Studio using the LmStudioProvider. Learn about server discovery, model queries, and asynchronous downloads via the local REST API.

- Tags: deep-dive
- Published: 2026-08-19

### [Using llmfit with MLX on Apple Silicon: Setup and Configuration Guide](/AlexsJones/llmfit/can-llmfit-be-used-with-mlx-on-apple-silicon)

Learn how to use llmfit with MLX on Apple Silicon. This guide details setup and configuration for seamless MLX inference on compatible hardware.

- Tags: getting-started
- Published: 2026-08-19

### [How to Set Up llmfit with llama.cpp: Complete Integration Guide](/AlexsJones/llmfit/how-to-set-up-llmfit-with-llamacpp)

Easily set up llmfit with llama.cpp using our complete integration guide. Discover how to drive recommendations with the --force-runtime llamacpp flag for seamless performance.

- Tags: how-to-guide
- Published: 2026-08-19

### [Which LLM Runtime Providers Are Supported by llmfit? A Complete Guide](/AlexsJones/llmfit/which-llm-runtime-providers-are-supported-by-llmfit)

Discover which LLM runtime providers llmfit supports including Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio. Effortlessly switch backends with the TUI picker.

- Tags: deep-dive
- Published: 2026-08-19

### [How llmfit Performs Dynamic Quantization for LLMs](/AlexsJones/llmfit/how-does-llmfit-perform-dynamic-quantization-for-llms)

Discover how llmfit achieves dynamic quantization for LLMs. llmfit automatically selects optimal quantization levels to fit hardware memory, maximizing quality within your budget.

- Tags: deep-dive
- Published: 2026-08-19

### [LLM Run Modes in llmfit: A Complete Guide to GPU, CPU, and Distributed Execution](/AlexsJones/llmfit/what-are-the-different-llm-run-modes-supported-by-llmfit)

Explore llmfit's LLM run modes: Gpu, MoeOffload, CpuOffload, CpuOnly, and TensorParallel. Optimize inference by distributing model weights across GPU VRAM, RAM, and clusters for peak performance.

- Tags: deep-dive
- Published: 2026-08-19

### [How LLM Model Quality is Scored in llmfit: A Technical Deep Dive into the Benchmarking Engine](/AlexsJones/llmfit/how-is-llm-model-quality-scored-in-llmfit)

Discover how llmfit scores LLM model quality using regex rubrics and composite metrics. Learn about accuracy, inference speed, and automated routing in this technical deep dive.

- Tags: deep-dive
- Published: 2026-08-19

### [How llmfit Estimates LLM Inference Speed: A Memory-Bandwidth Roofline Model](/AlexsJones/llmfit/how-does-llmfit-estimate-llm-inference-speed)

Discover how llmfit estimates LLM inference speed using a memory-bandwidth roofline model. Learn its prediction methodology for tokens-per-second throughput.

- Tags: deep-dive
- Published: 2026-08-19

### [LLM Model Fitting Categories in llmfit: A Complete Guide to FitLevel Classification](/AlexsJones/llmfit/what-are-the-llm-model-fitting-categories-in-llmfit)

Discover the four LLM model fitting categories in llmfit: Perfect, Good, Marginal, and TooTight. Learn how FitLevel classifies hardware compatibility for your LLM models.

- Tags: deep-dive
- Published: 2026-08-19

### [How to Override Detected Hardware Specs in llmfit: A Complete Guide](/AlexsJones/llmfit/how-can-i-override-detected-hardware-specs-in-llmfit)

Override llmfit hardware detection with --ram --memory and --cpu-cores flags. Simulate any system for model fit testing in this complete guide.

- Tags: how-to-guide
- Published: 2026-08-19

### [What Hardware Information Does llmfit Collect? A Complete Technical Guide](/AlexsJones/llmfit/what-hardware-information-does-llmfit-collect)

Discover what hardware information llmfit collects including RAM CPU and GPU details. Learn about system specs and telemetry for optimal performance.

- Tags: deep-dive
- Published: 2026-08-19

### [How llmfit Detects Hardware Specifications: RAM, CPU, and GPU Detection Explained](/AlexsJones/llmfit/how-does-llmfit-detect-my-hardware-specifications)

Discover how llmfit detects hardware specifications like RAM, CPU, and GPU. Learn about the platform-specific probes and tools used by llmfit for accurate system profiling.

- Tags: internals
- Published: 2026-08-19

### [How to Run llmfit in Docker Interactively](/AlexsJones/llmfit/run-llmfit-docker-interactively)

Learn to run llmfit in Docker interactively using the terminal UI. Explore LLM model fitting with hardware detection, just like a native build. Get started now.

- Tags: how-to-guide
- Published: 2026-07-23

### [llmfit recommend Filtering Options: A Complete CLI Reference](/AlexsJones/llmfit/filtering-options-llmfit-recommend-command)

Explore llmfit recommend filtering options in this CLI reference. Discover eight flags to filter AI models by fit, use case, runtime, capabilities, and license. Maximize your model selection.

- Tags: api-reference
- Published: 2026-07-23

### [How llmfit Handles Unified Memory on Apple Silicon: Detection and Optimization Strategy](/AlexsJones/llmfit/llmfit-handle-unified-memory-apple-silicon)

Discover how llmfit masters unified memory on Apple Silicon. Learn its detection methods and optimization strategies for efficient CPU and GPU inference.

- Tags: internals
- Published: 2026-07-23

### [How to Add Custom Models to llmfit Without Recompiling](/AlexsJones/llmfit/add-custom-models-llmfit-without-recompiling)

Learn how to add custom models to llmfit without recompiling. Use a custom_models.json overlay file to easily integrate new models with the runtime loader.

- Tags: how-to-guide
- Published: 2026-07-23

### [How the llmfit Dynamic Quantization Selection Algorithm Works](/AlexsJones/llmfit/llmfit-dynamic-quantization-selection-algorithm)

Discover how the llmfit dynamic quantization selection algorithm automatically finds the best quantization format for your LLM within memory limits. Learn more about this efficient technique.

- Tags: internals
- Published: 2026-07-23

