whichllm

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

25 articles 3.9k View on GitHub ↗
25 articles
How to Display Detected Hardware Using WhichLLM

Learn how to display detected hardware using WhichLLM. This guide shows you how to import and call functions to easily view GPU, CPU, RAM, and OS details.

how-to-guide
Jun 10, 2026
How to Download and Chat with a Model Using the WhichLLM Run Command

Easily download and chat with LLMs using the whichllm run command. This tool handles downloads, dependencies, and launches chat sessions instantly. Get started now!

how-to-guide
Jun 10, 2026
How to Compare GPU Upgrade Options with WhichLLM: A Complete Guide

Compare GPU upgrade options with WhichLLM. Simulate graphics cards to see how they impact language model recommendations, quality, and inference speed. Optimize your hardware now.

how-to-guide
Jun 10, 2026
Complete Guide to WhichLLM CLI Commands: Usage and Examples

Discover WhichLLM CLI commands like default, plan, upgrade, run, snippet, and hardware. Learn hardware detection, model ranking, and interactive execution with this comprehensive guide.

how-to-guide
Jun 10, 2026
Bytes per Weight Values for Quantization Types in WhichLLM: Complete Reference Guide

Discover bytes per weight for WhichLLM quantization types. This guide details the QUANT_BYTES_PER_WEIGHT constant for accurate model storage and VRAM calculations.

api-reference
Jun 10, 2026
How WhichLLM's Ranking Engine Scores Models: The Complete Algorithm Guide

Discover how WhichLLM's ranking engine scores models using seven signals including benchmarks speed and popularity. Understand the complete algorithm to find optimal LLM variants for your hardware.

deep-dive
Jun 10, 2026
What Is the TTL for WhichLLM's Model Cache? A 6-Hour Technical Deep Dive

Discover the exact 6-hour TTL for WhichLLM's model cache, detailed in src/whichllm/models/cache.py. Understand cache expiration for optimal performance.

deep-dive
Jun 10, 2026
How Does WhichLLM Group Similar LLM Models? A Two-Pass Clustering Approach

Discover how WhichLLM groups similar LLM models with its innovative two-pass clustering approach. It leverages Hugging Face metadata and name normalization for accurate model family identification.

internals
Jun 10, 2026
What Parameters Does WhichLLM Extract from HuggingFace Models?

Discover what parameters WhichLLM extracts from HuggingFace models, including technical specs, licensing, popularity, quantization, and benchmarks. Access detailed ModelInfo.

deep-dive
Jun 10, 2026
How to Use the WhichLLM GPU Simulator for Planning: Complete Guide with Code Examples

Plan LLM deployments with the WhichLLM GPU simulator. Emulate any GPU configuration virtually, saving costs and ensuring compatibility before hardware purchase. Get the complete guide with code.

how-to-guide
Jun 10, 2026
What CPU Information Does WhichLLM Gather? A Technical Deep Dive

Discover what CPU info WhichLLM collects Including model name core count and instruction set support to optimize LLM deployment.

deep-dive
Jun 10, 2026
Which Libraries Does WhichLLM Use for NVIDIA GPU Detection?

Discover which libraries WhichLLM uses for NVIDIA GPU detection. Learn how pynvml and nvidia-smi enable efficient GPU monitoring for your LLM projects.

internals
Jun 10, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →