What Is the llmfit HTTP API? A Complete Endpoint Reference
The llmfit HTTP API is a REST-style JSON web interface built with the Axum framework that lets you query hardware specs, discover compatible LLM models, trigger model downloads, and calculate resource-aware execution plans from any local client.
The AlexsJones/llmfit project bundles a lightweight, self-contained web server inside the llmfit-tui crate. It is designed so external tools, scripts, or a browser can interact with llmfit's core logic without needing to touch the TUI. This article walks through the llmfit HTTP API surface, every endpoint, its parameters, response formats, and working code examples, all grounded directly in the source code.
Understanding the llmfit HTTP API Architecture
The server entry point lives in llmfit-tui/src/serve_api.rs. The run_serve function starts the server, build_router registers all routes, and the server binds to either a TCP socket or a Unix domain socket. All handlers share a common AppState (defined at lines 31-40) that holds:
- Node name and OS metadata
- Detected system specifications
- The full model list
- Optional context limit
- Download-tracking state
The router registers every API endpoint under the /api/v1/ namespace alongside UI fallback routes. Handlers return JSON wrapped in an ApiEnvelope that includes node info, system specs, pagination data, active filters, and an array of model objects. Errors are standardized through the ApiError type, which serializes to a JSON payload with a proper HTTP status code.
All llmfit HTTP API Endpoints and Their Parameters
The llmfit HTTP API exposes eight primary endpoints. Here's the full break down.
GET /health — Liveness Check
A minimal availability probe. Returns { "status": "ok", "node": {...} } from the health handler (lines 66-74 of serve_api.rs).
GET /api/v1/system — Hardware and Node Information
Reports node metadata and detected hardware specs. Supports ram_gb, vram_gb, and cpu_cores as optional query parameters to override auto-detected values.
GET /api/v1/system?ram_gb=64&vram_gb=24&cpu_cores=16
GET /api/v1/models — List, Filter, and Sort Models
The workhorse endpoint. It accepts a rich set of query parameters:
| Parameter | Meaning |
|---|---|
limit |
Page size for returned models |
top |
Number of top results to consider |
perfect |
Filter to only models with perfect fit |
min_fit |
Minimum fit level (e.g., good) |
runtime |
Target runtime (e.g., vllm) |
use_case |
Use case filter (e.g., coding) |
provider |
Model provider |
search |
Text search across model names |
sort |
Sort key (e.g., score) |
include_too_tight |
Include models that exceed memory budget |
max_context |
Maximum context window |
force_runtime |
Force a particular runtime |
license |
License filter |
ram_gb, vram_gb, cpu_cores |
Hardware overrides |
Example:
GET /api/v1/models?limit=20&min_fit=good&sort=score&runtime=vllm
GET /api/v1/models/top — Best-Fit Models Only
Same parameters as /models, but returns only the best-fitting models. Default top=5.
GET /api/v1/models/{name} — Retrieve a Single Model
Inherits all ModelsQuery parameters; the search parameter is effectively forced to the supplied name.
GET /api/v1/models/llama-2-7b
GET /api/v1/runtimes — Discover Installed Runtimes
Returns a list of runtime names (Ollama, MLX, llama.cpp, Docker, LM Studio, vLLM) with a boolean installed flag for each.
GET /api/v1/installed — List Installed Models
Returns a list of {name, runtime} objects for models already present on the host across every detected runtime.
POST /api/v1/download — Start a Model Pull
Starts a model download. Only reachable from localhost for safety. The JSON body is { model: string, runtime: string }. Returns a download ID for polling.
GET /api/v1/download/{id}/status — Poll Download Progress
Returns the current pull status, including a progress_pct field and a human-readable message.
POST /api/v1/plan — Estimate an Execution Plan
Computes a resource-aware plan for running a given model at a given context size. The JSON body accepts:
{
"model": "llama-2-7b",
"context": 8192,
"quant": "q4_0",
"target_tps": 10,
"kv_quant": "q8_0",
"ram_gb": 48,
"vram_gb": 12,
"cpu_cores": 8
}
The response is the raw PlanEstimate JSON from llmfit_core::plan.
llmfit HTTP API Response Shapes
Every successful response is wrapped in a standard envelope:
{
"node": { "name": "...", "os": "..." },
"system": { ... },
"total_models": 123,
"returned_models": 20,
"filters": { ... },
"models": [
{
"name": "...",
"provider": "...",
"fit_level": "...",
"runtime": "...",
"score": ...,
"throughput_tps": ...,
"memory_gb": ...,
...
}
]
}
Special responses are used for the simpler endpoints:
/health→{ "status": "ok", "node": {...} }/runtimes→{ "runtimes": [ { "name": "ollama", "installed": true } ], "warnings": [] }/installed→{ "models": [ { "name": "llama-2-7b", "runtime": "ollama" } ], "warnings": [] }- Download start →
{ "id": "dl-0", "model": "...", "runtime": "...", "status": "pulling" }
Practical Code Examples for the llmfit HTTP API
Here are seven working Rust snippets that demonstrate typical client usage. All assume the server is listening on http://localhost:8080.
1. Health Check
let resp = reqwest::get("http://localhost:8080/health")
.await?.json::<serde_json::Value>().await?;
println!("{:#}", resp);
2. Query System Specs with Hardware Overrides
let url = "http://localhost:8080/api/v1/system?ram_gb=64&vram_gb=24&cpu_cores=16";
let sys = reqwest::get(url).await?.json::<serde_json::Value>().await?;
println!("System: {}", sys["system"]);
3. List Top 5 Coding Models (Good Fit)
let url = "http://localhost:8080/api/v1/models/top?use_case=coding&min_fit=good";
let list = reqwest::get(url).await?.json::<serde_json::Value>().await?;
for m in list["models"].as_array().unwrap() {
println!("{} – score {}", m["name"], m["score"]);
}
4. Retrieve a Specific Model
let name = "llama-2-7b";
let url = format!("http://localhost:8080/api/v1/models/{}", name);
let model = reqwest::get(&url).await?.json::<serde_json::Value>().await?;
println!("Details for {}: {:#?}", name, model["models"]);
5. Discover Installed Runtimes
let runtimes = reqwest::get("http://localhost:8080/api/v1/runtimes")
.await?.json::<serde_json::Value>().await?;
println!("Available runtimes: {:?}", runtimes["runtimes"]);
6. Start a Model Download and Poll Status
let client = reqwest::Client::new();
let body = serde_json::json!({ "model": "llama-2-7b", "runtime": "ollama" });
let start = client.post("http://localhost:8080/api/v1/download")
.json(&body).send().await?.json::<serde_json::Value>().await?;
let id = start["id"].as_str().unwrap();
let status_url = format!("http://localhost:8080/api/v1/download/{}/status", id);
let status = reqwest::get(&status_url).await?.json::<serde_json::Value>().await?;
println!("Download status: {}", status["status"]);
7. Estimate a Plan for a Model
let plan_body = serde_json::json!({
"model": "llama-2-7b",
"context": 8192,
"quant": "q4_0",
"ram_gb": 48,
"vram_gb": 12,
"cpu_cores": 8
});
let estimate = client.post("http://localhost:8080/api/v1/plan")
.json(&plan_body).send().await?.json::<serde_json::Value>().await?;
println!("Estimated TPS: {}", estimate["tps"]);
Key Source Files Behind the llmfit HTTP API
| File | Purpose |
|---|---|
llmfit-tui/src/serve_api.rs |
Implements the HTTP server, routes, request parsing, and response formatting |
llmfit-tui/src/serve_shared.rs |
Helpers like system_json and fit_to_json that serialize core structures |
llmfit-core/src/models.rs |
Defines the LlmModel metadata used by the API |
llmfit-core/src/fit.rs |
Contains ModelFit, FitLevel, and ranking logic |
llmfit-core/src/providers.rs |
Implements runtime providers queried by /runtimes and /installed |
llmfit-core/src/plan.rs |
Performs the planning calculation for /plan |
Summary
- The llmfit HTTP API is an Axum-based REST server defined in
llmfit-tui/src/serve_api.rs, started byrun_serveand routed bybuild_router. - Endpoints live under
/api/v1/and cover system info, model listing/filtering, top picks, install discoveries, runtime detection, download control, and plan estimation. - All responses are wrapped in a consistent
ApiEnvelopeJSON structure, and errors use theApiErrortype with proper HTTP status codes. - The download endpoint is localhost-only, so remote access is deliberately restricted for security.
- The
/planendpoint performs full resource-aware estimation usingllmfit_core::plan, including memory usage and throughput projections.
Frequently Asked Questions
Is the llmfit HTTP API available over the network?
No. Only the download endpoint (POST /api/v1/download) is explicitly restricted to localhost, but the API is fundamentally designed for local tooling. You can bind it to a TCP socket or Unix socket, but remote-exposing it is not a use the source code encourages.
What runtimes does the llmfit HTTP API detect?
The /api/v1/runtimes endpoint detects, in no particular order: Ollama, MLX, llama.cpp, Docker, LM Studio, and vLLM. Each one is returned with a boolean installed flag.
How do I filter models with the llmfit API?
Use the /api/v1/models endpoint with query parameters such as min_fit=good, runtime=vllm, and sort=score. For a single model, query /api/v1/models/{name}.
Can I override my hardware specs in the API response?
Yes. The /api/v1/system and /api/v1/models endpoints both accept ram_gb, vram_gb, and cpu_cores query parameters, which override the detected hardware before any fit or plan calculation runs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →