What Is the llmfit HTTP API? A Complete Endpoint Reference

The llmfit HTTP API is a REST-style JSON web interface built with the Axum framework that lets you query hardware specs, discover compatible LLM models, trigger model downloads, and calculate resource-aware execution plans from any local client.

The AlexsJones/llmfit project bundles a lightweight, self-contained web server inside the llmfit-tui crate. It is designed so external tools, scripts, or a browser can interact with llmfit's core logic without needing to touch the TUI. This article walks through the llmfit HTTP API surface, every endpoint, its parameters, response formats, and working code examples, all grounded directly in the source code.

Understanding the llmfit HTTP API Architecture

The server entry point lives in llmfit-tui/src/serve_api.rs. The run_serve function starts the server, build_router registers all routes, and the server binds to either a TCP socket or a Unix domain socket. All handlers share a common AppState (defined at lines 31-40) that holds:

  • Node name and OS metadata
  • Detected system specifications
  • The full model list
  • Optional context limit
  • Download-tracking state

The router registers every API endpoint under the /api/v1/ namespace alongside UI fallback routes. Handlers return JSON wrapped in an ApiEnvelope that includes node info, system specs, pagination data, active filters, and an array of model objects. Errors are standardized through the ApiError type, which serializes to a JSON payload with a proper HTTP status code.

All llmfit HTTP API Endpoints and Their Parameters

The llmfit HTTP API exposes eight primary endpoints. Here's the full break down.

GET /health — Liveness Check

A minimal availability probe. Returns { "status": "ok", "node": {...} } from the health handler (lines 66-74 of serve_api.rs).

GET /api/v1/system — Hardware and Node Information

Reports node metadata and detected hardware specs. Supports ram_gb, vram_gb, and cpu_cores as optional query parameters to override auto-detected values.


GET /api/v1/system?ram_gb=64&vram_gb=24&cpu_cores=16

GET /api/v1/models — List, Filter, and Sort Models

The workhorse endpoint. It accepts a rich set of query parameters:

Parameter Meaning
limit Page size for returned models
top Number of top results to consider
perfect Filter to only models with perfect fit
min_fit Minimum fit level (e.g., good)
runtime Target runtime (e.g., vllm)
use_case Use case filter (e.g., coding)
provider Model provider
search Text search across model names
sort Sort key (e.g., score)
include_too_tight Include models that exceed memory budget
max_context Maximum context window
force_runtime Force a particular runtime
license License filter
ram_gb, vram_gb, cpu_cores Hardware overrides

Example:


GET /api/v1/models?limit=20&min_fit=good&sort=score&runtime=vllm

GET /api/v1/models/top — Best-Fit Models Only

Same parameters as /models, but returns only the best-fitting models. Default top=5.

GET /api/v1/models/{name} — Retrieve a Single Model

Inherits all ModelsQuery parameters; the search parameter is effectively forced to the supplied name.


GET /api/v1/models/llama-2-7b

GET /api/v1/runtimes — Discover Installed Runtimes

Returns a list of runtime names (Ollama, MLX, llama.cpp, Docker, LM Studio, vLLM) with a boolean installed flag for each.

GET /api/v1/installed — List Installed Models

Returns a list of {name, runtime} objects for models already present on the host across every detected runtime.

POST /api/v1/download — Start a Model Pull

Starts a model download. Only reachable from localhost for safety. The JSON body is { model: string, runtime: string }. Returns a download ID for polling.

GET /api/v1/download/{id}/status — Poll Download Progress

Returns the current pull status, including a progress_pct field and a human-readable message.

POST /api/v1/plan — Estimate an Execution Plan

Computes a resource-aware plan for running a given model at a given context size. The JSON body accepts:

{
  "model": "llama-2-7b",
  "context": 8192,
  "quant": "q4_0",
  "target_tps": 10,
  "kv_quant": "q8_0",
  "ram_gb": 48,
  "vram_gb": 12,
  "cpu_cores": 8
}

The response is the raw PlanEstimate JSON from llmfit_core::plan.

llmfit HTTP API Response Shapes

Every successful response is wrapped in a standard envelope:

{
  "node": { "name": "...", "os": "..." },
  "system": { ... },
  "total_models": 123,
  "returned_models": 20,
  "filters": { ... },
  "models": [
    {
      "name": "...",
      "provider": "...",
      "fit_level": "...",
      "runtime": "...",
      "score": ...,
      "throughput_tps": ...,
      "memory_gb": ...,
      ...
    }
  ]
}

Special responses are used for the simpler endpoints:

  • /health → { "status": "ok", "node": {...} }
  • /runtimes → { "runtimes": [ { "name": "ollama", "installed": true } ], "warnings": [] }
  • /installed → { "models": [ { "name": "llama-2-7b", "runtime": "ollama" } ], "warnings": [] }
  • Download start → { "id": "dl-0", "model": "...", "runtime": "...", "status": "pulling" }

Practical Code Examples for the llmfit HTTP API

Here are seven working Rust snippets that demonstrate typical client usage. All assume the server is listening on http://localhost:8080.

1. Health Check

let resp = reqwest::get("http://localhost:8080/health")
    .await?.json::<serde_json::Value>().await?;
println!("{:#}", resp);

2. Query System Specs with Hardware Overrides

let url = "http://localhost:8080/api/v1/system?ram_gb=64&vram_gb=24&cpu_cores=16";
let sys = reqwest::get(url).await?.json::<serde_json::Value>().await?;
println!("System: {}", sys["system"]);

3. List Top 5 Coding Models (Good Fit)

let url = "http://localhost:8080/api/v1/models/top?use_case=coding&min_fit=good";
let list = reqwest::get(url).await?.json::<serde_json::Value>().await?;
for m in list["models"].as_array().unwrap() {
    println!("{} – score {}", m["name"], m["score"]);
}

4. Retrieve a Specific Model

let name = "llama-2-7b";
let url = format!("http://localhost:8080/api/v1/models/{}", name);
let model = reqwest::get(&url).await?.json::<serde_json::Value>().await?;
println!("Details for {}: {:#?}", name, model["models"]);

5. Discover Installed Runtimes

let runtimes = reqwest::get("http://localhost:8080/api/v1/runtimes")
    .await?.json::<serde_json::Value>().await?;
println!("Available runtimes: {:?}", runtimes["runtimes"]);

6. Start a Model Download and Poll Status

let client = reqwest::Client::new();
let body = serde_json::json!({ "model": "llama-2-7b", "runtime": "ollama" });
let start = client.post("http://localhost:8080/api/v1/download")
    .json(&body).send().await?.json::<serde_json::Value>().await?;
let id = start["id"].as_str().unwrap();

let status_url = format!("http://localhost:8080/api/v1/download/{}/status", id);
let status = reqwest::get(&status_url).await?.json::<serde_json::Value>().await?;
println!("Download status: {}", status["status"]);

7. Estimate a Plan for a Model

let plan_body = serde_json::json!({
    "model": "llama-2-7b",
    "context": 8192,
    "quant": "q4_0",
    "ram_gb": 48,
    "vram_gb": 12,
    "cpu_cores": 8
});
let estimate = client.post("http://localhost:8080/api/v1/plan")
    .json(&plan_body).send().await?.json::<serde_json::Value>().await?;
println!("Estimated TPS: {}", estimate["tps"]);

Key Source Files Behind the llmfit HTTP API

File Purpose
llmfit-tui/src/serve_api.rs Implements the HTTP server, routes, request parsing, and response formatting
llmfit-tui/src/serve_shared.rs Helpers like system_json and fit_to_json that serialize core structures
llmfit-core/src/models.rs Defines the LlmModel metadata used by the API
llmfit-core/src/fit.rs Contains ModelFit, FitLevel, and ranking logic
llmfit-core/src/providers.rs Implements runtime providers queried by /runtimes and /installed
llmfit-core/src/plan.rs Performs the planning calculation for /plan

Summary

  • The llmfit HTTP API is an Axum-based REST server defined in llmfit-tui/src/serve_api.rs, started by run_serve and routed by build_router.
  • Endpoints live under /api/v1/ and cover system info, model listing/filtering, top picks, install discoveries, runtime detection, download control, and plan estimation.
  • All responses are wrapped in a consistent ApiEnvelope JSON structure, and errors use the ApiError type with proper HTTP status codes.
  • The download endpoint is localhost-only, so remote access is deliberately restricted for security.
  • The /plan endpoint performs full resource-aware estimation using llmfit_core::plan, including memory usage and throughput projections.

Frequently Asked Questions

Is the llmfit HTTP API available over the network?

No. Only the download endpoint (POST /api/v1/download) is explicitly restricted to localhost, but the API is fundamentally designed for local tooling. You can bind it to a TCP socket or Unix socket, but remote-exposing it is not a use the source code encourages.

What runtimes does the llmfit HTTP API detect?

The /api/v1/runtimes endpoint detects, in no particular order: Ollama, MLX, llama.cpp, Docker, LM Studio, and vLLM. Each one is returned with a boolean installed flag.

How do I filter models with the llmfit API?

Use the /api/v1/models endpoint with query parameters such as min_fit=good, runtime=vllm, and sort=score. For a single model, query /api/v1/models/{name}.

Can I override my hardware specs in the API response?

Yes. The /api/v1/system and /api/v1/models endpoints both accept ram_gb, vram_gb, and cpu_cores query parameters, which override the detected hardware before any fit or plan calculation runs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →