# What Is the llmfit HTTP API? A Complete Endpoint Reference

> Explore the llmfit HTTP API to query hardware, find LLM models, download them, and calculate execution plans. This Axum-based REST interface empowers local clients with full control over your LLM environment.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: api-reference
- Published: 2026-08-23

---

**The llmfit HTTP API is a REST-style JSON web interface built with the Axum framework that lets you query hardware specs, discover compatible LLM models, trigger model downloads, and calculate resource-aware execution plans from any local client.**

The [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) project bundles a lightweight, self-contained web server inside the `llmfit-tui` crate. It is designed so external tools, scripts, or a browser can interact with llmfit's core logic without needing to touch the TUI. This article walks through the llmfit HTTP API surface, every endpoint, its parameters, response formats, and working code examples, all grounded directly in the source code.

## Understanding the llmfit HTTP API Architecture

The server entry point lives in [`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs). The `run_serve` function starts the server, `build_router` registers all routes, and the server binds to either a TCP socket or a Unix domain socket. All handlers share a common `AppState` (defined at lines 31-40) that holds:

- Node name and OS metadata
- Detected system specifications
- The full model list
- Optional context limit
- Download-tracking state

The router registers every API endpoint under the `/api/v1/` namespace alongside UI fallback routes. Handlers return JSON wrapped in an `ApiEnvelope` that includes node info, system specs, pagination data, active filters, and an array of model objects. Errors are standardized through the `ApiError` type, which serializes to a JSON payload with a proper HTTP status code.

## All llmfit HTTP API Endpoints and Their Parameters

The llmfit HTTP API exposes eight primary endpoints. Here's the full break down.

### GET /health — Liveness Check

A minimal availability probe. Returns `{ "status": "ok", "node": {...} }` from the `health` handler (lines 66-74 of [`serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/serve_api.rs)).

### GET /api/v1/system — Hardware and Node Information

Reports node metadata and detected hardware specs. Supports `ram_gb`, `vram_gb`, and `cpu_cores` as optional query parameters to override auto-detected values.

```

GET /api/v1/system?ram_gb=64&vram_gb=24&cpu_cores=16

```

### GET /api/v1/models — List, Filter, and Sort Models

The workhorse endpoint. It accepts a rich set of query parameters:

| Parameter | Meaning |
|-----------|---------|
| `limit` | Page size for returned models |
| `top` | Number of top results to consider |
| `perfect` | Filter to only models with perfect fit |
| `min_fit` | Minimum fit level (e.g., `good`) |
| `runtime` | Target runtime (e.g., `vllm`) |
| `use_case` | Use case filter (e.g., `coding`) |
| `provider` | Model provider |
| `search` | Text search across model names |
| `sort` | Sort key (e.g., `score`) |
| `include_too_tight` | Include models that exceed memory budget |
| `max_context` | Maximum context window |
| `force_runtime` | Force a particular runtime |
| `license` | License filter |
| `ram_gb`, `vram_gb`, `cpu_cores` | Hardware overrides |

Example:

```

GET /api/v1/models?limit=20&min_fit=good&sort=score&runtime=vllm

```

### GET /api/v1/models/top — Best-Fit Models Only

Same parameters as `/models`, but returns only the best-fitting models. Default `top=5`.

### GET /api/v1/models/{name} — Retrieve a Single Model

Inherits all `ModelsQuery` parameters; the `search` parameter is effectively forced to the supplied name.

```

GET /api/v1/models/llama-2-7b

```

### GET /api/v1/runtimes — Discover Installed Runtimes

Returns a list of runtime names (Ollama, MLX, llama.cpp, Docker, LM Studio, vLLM) with a boolean `installed` flag for each.

### GET /api/v1/installed — List Installed Models

Returns a list of `{name, runtime}` objects for models already present on the host across every detected runtime.

### POST /api/v1/download — Start a Model Pull

Starts a model download. Only reachable from localhost for safety. The JSON body is `{ model: string, runtime: string }`. Returns a download ID for polling.

### GET /api/v1/download/{id}/status — Poll Download Progress

Returns the current pull status, including a `progress_pct` field and a human-readable message.

### POST /api/v1/plan — Estimate an Execution Plan

Computes a resource-aware plan for running a given model at a given context size. The JSON body accepts:

```json
{
  "model": "llama-2-7b",
  "context": 8192,
  "quant": "q4_0",
  "target_tps": 10,
  "kv_quant": "q8_0",
  "ram_gb": 48,
  "vram_gb": 12,
  "cpu_cores": 8
}

```

The response is the raw `PlanEstimate` JSON from `llmfit_core::plan`.

## llmfit HTTP API Response Shapes

Every successful response is wrapped in a standard envelope:

```json
{
  "node": { "name": "...", "os": "..." },
  "system": { ... },
  "total_models": 123,
  "returned_models": 20,
  "filters": { ... },
  "models": [
    {
      "name": "...",
      "provider": "...",
      "fit_level": "...",
      "runtime": "...",
      "score": ...,
      "throughput_tps": ...,
      "memory_gb": ...,
      ...
    }
  ]
}

```

Special responses are used for the simpler endpoints:

- `/health` → `{ "status": "ok", "node": {...} }`
- `/runtimes` → `{ "runtimes": [ { "name": "ollama", "installed": true } ], "warnings": [] }`
- `/installed` → `{ "models": [ { "name": "llama-2-7b", "runtime": "ollama" } ], "warnings": [] }`
- Download start → `{ "id": "dl-0", "model": "...", "runtime": "...", "status": "pulling" }`

## Practical Code Examples for the llmfit HTTP API

Here are seven working Rust snippets that demonstrate typical client usage. All assume the server is listening on `http://localhost:8080`.

### 1. Health Check

```rust
let resp = reqwest::get("http://localhost:8080/health")
    .await?.json::<serde_json::Value>().await?;
println!("{:#}", resp);

```

### 2. Query System Specs with Hardware Overrides

```rust
let url = "http://localhost:8080/api/v1/system?ram_gb=64&vram_gb=24&cpu_cores=16";
let sys = reqwest::get(url).await?.json::<serde_json::Value>().await?;
println!("System: {}", sys["system"]);

```

### 3. List Top 5 Coding Models (Good Fit)

```rust
let url = "http://localhost:8080/api/v1/models/top?use_case=coding&min_fit=good";
let list = reqwest::get(url).await?.json::<serde_json::Value>().await?;
for m in list["models"].as_array().unwrap() {
    println!("{} – score {}", m["name"], m["score"]);
}

```

### 4. Retrieve a Specific Model

```rust
let name = "llama-2-7b";
let url = format!("http://localhost:8080/api/v1/models/{}", name);
let model = reqwest::get(&url).await?.json::<serde_json::Value>().await?;
println!("Details for {}: {:#?}", name, model["models"]);

```

### 5. Discover Installed Runtimes

```rust
let runtimes = reqwest::get("http://localhost:8080/api/v1/runtimes")
    .await?.json::<serde_json::Value>().await?;
println!("Available runtimes: {:?}", runtimes["runtimes"]);

```

### 6. Start a Model Download and Poll Status

```rust
let client = reqwest::Client::new();
let body = serde_json::json!({ "model": "llama-2-7b", "runtime": "ollama" });
let start = client.post("http://localhost:8080/api/v1/download")
    .json(&body).send().await?.json::<serde_json::Value>().await?;
let id = start["id"].as_str().unwrap();

let status_url = format!("http://localhost:8080/api/v1/download/{}/status", id);
let status = reqwest::get(&status_url).await?.json::<serde_json::Value>().await?;
println!("Download status: {}", status["status"]);

```

### 7. Estimate a Plan for a Model

```rust
let plan_body = serde_json::json!({
    "model": "llama-2-7b",
    "context": 8192,
    "quant": "q4_0",
    "ram_gb": 48,
    "vram_gb": 12,
    "cpu_cores": 8
});
let estimate = client.post("http://localhost:8080/api/v1/plan")
    .json(&plan_body).send().await?.json::<serde_json::Value>().await?;
println!("Estimated TPS: {}", estimate["tps"]);

```

## Key Source Files Behind the llmfit HTTP API

| File | Purpose |
|------|---------|
| [`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs) | Implements the HTTP server, routes, request parsing, and response formatting |
| [`llmfit-tui/src/serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs) | Helpers like `system_json` and `fit_to_json` that serialize core structures |
| [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) | Defines the `LlmModel` metadata used by the API |
| [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) | Contains `ModelFit`, `FitLevel`, and ranking logic |
| [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) | Implements runtime providers queried by `/runtimes` and `/installed` |
| [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) | Performs the planning calculation for `/plan` |

## Summary

- The **llmfit HTTP API** is an Axum-based REST server defined in [`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs), started by `run_serve` and routed by `build_router`.
- Endpoints live under `/api/v1/` and cover **system info**, **model listing/filtering**, **top picks**, **install discoveries**, **runtime detection**, **download control**, and **plan estimation**.
- All responses are wrapped in a consistent `ApiEnvelope` JSON structure, and errors use the `ApiError` type with proper HTTP status codes.
- The download endpoint is **localhost-only**, so remote access is deliberately restricted for security.
- The `/plan` endpoint performs full resource-aware estimation using `llmfit_core::plan`, including memory usage and throughput projections.

## Frequently Asked Questions

### Is the llmfit HTTP API available over the network?

No. Only the download endpoint (`POST /api/v1/download`) is explicitly restricted to localhost, but the API is fundamentally designed for local tooling. You can bind it to a TCP socket or Unix socket, but remote-exposing it is not a use the source code encourages.

### What runtimes does the llmfit HTTP API detect?

The `/api/v1/runtimes` endpoint detects, in no particular order: **Ollama, MLX, llama.cpp, Docker, LM Studio, and vLLM**. Each one is returned with a boolean `installed` flag.

### How do I filter models with the llmfit API?

Use the `/api/v1/models` endpoint with query parameters such as `min_fit=good`, `runtime=vllm`, and `sort=score`. For a single model, query `/api/v1/models/{name}`.

### Can I override my hardware specs in the API response?

Yes. The `/api/v1/system` and `/api/v1/models` endpoints both accept `ram_gb`, `vram_gb`, and `cpu_cores` query parameters, which override the detected hardware before any fit or plan calculation runs.