# How the llmfit-tui MCP Server Exposes Hardware, Model, Runtime, and Planning Tools Over stdio

> Discover how llmfit-tui mcp_server uses stdio to expose hardware, model, runtime, and planning tools. Access powerful features from any language.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**The llmfit-tui MCP server turns a set of annotated Rust functions into JSON-encoded remote procedure calls that stream over standard input/output, allowing any language or script to query hardware specs, recommend models, discover runtimes, and calculate hardware plans.**

The **Model Control Protocol (MCP)** implementation in the `AlexsJones/llmfit` repository provides an asynchronous, stdio-based RPC layer that exposes the core `llmfit` analysis engine to external clients. By leveraging the `rmcp` crate and procedural macros, the server transforms local hardware introspection and model planning capabilities into a language-agnostic interface that requires no network configuration—only piped text streams.


## Architecture of the stdio MCP Layer

The server architecture centers on a declarative approach where Rust methods become remotely callable endpoints through compile-time code generation.

### The LlmfitMcpServer Struct and ToolRouter

At the heart of the implementation lies the **`LlmfitMcpServer`** struct defined in [`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs). This struct encapsulates the server's state, including:

- Detected **`SystemSpecs`** (hardware inventory)
- The loaded model catalog
- Optional calculation configuration profiles
- A unique node identifier
- The **`ToolRouter`** that dispatches incoming RPC names to handler methods

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L72-L82)

### Declarative Tool Definitions with #[tool]

Each capability exposed over MCP is defined as an async method annotated with the **`#[tool(name = "...", description = "...")]`** macro. This macro generates the serialization glue that:

1. Deserializes JSON arguments from the transport
2. Executes the method
3. Serializes the return value back to JSON

The **`#[tool_router]`** macro then implements a `tool_router()` method that aggregates all annotated tools into a `ToolRouter<Self>`, which the MCP service uses for request dispatch.

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L87-L92)

### stdio Transport Implementation

The transport layer uses **`rmcp::transport::io::stdio()`**, which creates a bidirectional channel reading line-delimited JSON requests from stdin and writing responses to stdout. This design makes the server immediately accessible from shell pipelines, Python scripts, or any process that can spawn the `llmfit` binary.

In `run_mcp_server()` (lines 75–105), the entry point constructs a Tokio runtime, instantiates `LlmfitMcpServer`, and binds it to the stdio transport. The `serve` method, provided by the `ServerHandler` trait (automatically implemented via `#[tool_handler]`), enters the async request-processing loop.

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L75-L105)


## Core MCP Tools and Capabilities

The server exposes five primary tool categories that mirror the CLI's functionality, each implemented as a distinct method on `LlmfitMcpServer`.

### Hardware Discovery via get_system_specs

The **`get_system_specs`** tool marshals hardware data detected by the `sysinfo` crate into a JSON payload. It utilizes the helper function `serve_shared::system_json` to translate the `SystemSpecs` struct into a serializable format containing RAM, GPU VRAM, CPU count, and platform identifiers.

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L106-L118)

### Model Recommendations with recommend_models

The **`recommend_models`** endpoint filters and ranks available models based on the host's hardware constraints. It invokes `self.analyze_all()`, which internally calls `llmfit_core::analysis::rankable_models`, then applies user-supplied `RecommendModelsParams` (limit, use-case, minimum fit quality). The response includes a total count, the number of results returned, and an array of model fit objects.

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L119-L164)

### Runtime and Installed Model Discovery

Two tools expose the state of local inference backends:

1. **`get_runtimes`** – Spawns blocking detection tasks for each supported provider (Ollama, MLX, llama.cpp, Docker, LM Studio, vLLM, RamaLama) and returns a JSON list of available runtimes.

2. **`get_installed_models`** – Queries each detected runtime for locally cached models, aggregating results into a unified JSON list.

*Sources:* `get_runtimes` [`L35-L80`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L35-L80) and `get_installed_models` [`L82-L34`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L82-L34)

### Planning with plan_hardware

The **`plan_hardware`** tool executes the planning algorithm for a specific model configuration. It:

1. Looks up the requested model by identifier
2. Constructs a `PlanRequest` with the specified context length and batch size
3. Selects the calculation config (from `--profile` overrides or defaults)
4. Calls `estimate_model_plan_with_config` from [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs)
5. Returns the raw estimate struct as JSON

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L200-L233)

### Safety and Filtering Logic

The underlying `analyze_all()` method performs platform-specific filtering, automatically removing **MLX-only** models when running on non-Apple Silicon hardware. It also respects bandwidth overrides from calculation profiles, ensuring the MCP-exposed data remains consistent with the native CLI output.

*Source:* [[`llmfit-tui/src/mcp_server.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L38-L57)


## Connecting to the MCP Server

You can interact with the llmfit-tui MCP server using the built-in CLI flag, generic MCP clients, or custom scripts.

### Starting the Server

```bash

# Launch the stdio MCP server

llmfit --mcp

```

The process now reads JSON-RPC requests from stdin and writes responses to stdout until terminated.

### Rust Client Example

```rust
use rmcp::client::Client;
use rmcp::transport::io::stdio;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Connect to the running server over stdio
    let client = Client::new(stdio()).await?;
    
    // Query hardware specs
    let specs = client
        .call("get_system_specs", serde_json::json!({}))
        .await?;
    println!("Hardware: {}", specs);
    
    // Get model recommendations
    let rec = client
        .call(
            "recommend_models",
            serde_json::json!({ "limit": 5, "use_case": "coding" }),
        )
        .await?;
    println!("Recommendations: {}", rec);
    
    // Plan hardware for a specific model
    let plan = client
        .call(
            "plan_hardware",
            serde_json::json!({ "model": "llama-2-7b", "context": 4096 }),
        )
        .await?;
    println!("Plan: {}", plan);
    
    Ok(())
}

```

### Python Client Example

```python
import json
import subprocess
import sys

proc = subprocess.Popen(
    ["llmfit", "--mcp"],
    stdin=subprocess.PIPE,
    stdout=subprocess.PIPE,
    text=True,
)

def call(method, params):
    req = {"method": method, "params": params}
    proc.stdin.write(json.dumps(req) + "\n")
    proc.stdin.flush()
    resp = proc.stdout.readline()
    return json.loads(resp)

# Discover available runtimes

print(call("get_runtimes", {}))

# Recommend up to 3 models with good fit or better

print(call("recommend_models", {"limit": 3, "min_fit": "good"}))

```

### Shell Pipeline Example

```bash

# Using rmcp-cli to decode responses

echo '{"method":"get_system_specs","params":{}}' | llmfit --mcp | rmcp-cli --decode

```


## Summary

- The **llmfit-tui MCP server** exposes hardware introspection, model recommendations, runtime discovery, and planning tools through a stdio-based JSON-RPC interface.
- The **`LlmfitMcpServer`** struct manages state and routes requests via the **`ToolRouter`**, generated by the `#[tool_router]` macro.
- Methods annotated with **`#[tool]`** automatically handle JSON serialization, making Rust functions remotely callable without boilerplate.
- The **`rmcp::transport::io::stdio()`** transport enables immediate integration with shell pipelines, Python scripts, and other language runtimes.
- Core capabilities include **`get_system_specs`**, **`recommend_models`**, **`get_runtimes`**, **`get_installed_models`**, and **`plan_hardware`**, each backed by the `llmfit-core` analysis engine.


## Frequently Asked Questions

### What transport protocol does the llmfit-tui MCP server use?

The server uses **standard input/output (stdio)** as its transport layer. It reads line-delimited JSON-RPC requests from stdin and writes JSON responses to stdout. This approach, implemented via `rmcp::transport::io::stdio()`, allows the server to operate without network ports, making it ideal for local tool composition and sandboxed environments.

### How are Rust methods converted into MCP-callable tools?

The conversion uses **procedural macros**. The `#[tool(name = "...", description = "...")]` attribute marks a method as an MCP endpoint, while `#[tool_router]` generates the routing table. At runtime, the `ToolRouter` deserializes incoming JSON parameters into the method's arguments, invokes the async function, and serializes the return value back to JSON for the stdio transport.

### Can I use the MCP server with languages other than Rust?

Yes. Because the protocol is JSON-RPC over text streams, any language that can spawn subprocesses and read/write lines can communicate with the server. The repository includes examples in **Python** and **Bash**, and the `rmcp` ecosystem provides clients for multiple languages that handle the stdio plumbing automatically.

### Where is the hardware detection logic implemented?

Hardware detection relies on the `sysinfo` crate, with the MCP-specific serialization located in **[`llmfit-tui/src/serve_shared.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_shared.rs)**. The `SystemSpecs` struct is populated during server initialization and exposed through the `get_system_specs` tool, ensuring clients receive consistent hardware data that accounts for platform-specific limitations (such as MLX compatibility on Apple Silicon).