How the llmfit-tui MCP Server Exposes Hardware, Model, Runtime, and Planning Tools Over stdio

The llmfit-tui MCP server turns a set of annotated Rust functions into JSON-encoded remote procedure calls that stream over standard input/output, allowing any language or script to query hardware specs, recommend models, discover runtimes, and calculate hardware plans.

The Model Control Protocol (MCP) implementation in the AlexsJones/llmfit repository provides an asynchronous, stdio-based RPC layer that exposes the core llmfit analysis engine to external clients. By leveraging the rmcp crate and procedural macros, the server transforms local hardware introspection and model planning capabilities into a language-agnostic interface that requires no network configuration—only piped text streams.

Architecture of the stdio MCP Layer

The server architecture centers on a declarative approach where Rust methods become remotely callable endpoints through compile-time code generation.

The LlmfitMcpServer Struct and ToolRouter

At the heart of the implementation lies the LlmfitMcpServer struct defined in llmfit-tui/src/mcp_server.rs. This struct encapsulates the server's state, including:

  • Detected SystemSpecs (hardware inventory)
  • The loaded model catalog
  • Optional calculation configuration profiles
  • A unique node identifier
  • The ToolRouter that dispatches incoming RPC names to handler methods

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L72-L82)

Declarative Tool Definitions with #[tool]

Each capability exposed over MCP is defined as an async method annotated with the #[tool(name = "...", description = "...")] macro. This macro generates the serialization glue that:

  1. Deserializes JSON arguments from the transport
  2. Executes the method
  3. Serializes the return value back to JSON

The #[tool_router] macro then implements a tool_router() method that aggregates all annotated tools into a ToolRouter<Self>, which the MCP service uses for request dispatch.

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L87-L92)

stdio Transport Implementation

The transport layer uses rmcp::transport::io::stdio(), which creates a bidirectional channel reading line-delimited JSON requests from stdin and writing responses to stdout. This design makes the server immediately accessible from shell pipelines, Python scripts, or any process that can spawn the llmfit binary.

In run_mcp_server() (lines 75–105), the entry point constructs a Tokio runtime, instantiates LlmfitMcpServer, and binds it to the stdio transport. The serve method, provided by the ServerHandler trait (automatically implemented via #[tool_handler]), enters the async request-processing loop.

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L75-L105)

Core MCP Tools and Capabilities

The server exposes five primary tool categories that mirror the CLI's functionality, each implemented as a distinct method on LlmfitMcpServer.

Hardware Discovery via get_system_specs

The get_system_specs tool marshals hardware data detected by the sysinfo crate into a JSON payload. It utilizes the helper function serve_shared::system_json to translate the SystemSpecs struct into a serializable format containing RAM, GPU VRAM, CPU count, and platform identifiers.

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L106-L118)

Model Recommendations with recommend_models

The recommend_models endpoint filters and ranks available models based on the host's hardware constraints. It invokes self.analyze_all(), which internally calls llmfit_core::analysis::rankable_models, then applies user-supplied RecommendModelsParams (limit, use-case, minimum fit quality). The response includes a total count, the number of results returned, and an array of model fit objects.

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L119-L164)

Runtime and Installed Model Discovery

Two tools expose the state of local inference backends:

  1. get_runtimes – Spawns blocking detection tasks for each supported provider (Ollama, MLX, llama.cpp, Docker, LM Studio, vLLM, RamaLama) and returns a JSON list of available runtimes.

  2. get_installed_models – Queries each detected runtime for locally cached models, aggregating results into a unified JSON list.

Sources: get_runtimes L35-L80 and get_installed_models L82-L34

Planning with plan_hardware

The plan_hardware tool executes the planning algorithm for a specific model configuration. It:

  1. Looks up the requested model by identifier
  2. Constructs a PlanRequest with the specified context length and batch size
  3. Selects the calculation config (from --profile overrides or defaults)
  4. Calls estimate_model_plan_with_config from llmfit-core/src/plan.rs
  5. Returns the raw estimate struct as JSON

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L200-L233)

Safety and Filtering Logic

The underlying analyze_all() method performs platform-specific filtering, automatically removing MLX-only models when running on non-Apple Silicon hardware. It also respects bandwidth overrides from calculation profiles, ensuring the MCP-exposed data remains consistent with the native CLI output.

Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L38-L57)

Connecting to the MCP Server

You can interact with the llmfit-tui MCP server using the built-in CLI flag, generic MCP clients, or custom scripts.

Starting the Server


# Launch the stdio MCP server

llmfit --mcp

The process now reads JSON-RPC requests from stdin and writes responses to stdout until terminated.

Rust Client Example

use rmcp::client::Client;
use rmcp::transport::io::stdio;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Connect to the running server over stdio
    let client = Client::new(stdio()).await?;
    
    // Query hardware specs
    let specs = client
        .call("get_system_specs", serde_json::json!({}))
        .await?;
    println!("Hardware: {}", specs);
    
    // Get model recommendations
    let rec = client
        .call(
            "recommend_models",
            serde_json::json!({ "limit": 5, "use_case": "coding" }),
        )
        .await?;
    println!("Recommendations: {}", rec);
    
    // Plan hardware for a specific model
    let plan = client
        .call(
            "plan_hardware",
            serde_json::json!({ "model": "llama-2-7b", "context": 4096 }),
        )
        .await?;
    println!("Plan: {}", plan);
    
    Ok(())
}

Python Client Example

import json
import subprocess
import sys

proc = subprocess.Popen(
    ["llmfit", "--mcp"],
    stdin=subprocess.PIPE,
    stdout=subprocess.PIPE,
    text=True,
)

def call(method, params):
    req = {"method": method, "params": params}
    proc.stdin.write(json.dumps(req) + "\n")
    proc.stdin.flush()
    resp = proc.stdout.readline()
    return json.loads(resp)

# Discover available runtimes

print(call("get_runtimes", {}))

# Recommend up to 3 models with good fit or better

print(call("recommend_models", {"limit": 3, "min_fit": "good"}))

Shell Pipeline Example


# Using rmcp-cli to decode responses

echo '{"method":"get_system_specs","params":{}}' | llmfit --mcp | rmcp-cli --decode

Summary

  • The llmfit-tui MCP server exposes hardware introspection, model recommendations, runtime discovery, and planning tools through a stdio-based JSON-RPC interface.
  • The LlmfitMcpServer struct manages state and routes requests via the ToolRouter, generated by the #[tool_router] macro.
  • Methods annotated with #[tool] automatically handle JSON serialization, making Rust functions remotely callable without boilerplate.
  • The rmcp::transport::io::stdio() transport enables immediate integration with shell pipelines, Python scripts, and other language runtimes.
  • Core capabilities include get_system_specs, recommend_models, get_runtimes, get_installed_models, and plan_hardware, each backed by the llmfit-core analysis engine.

Frequently Asked Questions

What transport protocol does the llmfit-tui MCP server use?

The server uses standard input/output (stdio) as its transport layer. It reads line-delimited JSON-RPC requests from stdin and writes JSON responses to stdout. This approach, implemented via rmcp::transport::io::stdio(), allows the server to operate without network ports, making it ideal for local tool composition and sandboxed environments.

How are Rust methods converted into MCP-callable tools?

The conversion uses procedural macros. The #[tool(name = "...", description = "...")] attribute marks a method as an MCP endpoint, while #[tool_router] generates the routing table. At runtime, the ToolRouter deserializes incoming JSON parameters into the method's arguments, invokes the async function, and serializes the return value back to JSON for the stdio transport.

Can I use the MCP server with languages other than Rust?

Yes. Because the protocol is JSON-RPC over text streams, any language that can spawn subprocesses and read/write lines can communicate with the server. The repository includes examples in Python and Bash, and the rmcp ecosystem provides clients for multiple languages that handle the stdio plumbing automatically.

Where is the hardware detection logic implemented?

Hardware detection relies on the sysinfo crate, with the MCP-specific serialization located in llmfit-tui/src/serve_shared.rs. The SystemSpecs struct is populated during server initialization and exposed through the get_system_specs tool, ensuring clients receive consistent hardware data that accounts for platform-specific limitations (such as MLX compatibility on Apple Silicon).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →