How the llmfit-tui MCP Server Exposes Hardware, Model, Runtime, and Planning Tools Over stdio
The llmfit-tui MCP server turns a set of annotated Rust functions into JSON-encoded remote procedure calls that stream over standard input/output, allowing any language or script to query hardware specs, recommend models, discover runtimes, and calculate hardware plans.
The Model Control Protocol (MCP) implementation in the AlexsJones/llmfit repository provides an asynchronous, stdio-based RPC layer that exposes the core llmfit analysis engine to external clients. By leveraging the rmcp crate and procedural macros, the server transforms local hardware introspection and model planning capabilities into a language-agnostic interface that requires no network configuration—only piped text streams.
Architecture of the stdio MCP Layer
The server architecture centers on a declarative approach where Rust methods become remotely callable endpoints through compile-time code generation.
The LlmfitMcpServer Struct and ToolRouter
At the heart of the implementation lies the LlmfitMcpServer struct defined in llmfit-tui/src/mcp_server.rs. This struct encapsulates the server's state, including:
- Detected
SystemSpecs(hardware inventory) - The loaded model catalog
- Optional calculation configuration profiles
- A unique node identifier
- The
ToolRouterthat dispatches incoming RPC names to handler methods
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L72-L82)
Declarative Tool Definitions with #[tool]
Each capability exposed over MCP is defined as an async method annotated with the #[tool(name = "...", description = "...")] macro. This macro generates the serialization glue that:
- Deserializes JSON arguments from the transport
- Executes the method
- Serializes the return value back to JSON
The #[tool_router] macro then implements a tool_router() method that aggregates all annotated tools into a ToolRouter<Self>, which the MCP service uses for request dispatch.
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L87-L92)
stdio Transport Implementation
The transport layer uses rmcp::transport::io::stdio(), which creates a bidirectional channel reading line-delimited JSON requests from stdin and writing responses to stdout. This design makes the server immediately accessible from shell pipelines, Python scripts, or any process that can spawn the llmfit binary.
In run_mcp_server() (lines 75–105), the entry point constructs a Tokio runtime, instantiates LlmfitMcpServer, and binds it to the stdio transport. The serve method, provided by the ServerHandler trait (automatically implemented via #[tool_handler]), enters the async request-processing loop.
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L75-L105)
Core MCP Tools and Capabilities
The server exposes five primary tool categories that mirror the CLI's functionality, each implemented as a distinct method on LlmfitMcpServer.
Hardware Discovery via get_system_specs
The get_system_specs tool marshals hardware data detected by the sysinfo crate into a JSON payload. It utilizes the helper function serve_shared::system_json to translate the SystemSpecs struct into a serializable format containing RAM, GPU VRAM, CPU count, and platform identifiers.
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L106-L118)
Model Recommendations with recommend_models
The recommend_models endpoint filters and ranks available models based on the host's hardware constraints. It invokes self.analyze_all(), which internally calls llmfit_core::analysis::rankable_models, then applies user-supplied RecommendModelsParams (limit, use-case, minimum fit quality). The response includes a total count, the number of results returned, and an array of model fit objects.
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L119-L164)
Runtime and Installed Model Discovery
Two tools expose the state of local inference backends:
-
get_runtimes– Spawns blocking detection tasks for each supported provider (Ollama, MLX, llama.cpp, Docker, LM Studio, vLLM, RamaLama) and returns a JSON list of available runtimes. -
get_installed_models– Queries each detected runtime for locally cached models, aggregating results into a unified JSON list.
Sources: get_runtimes L35-L80 and get_installed_models L82-L34
Planning with plan_hardware
The plan_hardware tool executes the planning algorithm for a specific model configuration. It:
- Looks up the requested model by identifier
- Constructs a
PlanRequestwith the specified context length and batch size - Selects the calculation config (from
--profileoverrides or defaults) - Calls
estimate_model_plan_with_configfromllmfit-core/src/plan.rs - Returns the raw estimate struct as JSON
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L200-L233)
Safety and Filtering Logic
The underlying analyze_all() method performs platform-specific filtering, automatically removing MLX-only models when running on non-Apple Silicon hardware. It also respects bandwidth overrides from calculation profiles, ensuring the MCP-exposed data remains consistent with the native CLI output.
Source: [llmfit-tui/src/mcp_server.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/mcp_server.rs#L38-L57)
Connecting to the MCP Server
You can interact with the llmfit-tui MCP server using the built-in CLI flag, generic MCP clients, or custom scripts.
Starting the Server
# Launch the stdio MCP server
llmfit --mcp
The process now reads JSON-RPC requests from stdin and writes responses to stdout until terminated.
Rust Client Example
use rmcp::client::Client;
use rmcp::transport::io::stdio;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
// Connect to the running server over stdio
let client = Client::new(stdio()).await?;
// Query hardware specs
let specs = client
.call("get_system_specs", serde_json::json!({}))
.await?;
println!("Hardware: {}", specs);
// Get model recommendations
let rec = client
.call(
"recommend_models",
serde_json::json!({ "limit": 5, "use_case": "coding" }),
)
.await?;
println!("Recommendations: {}", rec);
// Plan hardware for a specific model
let plan = client
.call(
"plan_hardware",
serde_json::json!({ "model": "llama-2-7b", "context": 4096 }),
)
.await?;
println!("Plan: {}", plan);
Ok(())
}
Python Client Example
import json
import subprocess
import sys
proc = subprocess.Popen(
["llmfit", "--mcp"],
stdin=subprocess.PIPE,
stdout=subprocess.PIPE,
text=True,
)
def call(method, params):
req = {"method": method, "params": params}
proc.stdin.write(json.dumps(req) + "\n")
proc.stdin.flush()
resp = proc.stdout.readline()
return json.loads(resp)
# Discover available runtimes
print(call("get_runtimes", {}))
# Recommend up to 3 models with good fit or better
print(call("recommend_models", {"limit": 3, "min_fit": "good"}))
Shell Pipeline Example
# Using rmcp-cli to decode responses
echo '{"method":"get_system_specs","params":{}}' | llmfit --mcp | rmcp-cli --decode
Summary
- The llmfit-tui MCP server exposes hardware introspection, model recommendations, runtime discovery, and planning tools through a stdio-based JSON-RPC interface.
- The
LlmfitMcpServerstruct manages state and routes requests via theToolRouter, generated by the#[tool_router]macro. - Methods annotated with
#[tool]automatically handle JSON serialization, making Rust functions remotely callable without boilerplate. - The
rmcp::transport::io::stdio()transport enables immediate integration with shell pipelines, Python scripts, and other language runtimes. - Core capabilities include
get_system_specs,recommend_models,get_runtimes,get_installed_models, andplan_hardware, each backed by thellmfit-coreanalysis engine.
Frequently Asked Questions
What transport protocol does the llmfit-tui MCP server use?
The server uses standard input/output (stdio) as its transport layer. It reads line-delimited JSON-RPC requests from stdin and writes JSON responses to stdout. This approach, implemented via rmcp::transport::io::stdio(), allows the server to operate without network ports, making it ideal for local tool composition and sandboxed environments.
How are Rust methods converted into MCP-callable tools?
The conversion uses procedural macros. The #[tool(name = "...", description = "...")] attribute marks a method as an MCP endpoint, while #[tool_router] generates the routing table. At runtime, the ToolRouter deserializes incoming JSON parameters into the method's arguments, invokes the async function, and serializes the return value back to JSON for the stdio transport.
Can I use the MCP server with languages other than Rust?
Yes. Because the protocol is JSON-RPC over text streams, any language that can spawn subprocesses and read/write lines can communicate with the server. The repository includes examples in Python and Bash, and the rmcp ecosystem provides clients for multiple languages that handle the stdio plumbing automatically.
Where is the hardware detection logic implemented?
Hardware detection relies on the sysinfo crate, with the MCP-specific serialization located in llmfit-tui/src/serve_shared.rs. The SystemSpecs struct is populated during server initialization and exposed through the get_system_specs tool, ensuring clients receive consistent hardware data that accounts for platform-specific limitations (such as MLX compatibility on Apple Silicon).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →