What Does the FreeLLMAPI Model Context Protocol (MCP) Server Do?
The FreeLLMAPI Model Context Protocol (MCP) server is a stateless JSON‑RPC gateway that lets AI agents query runtime routing information—available models, provider health, and cache statistics—without sending an inference request.
The FreeLLMAPI project provides a unified, OpenAI‑compatible router for free LLM providers. Its built‑in MCP server extends this gateway with machine‑readable introspection capabilities, enabling autonomous agents to make dynamic routing decisions and self‑diagnose infrastructure state.
MCP Server Architecture and Design
The MCP implementation in server/src/routes/mcp.ts follows a deliberately minimal, stateless design. Unlike persistent MCP transports such as Server‑Sent Events or stdio, FreeLLMAPI exposes a simple HTTP POST endpoint at /mrcp that accepts single JSON‑RPC requests and returns immediate responses.
# mcp.ts lines 13-21: core router structure
# Only five methods are implemented; no session state is maintained
This architecture eliminates connection management overhead. The server rejects batched JSON‑RPC calls with HTTP 405 and always reports protocol version 2025-06-18 in its capabilities.
Authentication
Authentication mirrors the standard /v1 OpenAI‑compatible endpoints. A unified API key is accepted via:
Authorization: Bearer YOUR_UNIFIED_KEYx-api-key: YOUR_UNIFIED_KEY
// mcp.ts lines 29-30: unified auth middleware
This ensures agents can reuse existing credentials without separate MCP‑specific configuration.
Five MCP Methods for Agent Introspection
listModels
Returns a catalog of free models currently available, including context windows, supported platforms, parameters, and a special "auto" entry that delegates model selection to the router.
Core implementation: mcp.ts lines 68‑84, delegating to server/src/services/model-listing.ts
curl -X POST http://localhost:3001/mcp \
-H "Authorization: Bearer YOUR_UNIFIED_KEY" \
-d '{"jsonrpc":"2.0","id":1,"method":"listModels","params":{}}'
providerHealth
Aggregates health status across configured providers: enabled API keys, rate‑limit cooldown states, and availability flags.
Core implementation: mcp.ts lines 87‑100
curl -X POST http://localhost:3001/mcp \
-H "Authorization: Bearer YOUR_UNIFIED_KEY" \
-d '{"jsonrpc":"2.0","id":2,"method":"providerHealth","params":{}}'
routingStrategy
Exposes the active routing strategy (e.g., fastest, balanced) and supports on‑the‑fly changes via setRoutingStrategy.
Core implementation: mcp.ts lines 8‑9, strategy logic in server/src/services/router.ts
curl -X POST http://localhost:3001/mcp \
-H "Authorization: Bearer YOUR_UNIFIED_KEY" \
-d '{"jsonrpc":"2.0","id":3,"method":"setRoutingStrategy","params":{"strategy":"balanced"}}'
cacheStats
Reports live request‑caching metrics: hit/miss counts and estimated tokens saved.
Data source: server/src/services/cache.ts
// Referenced at mcp.ts line 10
compressionStats
Provides request‑body compression statistics: invocation counts, processing timings, and bandwidth savings estimates.
Data source: server/src/services/compression/stats.ts
// Referenced at mcp.ts line 10
Response Format and Tool Integration
All MCP methods return JSON‑RPC 2.0 responses with a structured content field containing pretty‑printed JSON. The server uses a toolJson helper to ensure consistent formatting suitable for agent consumption.
// mcp.ts lines 56-60: toolJson helper for uniform response shaping
This design aligns with how AI coding agents—such as Claude Code, Cursor, or Cline—expect tool outputs: self‑contained, machine‑parseable JSON that can drive subsequent decision logic.
Key Source Files
| File | Purpose |
|---|---|
server/src/routes/mcp.ts |
Core stateless MCP JSON‑RPC router and method handlers |
server/src/services/model-listing.ts |
Model catalog generation for listModels |
server/src/services/router.ts |
Routing strategy implementation |
server/src/services/cache.ts |
Cache statistics provider |
server/src/services/compression/stats.ts |
Compression metrics provider |
docs/clients.md |
MCP server documentation for agents and developers |
Summary
- FreeLLMAPI MCP server adds stateless introspection to an LLM routing gateway, implemented in
server/src/routes/mcp.ts - Five JSON‑RPC methods expose models, provider health, routing strategy, cache stats, and compression stats
- Unified Bearer token auth reuses existing
/v1endpoint credentials - Single‑request HTTP transport rejects batches, reports protocol version
2025-06-18 - Agent‑optimized responses via
toolJsonhelper enable dynamic, autonomous routing decisions
Frequently Asked Questions
What is Model Context Protocol (MCP) in FreeLLMAPI?
FreeLLMAPI's MCP server is a lightweight JSON‑RPC interface that lets AI agents query infrastructure state—what models are available, which providers are healthy—without performing actual LLM inference. It implements a subset of the MCP specification using plain HTTP requests rather than persistent connections.
How do I authenticate with the FreeLLMAPI MCP server?
Use the same unified API key required by the /v1/chat/completions endpoint. Pass it as Authorization: Bearer YOUR_UNIFIED_KEY or as x-api-key: YOUR_UNIFIED_KEY. The MCP router reuses the authentication middleware from the main API surface.
Can I change routing strategy through MCP?
Yes. Call setRoutingStrategy with parameters like {"strategy": "balanced"} or {"strategy": "fastest"}. The change takes effect immediately for subsequent inference requests. The implementation delegates to server/src/services/router.ts.
Why does FreeLLMAPI MCP reject batched requests?
The server enforces strict statelessness. Batched JSON‑RPC arrays return HTTP 405 because the implementation handles only single request/response cycles. Agents should issue individual calls to /mcp for each introspection query.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →