How the MemPalace closet_llm Module Interfaces with Local LLM Runtimes

The closet_llm module (implemented in mempalace/llm_client.py) provides a unified provider interface that abstracts interactions with local LLM runtimes such as Ollama and OpenAI-compatible servers, featuring automatic local endpoint detection across RFC 1918, Tailscale, and IPv6 unique-local address ranges to enable fully offline operation.

MemPalace is an open-source knowledge management system that prioritizes privacy through local-first architecture. The closet_llm module serves as the abstraction layer between the application's entity refinement pipeline and various large language model runtimes, enabling seamless integration with locally-hosted models via a consistent Python API defined in the core client file.

Provider Architecture in mempalace/llm_client.py

The module centers on a base LLMProvider class that standardizes how MemPalace communicates with language models. All concrete implementations inherit from this base and must provide two core methods: classify() for sending prompts and check_available() for health checks, plus an is_external_service property indicating whether the configured endpoint resides outside the local machine.

Three concrete providers ship with the module:

  • OllamaProvider – Interfaces with Ollama's HTTP API on localhost:11434
  • OpenAICompatProvider – Connects to any OpenAI-compatible local server (vLLM, LM Studio, etc.)
  • AnthropicProvider – Handles cloud-based Anthropic API calls (optional external service)

How Providers Connect to Local Runtimes

Each provider implements specific HTTP protocols to bridge MemPalace with local inference engines, while the LLMResponse object standardizes the return format containing text, model, provider, and raw attributes.

Ollama Integration (Default)

The OllamaProvider communicates via HTTP POST requests to http://localhost:11434/api/chat. The request body includes the model name, combined system and user messages, and optional flags such as format: "json" for structured output and think parameters. For health checks, it queries /api/tags to verify the server is responsive before attempting classification.

OpenAI-Compatible Servers

For self-hosted OpenAI-compatible runtimes like vLLM, the provider constructs URLs ending in /v1/chat/completions. It POSTs JSON containing model, messages, temperature, and response_format: {"type": "json_object"} when JSON mode is enabled. When running locally, no API key is required, though the provider supports adding Authorization: Bearer <key> headers when configured.

Anthropic (External)

While the Anthropic provider targets the hosted API at api.anthropic.com, the module still treats it uniformly, automatically appending JSON formatting instructions to system prompts when json_mode=True.

Local Endpoint Detection Logic

The module implements intelligent locality checking through the _endpoint_is_local() helper function. This utility examines configured endpoints to determine if they resolve to local addresses, enabling the CLI to warn users before contacting external services.

Detected local ranges include:

  • RFC 1918 private IPv4 subnets (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16)
  • Tailscale CGNAT (100.64.0.0/10)
  • IPv6 unique-local addresses (fc00::/7)
  • Localhost variants (localhost, 127.0.0.1, *.local)

This detection powers the is_external_service property, returning False for local runtimes and True for cloud providers.

Factory Pattern and CLI Integration

The get_provider() factory function at the bottom of mempalace/llm_client.py instantiates the appropriate provider class based on the name parameter looked up in the internal PROVIDERS dictionary. It passes through configuration parameters including model, endpoint, api_key, timeout, and provider-specific options like num_ctx for Ollama context sizing.

In mempalace/cli.py, the integration follows this workflow:

  1. Argument parsing – Reads --llm-provider, --llm-model, --llm-endpoint, and --llm-api-key flags
  2. Provider instantiation – Calls get_provider() with the parsed configuration
  3. Health verification – Invokes provider.check_available() during initialization (e.g., mempalace init)
  4. Classification – Executes provider.classify(system_prompt, user_prompt, json_mode=True) during entity refinement pipelines

Practical Implementation Examples

The following examples demonstrate direct usage of the closet_llm module's public API:

Direct Ollama Provider Usage

from mempalace.llm_client import OllamaProvider

ollama = OllamaProvider(model="llama3:8b", endpoint="http://localhost:11434")
ok, msg = ollama.check_available()
if ok:
    resp = ollama.classify(
        system="You are a helpful assistant.",
        user="Extract the people mentioned in this text.",
        json_mode=True,
        think=False,
    )
    print(resp.text)          # JSON string with the classification result

else:
    print("Ollama not reachable:", msg)

Runtime Provider Selection via Factory

from mempalace.llm_client import get_provider

provider = get_provider(
    name="openai-compat",
    model="gpt-4o-mini",
    endpoint="http://localhost:8000",   # a local vLLM server

    api_key=None,                       # no key needed for a local server

)

if provider.check_available()[0]:
    result = provider.classify(
        system="You are a factual extractor.",
        user="List all project names in the snippet.",
        json_mode=True,
    )
    print(result.text)

Checking Endpoint Locality

from mempalace.llm_client import get_provider

provider = get_provider("ollama", model="gemma:2b")
print("Will this call go off‑machine?", provider.is_external_service)

# → False, because the endpoint resolves to localhost

Summary

  • The closet_llm module centralizes LLM interactions in mempalace/llm_client.py, implementing a provider pattern that supports Ollama, OpenAI-compatible servers, and Anthropic through the LLMProvider base class.
  • Local runtime support requires no API keys and operates over HTTP to localhost:11434 (Ollama) or custom ports (vLLM), satisfying MemPalace's zero-API-required principle.
  • Automatic locality detection distinguishes local from external endpoints using IP range validation (RFC 1918, Tailscale CGNAT, IPv6 unique-local), preserving the application's offline-first guarantee.
  • The get_provider() factory enables consistent initialization across the CLI and entity refinement pipelines, while check_available() ensures local runtimes are reachable before inference begins.

Frequently Asked Questions

What file contains the closet_llm module implementation?

The closet_llm module is implemented in mempalace/llm_client.py, which defines the LLMProvider base class, three concrete provider implementations (Ollama, OpenAI-compat, Anthropic), the _endpoint_is_local() detection logic, and the get_provider() factory function.

How does the module determine if an LLM endpoint is running locally?

The _endpoint_is_local() helper function analyzes the configured endpoint URL against private IP ranges including RFC 1918 (10/8, 172.16/12, 192.168/16), Tailscale CGNAT (100.64.0.0/10), IPv6 unique-local (fc00::/7), and localhost variants. This boolean result sets the is_external_service property used by the CLI to warn about external calls.

Can I use MemPalace with vLLM or LM Studio instead of Ollama?

Yes. Configure the openai-compat provider with your local server endpoint (e.g., http://localhost:8000 for vLLM or http://localhost:1234 for LM Studio). The provider uses the standard /v1/chat/completions endpoint and requires no API key for local instances, though it supports bearer token authentication if your local server is configured with one.

What HTTP endpoints does the Ollama provider use internally?

The OllamaProvider sends POST requests to /api/chat for inference and /api/tags for health checks, both targeting http://localhost:11434 by default. The request body includes model names, message arrays, and optional parameters like format: "json" and num_ctx for context window sizing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →