# How the MemPalace closet_llm Module Interfaces with Local LLM Runtimes

> Discover how the closet_llm module interfaces with local LLM runtimes like Ollama for seamless offline operation. Learn about automatic endpoint detection and unified provider interfaces.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: internals
- Published: 2026-06-06

---

**The** `closet_llm` **module (implemented in** [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py)**) provides a unified provider interface that abstracts interactions with local LLM runtimes such as Ollama and OpenAI-compatible servers, featuring automatic local endpoint detection across RFC 1918, Tailscale, and IPv6 unique-local address ranges to enable fully offline operation.**

MemPalace is an open-source knowledge management system that prioritizes privacy through local-first architecture. The **closet_llm module** serves as the abstraction layer between the application's entity refinement pipeline and various large language model runtimes, enabling seamless integration with locally-hosted models via a consistent Python API defined in the core client file.

## Provider Architecture in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py)

The module centers on a base `LLMProvider` class that standardizes how MemPalace communicates with language models. All concrete implementations inherit from this base and must provide two core methods: `classify()` for sending prompts and `check_available()` for health checks, plus an `is_external_service` property indicating whether the configured endpoint resides outside the local machine.

Three concrete providers ship with the module:

- **OllamaProvider** – Interfaces with Ollama's HTTP API on `localhost:11434`
- **OpenAICompatProvider** – Connects to any OpenAI-compatible local server (vLLM, LM Studio, etc.)
- **AnthropicProvider** – Handles cloud-based Anthropic API calls (optional external service)

## How Providers Connect to Local Runtimes

Each provider implements specific HTTP protocols to bridge MemPalace with local inference engines, while the `LLMResponse` object standardizes the return format containing `text`, `model`, `provider`, and `raw` attributes.

### Ollama Integration (Default)

The **OllamaProvider** communicates via HTTP POST requests to `http://localhost:11434/api/chat`. The request body includes the model name, combined system and user messages, and optional flags such as `format: "json"` for structured output and `think` parameters. For health checks, it queries `/api/tags` to verify the server is responsive before attempting classification.

### OpenAI-Compatible Servers

For self-hosted OpenAI-compatible runtimes like vLLM, the provider constructs URLs ending in `/v1/chat/completions`. It POSTs JSON containing `model`, `messages`, `temperature`, and `response_format: {"type": "json_object"}` when JSON mode is enabled. When running locally, no API key is required, though the provider supports adding `Authorization: Bearer <key>` headers when configured.

### Anthropic (External)

While the Anthropic provider targets the hosted API at `api.anthropic.com`, the module still treats it uniformly, automatically appending JSON formatting instructions to system prompts when `json_mode=True`.

## Local Endpoint Detection Logic

The module implements intelligent locality checking through the `_endpoint_is_local()` helper function. This utility examines configured endpoints to determine if they resolve to local addresses, enabling the CLI to warn users before contacting external services.

Detected local ranges include:

- **RFC 1918** private IPv4 subnets (`10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16`)
- **Tailscale CGNAT** (`100.64.0.0/10`)
- **IPv6 unique-local** addresses (`fc00::/7`)
- **Localhost** variants (`localhost`, `127.0.0.1`, `*.local`)

This detection powers the `is_external_service` property, returning `False` for local runtimes and `True` for cloud providers.

## Factory Pattern and CLI Integration

The `get_provider()` factory function at the bottom of [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py) instantiates the appropriate provider class based on the `name` parameter looked up in the internal `PROVIDERS` dictionary. It passes through configuration parameters including `model`, `endpoint`, `api_key`, `timeout`, and provider-specific options like `num_ctx` for Ollama context sizing.

In [`mempalace/cli.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/cli.py), the integration follows this workflow:

1. **Argument parsing** – Reads `--llm-provider`, `--llm-model`, `--llm-endpoint`, and `--llm-api-key` flags
2. **Provider instantiation** – Calls `get_provider()` with the parsed configuration
3. **Health verification** – Invokes `provider.check_available()` during initialization (e.g., `mempalace init`)
4. **Classification** – Executes `provider.classify(system_prompt, user_prompt, json_mode=True)` during entity refinement pipelines

## Practical Implementation Examples

The following examples demonstrate direct usage of the closet_llm module's public API:

**Direct Ollama Provider Usage**

```python
from mempalace.llm_client import OllamaProvider

ollama = OllamaProvider(model="llama3:8b", endpoint="http://localhost:11434")
ok, msg = ollama.check_available()
if ok:
    resp = ollama.classify(
        system="You are a helpful assistant.",
        user="Extract the people mentioned in this text.",
        json_mode=True,
        think=False,
    )
    print(resp.text)          # JSON string with the classification result

else:
    print("Ollama not reachable:", msg)

```

**Runtime Provider Selection via Factory**

```python
from mempalace.llm_client import get_provider

provider = get_provider(
    name="openai-compat",
    model="gpt-4o-mini",
    endpoint="http://localhost:8000",   # a local vLLM server

    api_key=None,                       # no key needed for a local server

)

if provider.check_available()[0]:
    result = provider.classify(
        system="You are a factual extractor.",
        user="List all project names in the snippet.",
        json_mode=True,
    )
    print(result.text)

```

**Checking Endpoint Locality**

```python
from mempalace.llm_client import get_provider

provider = get_provider("ollama", model="gemma:2b")
print("Will this call go off‑machine?", provider.is_external_service)

# → False, because the endpoint resolves to localhost

```

## Summary

- The **closet_llm module** centralizes LLM interactions in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py), implementing a provider pattern that supports Ollama, OpenAI-compatible servers, and Anthropic through the `LLMProvider` base class.
- **Local runtime support** requires no API keys and operates over HTTP to `localhost:11434` (Ollama) or custom ports (vLLM), satisfying MemPalace's zero-API-required principle.
- **Automatic locality detection** distinguishes local from external endpoints using IP range validation (RFC 1918, Tailscale CGNAT, IPv6 unique-local), preserving the application's offline-first guarantee.
- The **`get_provider()` factory** enables consistent initialization across the CLI and entity refinement pipelines, while `check_available()` ensures local runtimes are reachable before inference begins.

## Frequently Asked Questions

### What file contains the closet_llm module implementation?

The closet_llm module is implemented in [`mempalace/llm_client.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/llm_client.py), which defines the `LLMProvider` base class, three concrete provider implementations (Ollama, OpenAI-compat, Anthropic), the `_endpoint_is_local()` detection logic, and the `get_provider()` factory function.

### How does the module determine if an LLM endpoint is running locally?

The `_endpoint_is_local()` helper function analyzes the configured endpoint URL against private IP ranges including RFC 1918 (10/8, 172.16/12, 192.168/16), Tailscale CGNAT (100.64.0.0/10), IPv6 unique-local (fc00::/7), and localhost variants. This boolean result sets the `is_external_service` property used by the CLI to warn about external calls.

### Can I use MemPalace with vLLM or LM Studio instead of Ollama?

Yes. Configure the `openai-compat` provider with your local server endpoint (e.g., `http://localhost:8000` for vLLM or `http://localhost:1234` for LM Studio). The provider uses the standard `/v1/chat/completions` endpoint and requires no API key for local instances, though it supports bearer token authentication if your local server is configured with one.

### What HTTP endpoints does the Ollama provider use internally?

The `OllamaProvider` sends POST requests to `/api/chat` for inference and `/api/tags` for health checks, both targeting `http://localhost:11434` by default. The request body includes model names, message arrays, and optional parameters like `format: "json"` and `num_ctx` for context window sizing.