# How to Configure Nemori for Different LLM Providers and Embedding Models

> Configure Nemori for various LLM and embedding providers. Easily switch models, override URLs, or inject custom clients for flexible integration.

- Repository: [Nemori AI/nemori](https://github.com/nemori-ai/nemori)
- Tags: how-to-guide
- Published: 2026-03-08

---

**Nemori supports multiple LLM and embedding providers through the `LLMClient` and `EmbeddingClient` wrapper classes, allowing configuration via model name changes, base URL overrides, or complete client injection in `DefaultProviders`.**

The `nemori-ai/nemori` repository provides a flexible memory management system that you can configure for different LLM providers or embedding models without modifying core logic. By leveraging the provider abstraction layer in [`src/services/providers.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/providers.py), you can point Nemori at OpenAI, Azure OpenAI, self-hosted endpoints, or entirely custom implementations.

## Understanding Nemori's LLM and Embedding Architecture

Nemori delegates all model interactions to two lightweight wrappers defined in [`src/utils/llm_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/llm_client.py) and [`src/utils/embedding_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/embedding_client.py). The **`LLMClient`** class handles chat completions via `chat_completion` and structured JSON generation via `generate_json_response`, while the **`EmbeddingClient`** class manages text embeddings through `embed_texts` and auto-detects output dimensions based on the model name.

The **`DefaultProviders`** factory in [`src/services/providers.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/providers.py) wires these clients into the service graph. By default, it reads the **`OPENAI_API_KEY`** environment variable and the model names defined in **`MemoryConfig`** (`llm_model` and `embedding_model`). You can override this behavior at three different levels of abstraction.

## Three Methods to Configure Nemori for Different Providers

### 1. Swap Model Names via MemoryConfig

The simplest way to configure Nemori for different LLM providers or embedding models is to change the model identifiers in the configuration dataclass. **`MemoryConfig`** in [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py) exposes `llm_model` and `embedding_model` fields that both wrapper classes read during instantiation.

This approach works for any provider that maintains OpenAI API compatibility, including OpenAI itself, Azure OpenAI, and most self-hosted inference servers. You only need to update the model strings to match the target provider's available models.

### 2. Override the Base URL for OpenAI-Compatible Endpoints

When targeting Azure OpenAI, a local LLM server, or any other OpenAI-compatible host, you must redirect the API calls to a custom endpoint. Both **`LLMClient`** and **`EmbeddingClient`** accept a **`base_url`** parameter that the underlying OpenAI SDK uses to route requests.

You can pass this parameter when manually constructing the clients, or extend `MemoryConfig` to include a custom base URL field. This method allows you to configure Nemori for different LLM providers without changing the core client logic, provided the alternative service adheres to the OpenAI API specification.

### 3. Inject Custom Client Implementations

For providers that do not follow the OpenAI API contract, or when you need fine-grained control over request handling, you can inject completely custom client implementations. The **`DefaultProviders`** constructor accepts optional **`llm_client`** and **`embedding_client`** parameters.

By creating a class that implements the same public methods as `LLMClient` (`chat_completion`, `generate_json_response`) or `EmbeddingClient` (`embed_texts`), you can integrate any SDK or local inference engine. This is the most flexible way to configure Nemori for different LLM providers or embedding models, supporting everything from local Llama.cpp servers to proprietary enterprise APIs.

## Configuration Examples

### Switch to a Different OpenAI Model via the Config

Change the model names in `MemoryConfig` to use different OpenAI models without touching client code.

```python
from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders

# Adjust the model names – no code changes elsewhere needed

config = MemoryConfig(
    llm_model="gpt-4o-mini",              # or "gpt-4o", "gpt-3.5-turbo"

    embedding_model="text-embedding-3-large" # or "text-embedding-ada-002"

)

providers = DefaultProviders(config)

```

### Use Azure OpenAI by Setting `base_url`

Redirect API calls to an Azure OpenAI endpoint or any other OpenAI-compatible host.

```python
from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders

config = MemoryConfig(
    llm_model="gpt-35-turbo",                     # Azure model name

    embedding_model="text-embedding-3-small",
)

# Azure OpenAI requires a custom endpoint like:

azure_endpoint = "https://my-azure-openai.openai.azure.com/v1"

# Build the clients manually with the custom base URL

from nemori.main.src.utils.llm_client import LLMClient
from nemori.main.src.utils.embedding_client import EmbeddingClient

llm = LLMClient(
    api_key=config.openai_api_key,
    model=config.llm_model,
    base_url=azure_endpoint,
)

embedder = EmbeddingClient(
    api_key=config.openai_api_key,
    model=config.embedding_model,
    base_url=azure_endpoint,
)

providers = DefaultProviders(config, llm_client=llm, embedding_client=embedder)

```

### Plug in a Completely Custom LLM Client

Integrate a local Llama.cpp server or any non-OpenAI API by implementing the required interface.

```python
from typing import List, Dict, Any
from nemori.main.src.services.providers import DefaultProviders
from nemori.main.src.config import MemoryConfig

class LlamaCppClient:
    """Very small wrapper around a locally hosted Llama.cpp HTTP API."""
    def __init__(self, endpoint: str):
        self.endpoint = endpoint

    def chat_completion(self, messages: List[Dict[str, str]], **kwargs) -> Any:
        # Perform a POST to the local server, return a dict mimicking OpenAI's response

        import requests, json
        payload = {"messages": messages, **kwargs}
        resp = requests.post(self.endpoint, json=payload, timeout=30)
        resp.raise_for_status()
        return resp.json()   # should contain .choices[0].message.content, .usage, etc.

# Build a thin adapter that satisfies the LLMClient interface used by Nemori

class LLMAdapter:
    def __init__(self, client: LlamaCppClient):
        self.client = client

    def chat_completion(self, messages, temperature=0.7, max_tokens=None, category=None, **kw):
        raw = self.client.chat_completion(messages, temperature=temperature,
                                          max_tokens=max_tokens, **kw)
        # Convert raw dict into the LLMResponse dataclass expected by Nemori

        from nemori.main.src.utils.llm_client import LLMResponse
        return LLMResponse(
            content=raw["choices"][0]["message"]["content"],
            usage=raw.get("usage", {}),
            model=raw.get("model", "llama-cpp"),
            finish_reason=raw["choices"][0].get("finish_reason", "stop"),
            response_time=0.0,   # optional timing info

        )

# Use the adapter in the provider graph

config = MemoryConfig()
llama_client = LLMAdapter(LlamaCppClient(endpoint="http://localhost:8080/v1/chat/completions"))

# Embedding can stay the default OpenAI one, or you could similarly wrap a local embedder.

providers = DefaultProviders(config, llm_client=llama_client)

```

### Override Only the Embedding Model

Switch to a different embedding model without affecting the LLM configuration.

```python
from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders

config = MemoryConfig(embedding_model="text-embedding-3-large")
providers = DefaultProviders(config)

# The embedding client automatically picks the correct dimension (3072) via its internal logic.

```

## Key Source Files for Nemori Configuration

| File | Role | Link |
|------|------|------|
| [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py) | Central configuration (model names, API key, optional `base_url` handling) | [View](https://github.com/nemori-ai/nemori/blob/main/src/config.py) |
| [`src/utils/llm_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/llm_client.py) | Wrapper around OpenAI's chat/completion endpoint; accepts `api_key`, `model`, `base_url` | [View](https://github.com/nemori-ai/nemori/blob/main/src/utils/llm_client.py) |
| [`src/utils/embedding_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/embedding_client.py) | Wrapper around OpenAI's embedding endpoint; auto-detects dimension from model name | [View](https://github.com/nemori-ai/nemori/blob/main/src/utils/embedding_client.py) |
| [`src/services/providers.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/providers.py) | Factory that builds the service graph; allows injection of custom LLM/embedding clients | [View](https://github.com/nemori-ai/nemori/blob/main/src/services/providers.py) |

## Summary

- Nemori uses **`LLMClient`** and **`EmbeddingClient`** wrappers in `src/utils/` to isolate provider-specific logic from the core memory system.
- You can configure Nemori for different LLM providers or embedding models by updating **`MemoryConfig`** in [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py), overriding the **`base_url`** for OpenAI-compatible endpoints, or injecting fully custom client instances.
- The **`DefaultProviders`** factory accepts optional `llm_client` and `embedding_client` parameters, enabling you to bypass the default OpenAI SDK integration entirely for local or proprietary models.
- **`EmbeddingClient`** automatically detects embedding dimensions based on the model name, simplifying configuration when switching between models like `text-embedding-3-large` and `text-embedding-ada-002`.

## Frequently Asked Questions

### How do I switch from OpenAI to Azure OpenAI in Nemori?

Override the `base_url` parameter when constructing `LLMClient` and `EmbeddingClient` to point to your Azure OpenAI endpoint (e.g., `https://my-azure-openai.openai.azure.com/v1`), then pass these custom instances to `DefaultProviders`. Ensure your `llm_model` value matches the Azure deployment name exactly.

### Can I use a local LLM like Llama.cpp with Nemori?

Yes. Create a custom client class that implements the `chat_completion` method to POST requests to your local Llama.cpp HTTP server, then wrap it in an adapter that returns an `LLMResponse` dataclass. Inject this adapter into `DefaultProviders` via the `llm_client` parameter to replace the default OpenAI client entirely.

### Where does Nemori store the default model configuration?

Default model names and API settings are defined in the `MemoryConfig` dataclass located in [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py). The `llm_model` and `embedding_model` fields specify which models `LLMClient` and `EmbeddingClient` use when calling the API, while the `openai_api_key` field reads from the `OPENAI_API_KEY` environment variable by default.

### How do I change only the embedding model without affecting the LLM?

Instantiate `MemoryConfig` with only the `embedding_model` parameter changed (e.g., `MemoryConfig(embedding_model="text-embedding-3-large")`), then pass this config to `DefaultProviders`. The `EmbeddingClient` automatically detects the correct output dimensions for the specified model, while the LLM client continues using the default chat model unless explicitly overridden.