How to Configure Nemori for Different LLM Providers and Embedding Models

Nemori supports multiple LLM and embedding providers through the LLMClient and EmbeddingClient wrapper classes, allowing configuration via model name changes, base URL overrides, or complete client injection in DefaultProviders.

The nemori-ai/nemori repository provides a flexible memory management system that you can configure for different LLM providers or embedding models without modifying core logic. By leveraging the provider abstraction layer in src/services/providers.py, you can point Nemori at OpenAI, Azure OpenAI, self-hosted endpoints, or entirely custom implementations.

Understanding Nemori's LLM and Embedding Architecture

Nemori delegates all model interactions to two lightweight wrappers defined in src/utils/llm_client.py and src/utils/embedding_client.py. The LLMClient class handles chat completions via chat_completion and structured JSON generation via generate_json_response, while the EmbeddingClient class manages text embeddings through embed_texts and auto-detects output dimensions based on the model name.

The DefaultProviders factory in src/services/providers.py wires these clients into the service graph. By default, it reads the OPENAI_API_KEY environment variable and the model names defined in MemoryConfig (llm_model and embedding_model). You can override this behavior at three different levels of abstraction.

Three Methods to Configure Nemori for Different Providers

1. Swap Model Names via MemoryConfig

The simplest way to configure Nemori for different LLM providers or embedding models is to change the model identifiers in the configuration dataclass. MemoryConfig in src/config.py exposes llm_model and embedding_model fields that both wrapper classes read during instantiation.

This approach works for any provider that maintains OpenAI API compatibility, including OpenAI itself, Azure OpenAI, and most self-hosted inference servers. You only need to update the model strings to match the target provider's available models.

2. Override the Base URL for OpenAI-Compatible Endpoints

When targeting Azure OpenAI, a local LLM server, or any other OpenAI-compatible host, you must redirect the API calls to a custom endpoint. Both LLMClient and EmbeddingClient accept a base_url parameter that the underlying OpenAI SDK uses to route requests.

You can pass this parameter when manually constructing the clients, or extend MemoryConfig to include a custom base URL field. This method allows you to configure Nemori for different LLM providers without changing the core client logic, provided the alternative service adheres to the OpenAI API specification.

3. Inject Custom Client Implementations

For providers that do not follow the OpenAI API contract, or when you need fine-grained control over request handling, you can inject completely custom client implementations. The DefaultProviders constructor accepts optional llm_client and embedding_client parameters.

By creating a class that implements the same public methods as LLMClient (chat_completion, generate_json_response) or EmbeddingClient (embed_texts), you can integrate any SDK or local inference engine. This is the most flexible way to configure Nemori for different LLM providers or embedding models, supporting everything from local Llama.cpp servers to proprietary enterprise APIs.

Configuration Examples

Switch to a Different OpenAI Model via the Config

Change the model names in MemoryConfig to use different OpenAI models without touching client code.

from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders

# Adjust the model names – no code changes elsewhere needed

config = MemoryConfig(
    llm_model="gpt-4o-mini",              # or "gpt-4o", "gpt-3.5-turbo"

    embedding_model="text-embedding-3-large" # or "text-embedding-ada-002"

)

providers = DefaultProviders(config)

Use Azure OpenAI by Setting base_url

Redirect API calls to an Azure OpenAI endpoint or any other OpenAI-compatible host.

from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders

config = MemoryConfig(
    llm_model="gpt-35-turbo",                     # Azure model name

    embedding_model="text-embedding-3-small",
)

# Azure OpenAI requires a custom endpoint like:

azure_endpoint = "https://my-azure-openai.openai.azure.com/v1"

# Build the clients manually with the custom base URL

from nemori.main.src.utils.llm_client import LLMClient
from nemori.main.src.utils.embedding_client import EmbeddingClient

llm = LLMClient(
    api_key=config.openai_api_key,
    model=config.llm_model,
    base_url=azure_endpoint,
)

embedder = EmbeddingClient(
    api_key=config.openai_api_key,
    model=config.embedding_model,
    base_url=azure_endpoint,
)

providers = DefaultProviders(config, llm_client=llm, embedding_client=embedder)

Plug in a Completely Custom LLM Client

Integrate a local Llama.cpp server or any non-OpenAI API by implementing the required interface.

from typing import List, Dict, Any
from nemori.main.src.services.providers import DefaultProviders
from nemori.main.src.config import MemoryConfig

class LlamaCppClient:
    """Very small wrapper around a locally hosted Llama.cpp HTTP API."""
    def __init__(self, endpoint: str):
        self.endpoint = endpoint

    def chat_completion(self, messages: List[Dict[str, str]], **kwargs) -> Any:
        # Perform a POST to the local server, return a dict mimicking OpenAI's response

        import requests, json
        payload = {"messages": messages, **kwargs}
        resp = requests.post(self.endpoint, json=payload, timeout=30)
        resp.raise_for_status()
        return resp.json()   # should contain .choices[0].message.content, .usage, etc.

# Build a thin adapter that satisfies the LLMClient interface used by Nemori

class LLMAdapter:
    def __init__(self, client: LlamaCppClient):
        self.client = client

    def chat_completion(self, messages, temperature=0.7, max_tokens=None, category=None, **kw):
        raw = self.client.chat_completion(messages, temperature=temperature,
                                          max_tokens=max_tokens, **kw)
        # Convert raw dict into the LLMResponse dataclass expected by Nemori

        from nemori.main.src.utils.llm_client import LLMResponse
        return LLMResponse(
            content=raw["choices"][0]["message"]["content"],
            usage=raw.get("usage", {}),
            model=raw.get("model", "llama-cpp"),
            finish_reason=raw["choices"][0].get("finish_reason", "stop"),
            response_time=0.0,   # optional timing info

        )

# Use the adapter in the provider graph

config = MemoryConfig()
llama_client = LLMAdapter(LlamaCppClient(endpoint="http://localhost:8080/v1/chat/completions"))

# Embedding can stay the default OpenAI one, or you could similarly wrap a local embedder.

providers = DefaultProviders(config, llm_client=llama_client)

Override Only the Embedding Model

Switch to a different embedding model without affecting the LLM configuration.

from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders

config = MemoryConfig(embedding_model="text-embedding-3-large")
providers = DefaultProviders(config)

# The embedding client automatically picks the correct dimension (3072) via its internal logic.

Key Source Files for Nemori Configuration

File Role Link
src/config.py Central configuration (model names, API key, optional base_url handling) View
src/utils/llm_client.py Wrapper around OpenAI's chat/completion endpoint; accepts api_key, model, base_url View
src/utils/embedding_client.py Wrapper around OpenAI's embedding endpoint; auto-detects dimension from model name View
src/services/providers.py Factory that builds the service graph; allows injection of custom LLM/embedding clients View

Summary

  • Nemori uses LLMClient and EmbeddingClient wrappers in src/utils/ to isolate provider-specific logic from the core memory system.
  • You can configure Nemori for different LLM providers or embedding models by updating MemoryConfig in src/config.py, overriding the base_url for OpenAI-compatible endpoints, or injecting fully custom client instances.
  • The DefaultProviders factory accepts optional llm_client and embedding_client parameters, enabling you to bypass the default OpenAI SDK integration entirely for local or proprietary models.
  • EmbeddingClient automatically detects embedding dimensions based on the model name, simplifying configuration when switching between models like text-embedding-3-large and text-embedding-ada-002.

Frequently Asked Questions

How do I switch from OpenAI to Azure OpenAI in Nemori?

Override the base_url parameter when constructing LLMClient and EmbeddingClient to point to your Azure OpenAI endpoint (e.g., https://my-azure-openai.openai.azure.com/v1), then pass these custom instances to DefaultProviders. Ensure your llm_model value matches the Azure deployment name exactly.

Can I use a local LLM like Llama.cpp with Nemori?

Yes. Create a custom client class that implements the chat_completion method to POST requests to your local Llama.cpp HTTP server, then wrap it in an adapter that returns an LLMResponse dataclass. Inject this adapter into DefaultProviders via the llm_client parameter to replace the default OpenAI client entirely.

Where does Nemori store the default model configuration?

Default model names and API settings are defined in the MemoryConfig dataclass located in src/config.py. The llm_model and embedding_model fields specify which models LLMClient and EmbeddingClient use when calling the API, while the openai_api_key field reads from the OPENAI_API_KEY environment variable by default.

How do I change only the embedding model without affecting the LLM?

Instantiate MemoryConfig with only the embedding_model parameter changed (e.g., MemoryConfig(embedding_model="text-embedding-3-large")), then pass this config to DefaultProviders. The EmbeddingClient automatically detects the correct output dimensions for the specified model, while the LLM client continues using the default chat model unless explicitly overridden.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →