How to Configure Nemori for Different LLM Providers and Embedding Models
Nemori supports multiple LLM and embedding providers through the LLMClient and EmbeddingClient wrapper classes, allowing configuration via model name changes, base URL overrides, or complete client injection in DefaultProviders.
The nemori-ai/nemori repository provides a flexible memory management system that you can configure for different LLM providers or embedding models without modifying core logic. By leveraging the provider abstraction layer in src/services/providers.py, you can point Nemori at OpenAI, Azure OpenAI, self-hosted endpoints, or entirely custom implementations.
Understanding Nemori's LLM and Embedding Architecture
Nemori delegates all model interactions to two lightweight wrappers defined in src/utils/llm_client.py and src/utils/embedding_client.py. The LLMClient class handles chat completions via chat_completion and structured JSON generation via generate_json_response, while the EmbeddingClient class manages text embeddings through embed_texts and auto-detects output dimensions based on the model name.
The DefaultProviders factory in src/services/providers.py wires these clients into the service graph. By default, it reads the OPENAI_API_KEY environment variable and the model names defined in MemoryConfig (llm_model and embedding_model). You can override this behavior at three different levels of abstraction.
Three Methods to Configure Nemori for Different Providers
1. Swap Model Names via MemoryConfig
The simplest way to configure Nemori for different LLM providers or embedding models is to change the model identifiers in the configuration dataclass. MemoryConfig in src/config.py exposes llm_model and embedding_model fields that both wrapper classes read during instantiation.
This approach works for any provider that maintains OpenAI API compatibility, including OpenAI itself, Azure OpenAI, and most self-hosted inference servers. You only need to update the model strings to match the target provider's available models.
2. Override the Base URL for OpenAI-Compatible Endpoints
When targeting Azure OpenAI, a local LLM server, or any other OpenAI-compatible host, you must redirect the API calls to a custom endpoint. Both LLMClient and EmbeddingClient accept a base_url parameter that the underlying OpenAI SDK uses to route requests.
You can pass this parameter when manually constructing the clients, or extend MemoryConfig to include a custom base URL field. This method allows you to configure Nemori for different LLM providers without changing the core client logic, provided the alternative service adheres to the OpenAI API specification.
3. Inject Custom Client Implementations
For providers that do not follow the OpenAI API contract, or when you need fine-grained control over request handling, you can inject completely custom client implementations. The DefaultProviders constructor accepts optional llm_client and embedding_client parameters.
By creating a class that implements the same public methods as LLMClient (chat_completion, generate_json_response) or EmbeddingClient (embed_texts), you can integrate any SDK or local inference engine. This is the most flexible way to configure Nemori for different LLM providers or embedding models, supporting everything from local Llama.cpp servers to proprietary enterprise APIs.
Configuration Examples
Switch to a Different OpenAI Model via the Config
Change the model names in MemoryConfig to use different OpenAI models without touching client code.
from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders
# Adjust the model names – no code changes elsewhere needed
config = MemoryConfig(
llm_model="gpt-4o-mini", # or "gpt-4o", "gpt-3.5-turbo"
embedding_model="text-embedding-3-large" # or "text-embedding-ada-002"
)
providers = DefaultProviders(config)
Use Azure OpenAI by Setting base_url
Redirect API calls to an Azure OpenAI endpoint or any other OpenAI-compatible host.
from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders
config = MemoryConfig(
llm_model="gpt-35-turbo", # Azure model name
embedding_model="text-embedding-3-small",
)
# Azure OpenAI requires a custom endpoint like:
azure_endpoint = "https://my-azure-openai.openai.azure.com/v1"
# Build the clients manually with the custom base URL
from nemori.main.src.utils.llm_client import LLMClient
from nemori.main.src.utils.embedding_client import EmbeddingClient
llm = LLMClient(
api_key=config.openai_api_key,
model=config.llm_model,
base_url=azure_endpoint,
)
embedder = EmbeddingClient(
api_key=config.openai_api_key,
model=config.embedding_model,
base_url=azure_endpoint,
)
providers = DefaultProviders(config, llm_client=llm, embedding_client=embedder)
Plug in a Completely Custom LLM Client
Integrate a local Llama.cpp server or any non-OpenAI API by implementing the required interface.
from typing import List, Dict, Any
from nemori.main.src.services.providers import DefaultProviders
from nemori.main.src.config import MemoryConfig
class LlamaCppClient:
"""Very small wrapper around a locally hosted Llama.cpp HTTP API."""
def __init__(self, endpoint: str):
self.endpoint = endpoint
def chat_completion(self, messages: List[Dict[str, str]], **kwargs) -> Any:
# Perform a POST to the local server, return a dict mimicking OpenAI's response
import requests, json
payload = {"messages": messages, **kwargs}
resp = requests.post(self.endpoint, json=payload, timeout=30)
resp.raise_for_status()
return resp.json() # should contain .choices[0].message.content, .usage, etc.
# Build a thin adapter that satisfies the LLMClient interface used by Nemori
class LLMAdapter:
def __init__(self, client: LlamaCppClient):
self.client = client
def chat_completion(self, messages, temperature=0.7, max_tokens=None, category=None, **kw):
raw = self.client.chat_completion(messages, temperature=temperature,
max_tokens=max_tokens, **kw)
# Convert raw dict into the LLMResponse dataclass expected by Nemori
from nemori.main.src.utils.llm_client import LLMResponse
return LLMResponse(
content=raw["choices"][0]["message"]["content"],
usage=raw.get("usage", {}),
model=raw.get("model", "llama-cpp"),
finish_reason=raw["choices"][0].get("finish_reason", "stop"),
response_time=0.0, # optional timing info
)
# Use the adapter in the provider graph
config = MemoryConfig()
llama_client = LLMAdapter(LlamaCppClient(endpoint="http://localhost:8080/v1/chat/completions"))
# Embedding can stay the default OpenAI one, or you could similarly wrap a local embedder.
providers = DefaultProviders(config, llm_client=llama_client)
Override Only the Embedding Model
Switch to a different embedding model without affecting the LLM configuration.
from nemori.main.src.config import MemoryConfig
from nemori.main.src.services.providers import DefaultProviders
config = MemoryConfig(embedding_model="text-embedding-3-large")
providers = DefaultProviders(config)
# The embedding client automatically picks the correct dimension (3072) via its internal logic.
Key Source Files for Nemori Configuration
| File | Role | Link |
|---|---|---|
src/config.py |
Central configuration (model names, API key, optional base_url handling) |
View |
src/utils/llm_client.py |
Wrapper around OpenAI's chat/completion endpoint; accepts api_key, model, base_url |
View |
src/utils/embedding_client.py |
Wrapper around OpenAI's embedding endpoint; auto-detects dimension from model name | View |
src/services/providers.py |
Factory that builds the service graph; allows injection of custom LLM/embedding clients | View |
Summary
- Nemori uses
LLMClientandEmbeddingClientwrappers insrc/utils/to isolate provider-specific logic from the core memory system. - You can configure Nemori for different LLM providers or embedding models by updating
MemoryConfiginsrc/config.py, overriding thebase_urlfor OpenAI-compatible endpoints, or injecting fully custom client instances. - The
DefaultProvidersfactory accepts optionalllm_clientandembedding_clientparameters, enabling you to bypass the default OpenAI SDK integration entirely for local or proprietary models. EmbeddingClientautomatically detects embedding dimensions based on the model name, simplifying configuration when switching between models liketext-embedding-3-largeandtext-embedding-ada-002.
Frequently Asked Questions
How do I switch from OpenAI to Azure OpenAI in Nemori?
Override the base_url parameter when constructing LLMClient and EmbeddingClient to point to your Azure OpenAI endpoint (e.g., https://my-azure-openai.openai.azure.com/v1), then pass these custom instances to DefaultProviders. Ensure your llm_model value matches the Azure deployment name exactly.
Can I use a local LLM like Llama.cpp with Nemori?
Yes. Create a custom client class that implements the chat_completion method to POST requests to your local Llama.cpp HTTP server, then wrap it in an adapter that returns an LLMResponse dataclass. Inject this adapter into DefaultProviders via the llm_client parameter to replace the default OpenAI client entirely.
Where does Nemori store the default model configuration?
Default model names and API settings are defined in the MemoryConfig dataclass located in src/config.py. The llm_model and embedding_model fields specify which models LLMClient and EmbeddingClient use when calling the API, while the openai_api_key field reads from the OPENAI_API_KEY environment variable by default.
How do I change only the embedding model without affecting the LLM?
Instantiate MemoryConfig with only the embedding_model parameter changed (e.g., MemoryConfig(embedding_model="text-embedding-3-large")), then pass this config to DefaultProviders. The EmbeddingClient automatically detects the correct output dimensions for the specified model, while the LLM client continues using the default chat model unless explicitly overridden.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →