# How to Integrate Custom LLM Providers (Mistral, HuggingFace) with Llama-GitHub

> Integrate custom LLM providers like Mistral and HuggingFace with Llama-GitHub using the LLMManager. Easily add your own models for enhanced functionality.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: how-to-guide
- Published: 2026-03-04

---

**Llama-GitHub integrates custom LLM providers through the `LLMManager` singleton in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py), which automatically instantiates LangChain-compatible models for Mistral via API keys or accepts pre-configured custom LLMs for HuggingFace and other providers.**

Integrating custom LLM providers into your RAG pipeline requires a flexible abstraction layer. The llama-github repository solves this through the `LLMManager` class, which handles provider selection based on authentication credentials and supports direct injection of LangChain-compatible models. This architecture allows you to switch between OpenAI, Mistral, and open-source HuggingFace models without rewriting your prompt handling logic.

## How LLMManager Handles Provider Selection

The `LLMManager` class in [`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py) implements a singleton pattern that examines initialization arguments to determine which provider to instantiate. The manager follows a cascading selection logic:

- **OpenAI** – Activated when an `openai_api_key` is supplied.
- **Mistral** – Activated when a non-empty `mistral_api_key` is supplied; the code imports `ChatMistralAI` from *langchain-mistralai* and creates two models (a full-size and a lightweight "simple" model) according to source lines 67-76.
- **HuggingFace / open-source models** – Activated when `open_source_models_hg_dir` is provided. The current implementation records the directory and sets `model_type` to `"Hubgingface"` (source lines 85-92), though the actual model loading requires extension.
- **Custom LangChain LLM** – Activated when you pass a pre-configured LLM via the `llm` argument, setting `model_type` to `"Custom_langchain_llm"` (source lines 64-67).

## Integrating Mistral AI

To integrate Mistral, install the LangChain integration and provide your API key. The manager automatically creates both a standard model and a lightweight variant for fast responses.

```python
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler

# Supply the Mistral API key – the manager creates ChatMistralAI internally

manager = LLMManager(mistral_api_key="YOUR_MISTRAL_KEY")
handler = LLMHandler(llm_manager=manager)

response = await handler.ainvoke(
    human_question="Explain the differences between Mistral and GPT-4.",
    chat_history=["User: Hi", "Bot: Hello!"],
    context=["Mistral is a lightweight model..."]
)
print(response)

```

Enable **simple mode** to use the lightweight model for fast replies:

```python
manager = LLMManager(mistral_api_key="...", simple_mode=True)
handler = LLMHandler(llm_manager=manager)

# Uses the lightweight model when simple_llm=True

response = await handler.ainvoke("Quick fact about Python.", simple_llm=True)

```

## Integrating HuggingFace and Open-Source Models

The repository includes a placeholder branch for HuggingFace integration. To complete the implementation, extend the `elif open_source_models_hg_dir` block in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py) to load models using `AutoModelForCausalLM` and `AutoTokenizer`.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from langchain_core.language_models import BaseChatModel
from llama_github.llm_integration.llm_handler import LLMHandler
from llama_github.llm_integration.initial_load import LLMManager

class HFChatModel(BaseChatModel):
    def __init__(self, model_name: str, device: str = "cpu"):
        self.tokenizer = AutoTokenizer.from_pretrained(model_name)
        self.model = AutoModelForCausalLM.from_pretrained(model_name).to(device)

    async def ainvoke(self, messages, **kwargs):
        prompt = "\n".join(m.content for m in messages)
        inputs = self.tokenizer(prompt, return_tensors="pt").to(self.model.device)
        output = self.model.generate(**inputs, max_new_tokens=150)
        return self.tokenizer.decode(output[0], skip_special_tokens=True)

# Instantiate and inject

custom_llm = HFChatModel(model_name="mistralai/Mistral-7B-Instruct-v0.1")
manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)

```

When using the `open_source_models_hg_dir` parameter, the manager sets `model_type` to `"Hubgingface"`, but you must implement the actual loading logic in the source file to enable true open-source model loading.

## Using a Custom LangChain-Compatible LLM

Bypass automatic provider detection entirely by passing a LangChain-compatible LLM directly to the `llm` parameter. This approach supports any model implementing the async `ainvoke` interface, including custom wrappers for local models or proprietary APIs not explicitly supported.

```python
from llama_github.llm_integration.initial_load import LLMManager

# Pass any BaseChatModel implementation

manager = LLMManager(llm=my_custom_chat_model)

```

When using this method, `LLMManager` sets `model_type` to `"Custom_langchain_llm"` and stores your object for use by `LLMHandler`.

## Building Prompts with LLMHandler

The `LLMHandler` in [`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py) consumes an `LLMManager` instance and constructs LangChain prompt pipelines. The handler is provider-agnostic: when `model_type` equals `"OpenAI"` or any other supported type, it creates a system-message prompt template, injects chat history, context, and the current human question, then invokes the model asynchronously via `ainvoke` (source lines 55-81).

The handler works with any provider as long as the underlying LLM implements the LangChain async `ainvoke` interface. No modifications to `LLMHandler` are required when adding new providers through `LLMManager`.

## Summary

- **LLMManager** auto-detects providers based on API keys (`openai_api_key`, `mistral_api_key`) or directory paths (`open_source_models_hg_dir`) during initialization.
- **Mistral integration** requires `pip install langchain-mistralai` and creates dual models (standard and simple) when `mistral_api_key` is provided.
- **HuggingFace support** currently exists as a placeholder in lines 85-92 of [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py); extend the `elif open_source_models_hg_dir` block to implement full local model loading.
- **Custom LLMs** bypass automatic selection when passed via the `llm` parameter, provided they implement the LangChain `BaseChatModel` interface with async `ainvoke`.
- **LLMHandler** remains provider-agnostic, handling prompt construction and asynchronous invocation regardless of the underlying model source.

## Frequently Asked Questions

### Do I need to modify LLMHandler to use a new provider?

No. `LLMHandler` is designed to be provider-agnostic. As long as your LLM is wrapped in a LangChain-compatible interface with an `ainvoke` method and passed through `LLMManager`, the handler will route prompts correctly without code changes.

### How do I switch between the full model and simple model in Mistral?

Initialize `LLMManager` with `simple_mode=True` to create the lightweight variant. When calling `handler.ainvoke()`, pass `simple_llm=True` to use the fast model for specific queries, or `simple_llm=False` (default) to use the full model.

### Can I use local HuggingFace models without an internet connection?

Yes, but you must implement the loading logic in the `elif open_source_models_hg_dir` block of [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py). Currently, this branch only records the directory path. Extend it to load `AutoModelForCausalLM` from a local cache directory specified in `open_source_models_hg_dir`.

### What interface must my custom LLM implement to work with Llama-GitHub?

Your custom LLM must inherit from LangChain's `BaseChatModel` and implement the asynchronous `ainvoke(self, messages, **kwargs)` method. The method should accept a list of message objects and return a string or AIMessage response, handling tokenization and generation internally.