How to Integrate Custom LLM Providers (Mistral, HuggingFace) with Llama-GitHub
Llama-GitHub integrates custom LLM providers through the LLMManager singleton in initial_load.py, which automatically instantiates LangChain-compatible models for Mistral via API keys or accepts pre-configured custom LLMs for HuggingFace and other providers.
Integrating custom LLM providers into your RAG pipeline requires a flexible abstraction layer. The llama-github repository solves this through the LLMManager class, which handles provider selection based on authentication credentials and supports direct injection of LangChain-compatible models. This architecture allows you to switch between OpenAI, Mistral, and open-source HuggingFace models without rewriting your prompt handling logic.
How LLMManager Handles Provider Selection
The LLMManager class in llama_github/llm_integration/initial_load.py implements a singleton pattern that examines initialization arguments to determine which provider to instantiate. The manager follows a cascading selection logic:
- OpenAI – Activated when an
openai_api_keyis supplied. - Mistral – Activated when a non-empty
mistral_api_keyis supplied; the code importsChatMistralAIfrom langchain-mistralai and creates two models (a full-size and a lightweight "simple" model) according to source lines 67-76. - HuggingFace / open-source models – Activated when
open_source_models_hg_diris provided. The current implementation records the directory and setsmodel_typeto"Hubgingface"(source lines 85-92), though the actual model loading requires extension. - Custom LangChain LLM – Activated when you pass a pre-configured LLM via the
llmargument, settingmodel_typeto"Custom_langchain_llm"(source lines 64-67).
Integrating Mistral AI
To integrate Mistral, install the LangChain integration and provide your API key. The manager automatically creates both a standard model and a lightweight variant for fast responses.
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler
# Supply the Mistral API key – the manager creates ChatMistralAI internally
manager = LLMManager(mistral_api_key="YOUR_MISTRAL_KEY")
handler = LLMHandler(llm_manager=manager)
response = await handler.ainvoke(
human_question="Explain the differences between Mistral and GPT-4.",
chat_history=["User: Hi", "Bot: Hello!"],
context=["Mistral is a lightweight model..."]
)
print(response)
Enable simple mode to use the lightweight model for fast replies:
manager = LLMManager(mistral_api_key="...", simple_mode=True)
handler = LLMHandler(llm_manager=manager)
# Uses the lightweight model when simple_llm=True
response = await handler.ainvoke("Quick fact about Python.", simple_llm=True)
Integrating HuggingFace and Open-Source Models
The repository includes a placeholder branch for HuggingFace integration. To complete the implementation, extend the elif open_source_models_hg_dir block in initial_load.py to load models using AutoModelForCausalLM and AutoTokenizer.
from transformers import AutoModelForCausalLM, AutoTokenizer
from langchain_core.language_models import BaseChatModel
from llama_github.llm_integration.llm_handler import LLMHandler
from llama_github.llm_integration.initial_load import LLMManager
class HFChatModel(BaseChatModel):
def __init__(self, model_name: str, device: str = "cpu"):
self.tokenizer = AutoTokenizer.from_pretrained(model_name)
self.model = AutoModelForCausalLM.from_pretrained(model_name).to(device)
async def ainvoke(self, messages, **kwargs):
prompt = "\n".join(m.content for m in messages)
inputs = self.tokenizer(prompt, return_tensors="pt").to(self.model.device)
output = self.model.generate(**inputs, max_new_tokens=150)
return self.tokenizer.decode(output[0], skip_special_tokens=True)
# Instantiate and inject
custom_llm = HFChatModel(model_name="mistralai/Mistral-7B-Instruct-v0.1")
manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)
When using the open_source_models_hg_dir parameter, the manager sets model_type to "Hubgingface", but you must implement the actual loading logic in the source file to enable true open-source model loading.
Using a Custom LangChain-Compatible LLM
Bypass automatic provider detection entirely by passing a LangChain-compatible LLM directly to the llm parameter. This approach supports any model implementing the async ainvoke interface, including custom wrappers for local models or proprietary APIs not explicitly supported.
from llama_github.llm_integration.initial_load import LLMManager
# Pass any BaseChatModel implementation
manager = LLMManager(llm=my_custom_chat_model)
When using this method, LLMManager sets model_type to "Custom_langchain_llm" and stores your object for use by LLMHandler.
Building Prompts with LLMHandler
The LLMHandler in llama_github/llm_integration/llm_handler.py consumes an LLMManager instance and constructs LangChain prompt pipelines. The handler is provider-agnostic: when model_type equals "OpenAI" or any other supported type, it creates a system-message prompt template, injects chat history, context, and the current human question, then invokes the model asynchronously via ainvoke (source lines 55-81).
The handler works with any provider as long as the underlying LLM implements the LangChain async ainvoke interface. No modifications to LLMHandler are required when adding new providers through LLMManager.
Summary
- LLMManager auto-detects providers based on API keys (
openai_api_key,mistral_api_key) or directory paths (open_source_models_hg_dir) during initialization. - Mistral integration requires
pip install langchain-mistralaiand creates dual models (standard and simple) whenmistral_api_keyis provided. - HuggingFace support currently exists as a placeholder in lines 85-92 of
initial_load.py; extend theelif open_source_models_hg_dirblock to implement full local model loading. - Custom LLMs bypass automatic selection when passed via the
llmparameter, provided they implement the LangChainBaseChatModelinterface with asyncainvoke. - LLMHandler remains provider-agnostic, handling prompt construction and asynchronous invocation regardless of the underlying model source.
Frequently Asked Questions
Do I need to modify LLMHandler to use a new provider?
No. LLMHandler is designed to be provider-agnostic. As long as your LLM is wrapped in a LangChain-compatible interface with an ainvoke method and passed through LLMManager, the handler will route prompts correctly without code changes.
How do I switch between the full model and simple model in Mistral?
Initialize LLMManager with simple_mode=True to create the lightweight variant. When calling handler.ainvoke(), pass simple_llm=True to use the fast model for specific queries, or simple_llm=False (default) to use the full model.
Can I use local HuggingFace models without an internet connection?
Yes, but you must implement the loading logic in the elif open_source_models_hg_dir block of initial_load.py. Currently, this branch only records the directory path. Extend it to load AutoModelForCausalLM from a local cache directory specified in open_source_models_hg_dir.
What interface must my custom LLM implement to work with Llama-GitHub?
Your custom LLM must inherit from LangChain's BaseChatModel and implement the asynchronous ainvoke(self, messages, **kwargs) method. The method should accept a list of message objects and return a string or AIMessage response, handling tokenization and generation internally.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →