How to Integrate Custom LLM Providers (Mistral, HuggingFace) with Llama-GitHub

Llama-GitHub integrates custom LLM providers through the LLMManager singleton in initial_load.py, which automatically instantiates LangChain-compatible models for Mistral via API keys or accepts pre-configured custom LLMs for HuggingFace and other providers.

Integrating custom LLM providers into your RAG pipeline requires a flexible abstraction layer. The llama-github repository solves this through the LLMManager class, which handles provider selection based on authentication credentials and supports direct injection of LangChain-compatible models. This architecture allows you to switch between OpenAI, Mistral, and open-source HuggingFace models without rewriting your prompt handling logic.

How LLMManager Handles Provider Selection

The LLMManager class in llama_github/llm_integration/initial_load.py implements a singleton pattern that examines initialization arguments to determine which provider to instantiate. The manager follows a cascading selection logic:

  • OpenAI – Activated when an openai_api_key is supplied.
  • Mistral – Activated when a non-empty mistral_api_key is supplied; the code imports ChatMistralAI from langchain-mistralai and creates two models (a full-size and a lightweight "simple" model) according to source lines 67-76.
  • HuggingFace / open-source models – Activated when open_source_models_hg_dir is provided. The current implementation records the directory and sets model_type to "Hubgingface" (source lines 85-92), though the actual model loading requires extension.
  • Custom LangChain LLM – Activated when you pass a pre-configured LLM via the llm argument, setting model_type to "Custom_langchain_llm" (source lines 64-67).

Integrating Mistral AI

To integrate Mistral, install the LangChain integration and provide your API key. The manager automatically creates both a standard model and a lightweight variant for fast responses.

from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler

# Supply the Mistral API key – the manager creates ChatMistralAI internally

manager = LLMManager(mistral_api_key="YOUR_MISTRAL_KEY")
handler = LLMHandler(llm_manager=manager)

response = await handler.ainvoke(
    human_question="Explain the differences between Mistral and GPT-4.",
    chat_history=["User: Hi", "Bot: Hello!"],
    context=["Mistral is a lightweight model..."]
)
print(response)

Enable simple mode to use the lightweight model for fast replies:

manager = LLMManager(mistral_api_key="...", simple_mode=True)
handler = LLMHandler(llm_manager=manager)

# Uses the lightweight model when simple_llm=True

response = await handler.ainvoke("Quick fact about Python.", simple_llm=True)

Integrating HuggingFace and Open-Source Models

The repository includes a placeholder branch for HuggingFace integration. To complete the implementation, extend the elif open_source_models_hg_dir block in initial_load.py to load models using AutoModelForCausalLM and AutoTokenizer.

from transformers import AutoModelForCausalLM, AutoTokenizer
from langchain_core.language_models import BaseChatModel
from llama_github.llm_integration.llm_handler import LLMHandler
from llama_github.llm_integration.initial_load import LLMManager

class HFChatModel(BaseChatModel):
    def __init__(self, model_name: str, device: str = "cpu"):
        self.tokenizer = AutoTokenizer.from_pretrained(model_name)
        self.model = AutoModelForCausalLM.from_pretrained(model_name).to(device)

    async def ainvoke(self, messages, **kwargs):
        prompt = "\n".join(m.content for m in messages)
        inputs = self.tokenizer(prompt, return_tensors="pt").to(self.model.device)
        output = self.model.generate(**inputs, max_new_tokens=150)
        return self.tokenizer.decode(output[0], skip_special_tokens=True)

# Instantiate and inject

custom_llm = HFChatModel(model_name="mistralai/Mistral-7B-Instruct-v0.1")
manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)

When using the open_source_models_hg_dir parameter, the manager sets model_type to "Hubgingface", but you must implement the actual loading logic in the source file to enable true open-source model loading.

Using a Custom LangChain-Compatible LLM

Bypass automatic provider detection entirely by passing a LangChain-compatible LLM directly to the llm parameter. This approach supports any model implementing the async ainvoke interface, including custom wrappers for local models or proprietary APIs not explicitly supported.

from llama_github.llm_integration.initial_load import LLMManager

# Pass any BaseChatModel implementation

manager = LLMManager(llm=my_custom_chat_model)

When using this method, LLMManager sets model_type to "Custom_langchain_llm" and stores your object for use by LLMHandler.

Building Prompts with LLMHandler

The LLMHandler in llama_github/llm_integration/llm_handler.py consumes an LLMManager instance and constructs LangChain prompt pipelines. The handler is provider-agnostic: when model_type equals "OpenAI" or any other supported type, it creates a system-message prompt template, injects chat history, context, and the current human question, then invokes the model asynchronously via ainvoke (source lines 55-81).

The handler works with any provider as long as the underlying LLM implements the LangChain async ainvoke interface. No modifications to LLMHandler are required when adding new providers through LLMManager.

Summary

  • LLMManager auto-detects providers based on API keys (openai_api_key, mistral_api_key) or directory paths (open_source_models_hg_dir) during initialization.
  • Mistral integration requires pip install langchain-mistralai and creates dual models (standard and simple) when mistral_api_key is provided.
  • HuggingFace support currently exists as a placeholder in lines 85-92 of initial_load.py; extend the elif open_source_models_hg_dir block to implement full local model loading.
  • Custom LLMs bypass automatic selection when passed via the llm parameter, provided they implement the LangChain BaseChatModel interface with async ainvoke.
  • LLMHandler remains provider-agnostic, handling prompt construction and asynchronous invocation regardless of the underlying model source.

Frequently Asked Questions

Do I need to modify LLMHandler to use a new provider?

No. LLMHandler is designed to be provider-agnostic. As long as your LLM is wrapped in a LangChain-compatible interface with an ainvoke method and passed through LLMManager, the handler will route prompts correctly without code changes.

How do I switch between the full model and simple model in Mistral?

Initialize LLMManager with simple_mode=True to create the lightweight variant. When calling handler.ainvoke(), pass simple_llm=True to use the fast model for specific queries, or simple_llm=False (default) to use the full model.

Can I use local HuggingFace models without an internet connection?

Yes, but you must implement the loading logic in the elif open_source_models_hg_dir block of initial_load.py. Currently, this branch only records the directory path. Extend it to load AutoModelForCausalLM from a local cache directory specified in open_source_models_hg_dir.

What interface must my custom LLM implement to work with Llama-GitHub?

Your custom LLM must inherit from LangChain's BaseChatModel and implement the asynchronous ainvoke(self, messages, **kwargs) method. The method should accept a list of message objects and return a string or AIMessage response, handling tokenization and generation internally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →