How to Use Custom LangChain LLM Objects with llama-github: A Complete Guide
You can inject any LangChain-compatible LLM into llama-github by passing it to the LLMManager constructor, which stores the instance and automatically routes all downstream requests through your custom model instead of the default API-key initialization logic.
The llama-github repository isolates all language-model interactions behind two core abstractions: the LLMManager singleton and the LLMHandler consumer. Because the handler never instantiates LLMs itself—instead fetching them exclusively via LLMManager.get_llm()—you can seamlessly integrate custom LangChain LLM objects such as ChatOpenAI, ChatMistralAI, or locally-hosted HuggingFace models by injecting them at initialization time.
Architecture Overview
llama-github decouples model management from inference execution through two specialized components:
-
LLMManager(defined inllama_github/llm_integration/initial_load.py): A singleton responsible for holding the actual LangChain LLM instance. When you supply a custom LLM via thellmparameter, the manager stores it internally, setsmodel_typeto"Custom_langchain_llm", and bypasses the built-in API-key initialization logic (lines 63-66). -
LLMHandler(defined inllama_github/llm_integration/llm_handler.py): The inference interface that constructs LangChainChatPromptTemplateobjects, manages chat history and context injection, and executes calls viaawait llm.ainvoke(...). The handler retrieves the active LLM exclusively throughLLMManager.get_llm(), meaning it will automatically use whatever model instance the manager holds (lines 15-27).
This design ensures that any LangChain-compatible object implementing the standard chat interface (ainvoke, invoke, etc.) becomes immediately available to all downstream processing pipelines, including the RAG processor in llama_github/rag_processing/rag_processor.py.
Step-by-Step Implementation
Follow these four steps to replace the default LLM with your custom LangChain object:
- Instantiate your LangChain LLM using any provider (OpenAI, Anthropic, local HuggingFace, or custom subclasses).
- Create
LLMManagerwith the custom LLM passed to thellmargument. - Initialize
LLMHandlerwith your configured manager instance. - Invoke the handler asynchronously with questions, chat history, and optional context.
Creating Your Custom LangChain LLM
Begin by constructing any LangChain chat model that conforms to the base chat interface. For example, using OpenAI with non-default parameters:
from langchain_openai import ChatOpenAI
custom_llm = ChatOpenAI(
api_key="YOUR_OPENAI_API_KEY",
model="gpt-4o-mini",
temperature=0.9,
)
Or for local inference using HuggingFace:
from langchain_community.chat_models import ChatHF
custom_llm = ChatHF(model="google/flan-t5-base", temperature=0.2)
Injecting the LLM into LLMManager
Pass your instantiated LLM to the LLMManager constructor. According to the source code in initial_load.py, this assigns your object to the internal state and sets the model type identifier:
from llama_github.llm_integration.initial_load import LLMManager
manager = LLMManager(llm=custom_llm) # model_type automatically set to "Custom_langchain_llm"
Initializing LLMHandler with Your Manager
Construct the handler by supplying the manager instance. The handler will now fetch your custom LLM for every inference request:
from llama_github.llm_integration.llm_handler import LLMHandler
handler = LLMHandler(llm_manager=manager)
Executing Requests
Use the asynchronous ainvoke method to process queries. The handler formats the prompt template, injects history and context, and calls your custom LLM's ainvoke method:
import asyncio
async def ask():
answer = await handler.ainvoke(
human_question="Explain the difference between zero-shot and few-shot prompting.",
chat_history=["User: What is LLM?", "Assistant: A large language model..."],
context=["Relevant doc snippet..."],
)
print(answer)
asyncio.run(ask())
Complete Code Examples
Using OpenAI Chat Models
This example demonstrates integrating a custom ChatOpenAI instance with modified temperature settings:
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler
from langchain_openai import ChatOpenAI
import asyncio
# Configure custom LangChain LLM
custom_llm = ChatOpenAI(
api_key="YOUR_OPENAI_API_KEY",
model="gpt-4o-mini",
temperature=0.9,
)
# Inject into llama-github's manager
manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)
# Execute query
async def main():
response = await handler.ainvoke(
human_question="What are the advantages of retrieval-augmented generation?",
chat_history=[],
context=["RAG combines external knowledge retrieval with language generation..."]
)
print(response)
asyncio.run(main())
Using Local HuggingFace Models
For offline or privacy-sensitive environments, pass a local HuggingFace chat model:
from langchain_community.chat_models import ChatHF
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler
# Local model configuration
custom_llm = ChatHF(model="google/flan-t5-base", temperature=0.2)
# Integration remains identical
manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)
# All subsequent handler calls route to the local HF model
Key Source Files and Implementation Details
Understanding these three files provides complete visibility into how custom LangChain LLM objects flow through the system:
| File | Purpose | Critical Implementation Detail |
|---|---|---|
llama_github/llm_integration/initial_load.py |
Defines LLMManager singleton |
Lines 63-66 handle custom LLM storage and skip API-key initialization when llm parameter is provided |
llama_github/llm_integration/llm_handler.py |
Defines LLMHandler inference interface |
Lines 15-27 implement the get_llm() call that retrieves your custom object from the manager |
llama_github/rag_processing/rag_processor.py |
High-level RAG orchestration | Demonstrates production usage pattern constructing both manager and handler for retrieval-augmented generation pipelines |
Summary
LLMManageracts as the single source of truth for LLM instances; passing a custom LangChain object to its constructor stores it as"Custom_langchain_llm"and disables default API-key logic.LLMHandlernever creates LLM instances itself, instead querying the manager viaget_llm(), ensuring automatic propagation of your custom model to all inference calls.- Any LangChain-compatible chat model—including
ChatOpenAI,ChatMistralAI,ChatHF, or bespoke subclasses—works immediately upon injection. - The integration requires only three lines of configuration code: import
LLMManager, instantiate withllm=your_custom_model, and pass toLLMHandler.
Frequently Asked Questions
Can I switch LLMs dynamically after initializing LLMManager?
No. According to the llama-github source code in initial_load.py, the LLMManager initializes the LLM instance during construction. To use a different model, you must create a new LLMManager instance with the alternative LLM and instantiate a fresh LLMHandler pointing to that new manager.
Does llama-github support both synchronous and asynchronous invocation with custom LLMs?
Yes. While the examples emphasize ainvoke() for asynchronous execution, LLMHandler supports standard LangChain invocation patterns. Your custom LLM must implement both invoke() and ainvoke() methods to support both sync and async workflows within the handler's prompt execution pipeline.
What happens if I don't provide a custom LLM to LLMManager?
If you instantiate LLMManager without the llm parameter, the class falls back to its built-in initialization logic, which attempts to configure default models using environment API keys. The source code in initial_load.py (lines 63-66) explicitly checks for the presence of a custom LLM to determine whether to skip this default behavior.
Can I use custom embedding or reranker models alongside my custom LLM?
Yes. LLMManager accepts additional parameters for tokenizer, embedding models, and rerankers alongside the primary llm argument. You can configure these components in the same constructor call where you inject your custom LangChain LLM, allowing full customization of the retrieval and generation pipeline while maintaining the singleton pattern.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →