How to Use Custom LangChain LLM Objects with llama-github: A Complete Guide

You can inject any LangChain-compatible LLM into llama-github by passing it to the LLMManager constructor, which stores the instance and automatically routes all downstream requests through your custom model instead of the default API-key initialization logic.

The llama-github repository isolates all language-model interactions behind two core abstractions: the LLMManager singleton and the LLMHandler consumer. Because the handler never instantiates LLMs itself—instead fetching them exclusively via LLMManager.get_llm()—you can seamlessly integrate custom LangChain LLM objects such as ChatOpenAI, ChatMistralAI, or locally-hosted HuggingFace models by injecting them at initialization time.

Architecture Overview

llama-github decouples model management from inference execution through two specialized components:

  • LLMManager (defined in llama_github/llm_integration/initial_load.py): A singleton responsible for holding the actual LangChain LLM instance. When you supply a custom LLM via the llm parameter, the manager stores it internally, sets model_type to "Custom_langchain_llm", and bypasses the built-in API-key initialization logic (lines 63-66).

  • LLMHandler (defined in llama_github/llm_integration/llm_handler.py): The inference interface that constructs LangChain ChatPromptTemplate objects, manages chat history and context injection, and executes calls via await llm.ainvoke(...). The handler retrieves the active LLM exclusively through LLMManager.get_llm(), meaning it will automatically use whatever model instance the manager holds (lines 15-27).

This design ensures that any LangChain-compatible object implementing the standard chat interface (ainvoke, invoke, etc.) becomes immediately available to all downstream processing pipelines, including the RAG processor in llama_github/rag_processing/rag_processor.py.

Step-by-Step Implementation

Follow these four steps to replace the default LLM with your custom LangChain object:

  1. Instantiate your LangChain LLM using any provider (OpenAI, Anthropic, local HuggingFace, or custom subclasses).
  2. Create LLMManager with the custom LLM passed to the llm argument.
  3. Initialize LLMHandler with your configured manager instance.
  4. Invoke the handler asynchronously with questions, chat history, and optional context.

Creating Your Custom LangChain LLM

Begin by constructing any LangChain chat model that conforms to the base chat interface. For example, using OpenAI with non-default parameters:

from langchain_openai import ChatOpenAI

custom_llm = ChatOpenAI(
    api_key="YOUR_OPENAI_API_KEY",
    model="gpt-4o-mini",
    temperature=0.9,
)

Or for local inference using HuggingFace:

from langchain_community.chat_models import ChatHF

custom_llm = ChatHF(model="google/flan-t5-base", temperature=0.2)

Injecting the LLM into LLMManager

Pass your instantiated LLM to the LLMManager constructor. According to the source code in initial_load.py, this assigns your object to the internal state and sets the model type identifier:

from llama_github.llm_integration.initial_load import LLMManager

manager = LLMManager(llm=custom_llm)  # model_type automatically set to "Custom_langchain_llm"

Initializing LLMHandler with Your Manager

Construct the handler by supplying the manager instance. The handler will now fetch your custom LLM for every inference request:

from llama_github.llm_integration.llm_handler import LLMHandler

handler = LLMHandler(llm_manager=manager)

Executing Requests

Use the asynchronous ainvoke method to process queries. The handler formats the prompt template, injects history and context, and calls your custom LLM's ainvoke method:

import asyncio

async def ask():
    answer = await handler.ainvoke(
        human_question="Explain the difference between zero-shot and few-shot prompting.",
        chat_history=["User: What is LLM?", "Assistant: A large language model..."],
        context=["Relevant doc snippet..."],
    )
    print(answer)

asyncio.run(ask())

Complete Code Examples

Using OpenAI Chat Models

This example demonstrates integrating a custom ChatOpenAI instance with modified temperature settings:

from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler
from langchain_openai import ChatOpenAI
import asyncio

# Configure custom LangChain LLM

custom_llm = ChatOpenAI(
    api_key="YOUR_OPENAI_API_KEY",
    model="gpt-4o-mini",
    temperature=0.9,
)

# Inject into llama-github's manager

manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)

# Execute query

async def main():
    response = await handler.ainvoke(
        human_question="What are the advantages of retrieval-augmented generation?",
        chat_history=[],
        context=["RAG combines external knowledge retrieval with language generation..."]
    )
    print(response)

asyncio.run(main())

Using Local HuggingFace Models

For offline or privacy-sensitive environments, pass a local HuggingFace chat model:

from langchain_community.chat_models import ChatHF
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler

# Local model configuration

custom_llm = ChatHF(model="google/flan-t5-base", temperature=0.2)

# Integration remains identical

manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)

# All subsequent handler calls route to the local HF model

Key Source Files and Implementation Details

Understanding these three files provides complete visibility into how custom LangChain LLM objects flow through the system:

File Purpose Critical Implementation Detail
llama_github/llm_integration/initial_load.py Defines LLMManager singleton Lines 63-66 handle custom LLM storage and skip API-key initialization when llm parameter is provided
llama_github/llm_integration/llm_handler.py Defines LLMHandler inference interface Lines 15-27 implement the get_llm() call that retrieves your custom object from the manager
llama_github/rag_processing/rag_processor.py High-level RAG orchestration Demonstrates production usage pattern constructing both manager and handler for retrieval-augmented generation pipelines

Summary

  • LLMManager acts as the single source of truth for LLM instances; passing a custom LangChain object to its constructor stores it as "Custom_langchain_llm" and disables default API-key logic.
  • LLMHandler never creates LLM instances itself, instead querying the manager via get_llm(), ensuring automatic propagation of your custom model to all inference calls.
  • Any LangChain-compatible chat model—including ChatOpenAI, ChatMistralAI, ChatHF, or bespoke subclasses—works immediately upon injection.
  • The integration requires only three lines of configuration code: import LLMManager, instantiate with llm=your_custom_model, and pass to LLMHandler.

Frequently Asked Questions

Can I switch LLMs dynamically after initializing LLMManager?

No. According to the llama-github source code in initial_load.py, the LLMManager initializes the LLM instance during construction. To use a different model, you must create a new LLMManager instance with the alternative LLM and instantiate a fresh LLMHandler pointing to that new manager.

Does llama-github support both synchronous and asynchronous invocation with custom LLMs?

Yes. While the examples emphasize ainvoke() for asynchronous execution, LLMHandler supports standard LangChain invocation patterns. Your custom LLM must implement both invoke() and ainvoke() methods to support both sync and async workflows within the handler's prompt execution pipeline.

What happens if I don't provide a custom LLM to LLMManager?

If you instantiate LLMManager without the llm parameter, the class falls back to its built-in initialization logic, which attempts to configure default models using environment API keys. The source code in initial_load.py (lines 63-66) explicitly checks for the presence of a custom LLM to determine whether to skip this default behavior.

Can I use custom embedding or reranker models alongside my custom LLM?

Yes. LLMManager accepts additional parameters for tokenizer, embedding models, and rerankers alongside the primary llm argument. You can configure these components in the same constructor call where you inject your custom LangChain LLM, allowing full customization of the retrieval and generation pipeline while maintaining the singleton pattern.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →