# How to Use Custom LangChain LLM Objects with llama-github: A Complete Guide

> Learn to integrate custom LangChain LLM objects with llama-github. Inject your LLM into the LLMManager for seamless custom model integration and control.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: how-to-guide
- Published: 2026-03-04

---

**You can inject any LangChain-compatible LLM into llama-github by passing it to the `LLMManager` constructor, which stores the instance and automatically routes all downstream requests through your custom model instead of the default API-key initialization logic.**

The `llama-github` repository isolates all language-model interactions behind two core abstractions: the `LLMManager` singleton and the `LLMHandler` consumer. Because the handler never instantiates LLMs itself—instead fetching them exclusively via `LLMManager.get_llm()`—you can seamlessly integrate custom LangChain LLM objects such as `ChatOpenAI`, `ChatMistralAI`, or locally-hosted HuggingFace models by injecting them at initialization time.

## Architecture Overview

`llama-github` decouples model management from inference execution through two specialized components:

- **`LLMManager`** (defined in [`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py)): A singleton responsible for holding the actual LangChain LLM instance. When you supply a custom LLM via the `llm` parameter, the manager stores it internally, sets `model_type` to `"Custom_langchain_llm"`, and bypasses the built-in API-key initialization logic (lines 63-66).

- **`LLMHandler`** (defined in [`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py)): The inference interface that constructs LangChain `ChatPromptTemplate` objects, manages chat history and context injection, and executes calls via `await llm.ainvoke(...)`. The handler retrieves the active LLM exclusively through `LLMManager.get_llm()`, meaning it will automatically use whatever model instance the manager holds (lines 15-27).

This design ensures that any LangChain-compatible object implementing the standard chat interface (`ainvoke`, `invoke`, etc.) becomes immediately available to all downstream processing pipelines, including the RAG processor in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py).

## Step-by-Step Implementation

Follow these four steps to replace the default LLM with your custom LangChain object:

1. **Instantiate your LangChain LLM** using any provider (OpenAI, Anthropic, local HuggingFace, or custom subclasses).
2. **Create `LLMManager` with the custom LLM** passed to the `llm` argument.
3. **Initialize `LLMHandler`** with your configured manager instance.
4. **Invoke the handler** asynchronously with questions, chat history, and optional context.

### Creating Your Custom LangChain LLM

Begin by constructing any LangChain chat model that conforms to the base chat interface. For example, using OpenAI with non-default parameters:

```python
from langchain_openai import ChatOpenAI

custom_llm = ChatOpenAI(
    api_key="YOUR_OPENAI_API_KEY",
    model="gpt-4o-mini",
    temperature=0.9,
)

```

Or for local inference using HuggingFace:

```python
from langchain_community.chat_models import ChatHF

custom_llm = ChatHF(model="google/flan-t5-base", temperature=0.2)

```

### Injecting the LLM into LLMManager

Pass your instantiated LLM to the `LLMManager` constructor. According to the source code in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py), this assigns your object to the internal state and sets the model type identifier:

```python
from llama_github.llm_integration.initial_load import LLMManager

manager = LLMManager(llm=custom_llm)  # model_type automatically set to "Custom_langchain_llm"

```

### Initializing LLMHandler with Your Manager

Construct the handler by supplying the manager instance. The handler will now fetch your custom LLM for every inference request:

```python
from llama_github.llm_integration.llm_handler import LLMHandler

handler = LLMHandler(llm_manager=manager)

```

### Executing Requests

Use the asynchronous `ainvoke` method to process queries. The handler formats the prompt template, injects history and context, and calls your custom LLM's `ainvoke` method:

```python
import asyncio

async def ask():
    answer = await handler.ainvoke(
        human_question="Explain the difference between zero-shot and few-shot prompting.",
        chat_history=["User: What is LLM?", "Assistant: A large language model..."],
        context=["Relevant doc snippet..."],
    )
    print(answer)

asyncio.run(ask())

```

## Complete Code Examples

### Using OpenAI Chat Models

This example demonstrates integrating a custom `ChatOpenAI` instance with modified temperature settings:

```python
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler
from langchain_openai import ChatOpenAI
import asyncio

# Configure custom LangChain LLM

custom_llm = ChatOpenAI(
    api_key="YOUR_OPENAI_API_KEY",
    model="gpt-4o-mini",
    temperature=0.9,
)

# Inject into llama-github's manager

manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)

# Execute query

async def main():
    response = await handler.ainvoke(
        human_question="What are the advantages of retrieval-augmented generation?",
        chat_history=[],
        context=["RAG combines external knowledge retrieval with language generation..."]
    )
    print(response)

asyncio.run(main())

```

### Using Local HuggingFace Models

For offline or privacy-sensitive environments, pass a local HuggingFace chat model:

```python
from langchain_community.chat_models import ChatHF
from llama_github.llm_integration.initial_load import LLMManager
from llama_github.llm_integration.llm_handler import LLMHandler

# Local model configuration

custom_llm = ChatHF(model="google/flan-t5-base", temperature=0.2)

# Integration remains identical

manager = LLMManager(llm=custom_llm)
handler = LLMHandler(llm_manager=manager)

# All subsequent handler calls route to the local HF model

```

## Key Source Files and Implementation Details

Understanding these three files provides complete visibility into how custom LangChain LLM objects flow through the system:

| File | Purpose | Critical Implementation Detail |
|------|---------|-------------------------------|
| [`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py) | Defines `LLMManager` singleton | Lines 63-66 handle custom LLM storage and skip API-key initialization when `llm` parameter is provided |
| [`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py) | Defines `LLMHandler` inference interface | Lines 15-27 implement the `get_llm()` call that retrieves your custom object from the manager |
| [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) | High-level RAG orchestration | Demonstrates production usage pattern constructing both manager and handler for retrieval-augmented generation pipelines |

## Summary

- **`LLMManager`** acts as the single source of truth for LLM instances; passing a custom LangChain object to its constructor stores it as `"Custom_langchain_llm"` and disables default API-key logic.
- **`LLMHandler`** never creates LLM instances itself, instead querying the manager via `get_llm()`, ensuring automatic propagation of your custom model to all inference calls.
- Any LangChain-compatible chat model—including `ChatOpenAI`, `ChatMistralAI`, `ChatHF`, or bespoke subclasses—works immediately upon injection.
- The integration requires only three lines of configuration code: import `LLMManager`, instantiate with `llm=your_custom_model`, and pass to `LLMHandler`.

## Frequently Asked Questions

### Can I switch LLMs dynamically after initializing LLMManager?

No. According to the `llama-github` source code in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py), the `LLMManager` initializes the LLM instance during construction. To use a different model, you must create a new `LLMManager` instance with the alternative LLM and instantiate a fresh `LLMHandler` pointing to that new manager.

### Does llama-github support both synchronous and asynchronous invocation with custom LLMs?

Yes. While the examples emphasize `ainvoke()` for asynchronous execution, `LLMHandler` supports standard LangChain invocation patterns. Your custom LLM must implement both `invoke()` and `ainvoke()` methods to support both sync and async workflows within the handler's prompt execution pipeline.

### What happens if I don't provide a custom LLM to LLMManager?

If you instantiate `LLMManager` without the `llm` parameter, the class falls back to its built-in initialization logic, which attempts to configure default models using environment API keys. The source code in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py) (lines 63-66) explicitly checks for the presence of a custom LLM to determine whether to skip this default behavior.

### Can I use custom embedding or reranker models alongside my custom LLM?

Yes. `LLMManager` accepts additional parameters for tokenizer, embedding models, and rerankers alongside the primary `llm` argument. You can configure these components in the same constructor call where you inject your custom LangChain LLM, allowing full customization of the retrieval and generation pipeline while maintaining the singleton pattern.