# How to Configure Embedding and Reranking Models in Llama-GitHub for Better Retrieval Accuracy

> Improve retrieval accuracy in llama-github by learning how to configure embedding and reranking models. Update config.json or use LLMManager for optimal results.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: how-to-guide
- Published: 2026-03-04

---

**To configure embedding and reranking models in llama-github, update the `default_embedding` and `default_reranker` keys in [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json) or pass explicit model IDs to the `LLMManager` constructor, ensuring `simple_mode=False` to load the models.**

The llama-github repository implements a **two-stage retrieval pipeline** that combines dense vector similarity with cross-encoder reranking to improve code context retrieval. According to the source code in [`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py), the `LLMManager` class centralizes model instantiation and loads default identifiers from [`config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config/config.json). Configuring these models correctly is essential for achieving high-precision retrieval when processing GitHub repositories.

## Understanding the Two-Stage Retrieval Pipeline

The `RAGProcessor` class orchestrates retrieval through two distinct scoring mechanisms:

1. **Reranking stage** – A sentence-pair classifier (`AutoModelForSequenceClassification`) scores how well each retrieved chunk matches the query. This happens in `RAGProcessor.retrieve_topn_contexts` via `reranker.compute_score(sentence_pairs)`.

2. **Embedding similarity stage** – A dense vector model (`AutoModel`) computes cosine similarity between the query (plus optional answer) and each candidate chunk using `embedding_model.encode()`.

The final ranking combines the reranker score, cosine similarity, and a lightweight LLM relevance score to select the top contexts. Both models are lazily loaded by `LLMManager` when `simple_mode` is disabled.

## Configuring Models via config.json

The simplest way to configure embedding and reranking models is by editing the configuration file at [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json). The `LLMManager` reads the keys `default_embedding` and `default_reranker` via `config.get()` calls (see lines 38-41 in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py)).

Update the JSON values to use domain-specific models for better code retrieval:

```json
{
  "default_embedding": "jinaai/jina-embeddings-v2-base-code",
  "default_reranker": "jinaai/jina-reranker-v2-base-multilingual",
  "top_n_contexts": 5
}

```

**Recommended model choices:**
- **Embedding models**: `jinaai/jina-embeddings-v2-base-code` (code-optimized), `sentence-transformers/all-MiniLM-L6-v2` (general purpose)
- **Rerankers**: `jinaai/jina-reranker-v2-base-multilingual` (multilingual code), `cross-encoder/ms-marco-MiniLM-L-12-v2` (MS MARCO trained)

Higher-quality embeddings provide more accurate cosine similarity calculations, while stronger rerankers improve the initial ranking before the embedding refinement stage.

## Runtime Configuration with LLMManager

For per-run experimentation without modifying repository files, instantiate `LLMManager` with explicit model parameters. This bypasses the JSON defaults and allows A/B testing different model combinations.

```python
from llama_github.llm_integration.initial_load import LLMManager

# Override defaults at runtime

manager = LLMManager(
    embedding_model="sentence-transformers/all-MiniLM-L6-v2",
    rerank_model="cross-encoder/ms-marco-MiniLM-L-12-v2",
    simple_mode=False  # Required to load models

)

# Use with RAGProcessor

from llama_github.rag_processing.rag_processor import RAGProcessor
from llama_github.data_retrieval.github_api import GitHubAPIHandler

github_api = GitHubAPIHandler(token="your_github_token")
rag_processor = RAGProcessor(
    github_api_handler=github_api,
    llm_manager=manager
)

```

The constructor parameters directly set the model identifiers that `LLMManager` passes to `AutoModel.from_pretrained()` and `AutoModelForSequenceClassification.from_pretrained()`.

## Critical Configuration Requirements

### Disable Simple Mode

You must ensure `simple_mode=False` (the default) when initializing `LLMManager`. When `simple_mode=True`, the initialization code at lines 93-95 skips both model instantiations, leaving `embedding_model` and `rerank_model` as `None`. This breaks the ranking pipeline entirely.

```python

# Correct initialization

manager = LLMManager(simple_mode=False)

# This will fail to rank contexts

manager = LLMManager(simple_mode=True)  # Avoid this

```

### Adjust Retrieval Parameters

Tune the `top_n_contexts` parameter (default: 4) to control precision versus recall:

- **Higher values** (8-10): Increase recall by including more candidate chunks
- **Lower values** (2-3): Improve precision and reduce latency by keeping only the highest-confidence contexts

Set this in [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json) or pass it through your `RAGProcessor` initialization chain.

## How Model Selection Affects Retrieval Accuracy

The source code reveals that model quality directly impacts three scoring components in `retrieve_topn_contexts`:

**Embedding models** determine semantic matching quality through cosine similarity. Code-specific embeddings (e.g., trained on GitHub repositories) better capture programming language semantics than general text models.

**Reranking models** function as binary classifiers (`num_labels=1`) that score query-chunk relevance. Models fine-tuned on code relevance datasets provide more accurate initial rankings before the embedding similarity calculation.

**Combined scoring** multiplies the reranker score with embedding similarity and LLM relevance scores. Weak performance in either the embedding or reranking stage propagates to the final context selection, reducing the quality of downstream LLM generation.

## Summary

- **Edit [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json)** to change the default `default_embedding` and `default_reranker` model identifiers permanently
- **Pass explicit parameters** to `LLMManager(embedding_model=..., rerank_model=...)` for temporary overrides
- **Always set `simple_mode=False`** to ensure models are actually loaded (lines 93-95 in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py))
- **Choose code-specific models** like `jinaai/jina-embeddings-v2-base-code` for better semantic matching of programming languages
- **Tune `top_n_contexts`** to balance between retrieval recall and precision

## Frequently Asked Questions

### What file controls the default embedding and reranking models?

The defaults are stored in [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json) under the keys `default_embedding` and `default_reranker`. The `LLMManager` class reads these values at initialization via the `Config` singleton implemented in [`llama_github/config/config.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.py).

### Can I use different models for different queries without restarting?

Yes. Create a new `LLMManager` instance with different `embedding_model` and `rerank_model` parameters for each query batch. Since models are loaded lazily, you can instantiate multiple managers with different configurations in the same Python process to compare retrieval quality across model combinations.

### Why is my retrieval pipeline returning unranked results?

Check that `simple_mode=False` in your `LLMManager` initialization. When `simple_mode=True`, the code at lines 93-95 skips model loading, leaving both `embedding_model` and `rerank_model` as `None`. Without these models, `RAGProcessor` cannot compute similarity scores or rerank contexts, resulting in random or unranked output.

### How do I know which embedding model is best for code retrieval?

Select models explicitly trained on code corpora. According to the implementation in [`initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/initial_load.py), the embedding model generates vectors for cosine similarity comparisons in `RAGProcessor`. Models like `jinaai/jina-embeddings-v2-base-code` capture programming language syntax and semantics better than general text models, resulting in more accurate similarity scores for code-related queries.