# How Llama-GitHub Implements Context Relevance Scoring Using LLM

> Discover how Llama-GitHub uses LLMs for context relevance scoring. It combines LLM output, vector similarity, and rerank scores for a 0-100 relevance ranking.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: deep-dive
- Published: 2026-03-04

---

**Llama-GitHub evaluates code snippet relevance by invoking a lightweight LLM with a structured scoring rubric, then combines the output with vector similarity and rerank scores to produce a final 0-100 relevance ranking.**

Llama-GitHub enhances its retrieval-augmented generation (RAG) pipeline by using a language model to judge how well candidate code snippets answer a developer's query. This **context relevance scoring** layer adds semantic evaluation beyond pure vector similarity, ensuring the final context contains genuinely useful code rather than just lexically similar text. The implementation spans the `RagProcessor` class, the `LLMHandler` module, and a configurable scoring prompt stored in [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json).

## The Five-Step Scoring Pipeline

The context relevance scoring mechanism operates as a secondary filter after initial retrieval, using an LLM to apply nuanced judgment to the top candidate snippets.

### 1. Candidate Selection via Hybrid Retrieval

Before the LLM evaluates anything, the system narrows the candidate pool using traditional retrieval signals. The `RagProcessor` first fetches raw files from GitHub, then filters them through **embedding similarity** and a **Jina reranker** to surface the most promising snippets. This cosine-similarity loop reduces the volume of data sent to the expensive LLM evaluation step.

### 2. LLM Invocation with Structured Prompts

For each top-ranked candidate, `RagProcessor.get_context_relevance_score` constructs a **system prompt** from `config["scoring_context_prompt"]` and sends the user query alongside the candidate snippet to a lightweight LLM retrieved via `self.llm_manager.get_llm_simple()`. To maintain pipeline performance, these calls execute asynchronously and in parallel using `asyncio.gather`, as implemented in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) on lines 451-456.

### 3. Forcing Structured Integer Output

The processor enforces strict output formatting through a Pydantic model named `_ContextRelevanceScore`, which defines a single integer field `score`. The `LLMHandler.ainvoke` method receives `output_structure=self._ContextRelevanceScore` and wraps the model in `with_structured_output`, instructing LangChain to parse the response into this schema. The prompt explicitly directs the model to return **only** the integer value, eliminating parsing ambiguity. This logic resides in [`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py) on lines 29-35 and 84-90, with prompt selection logic on lines 51-54.

### 4. The 0-100 Scoring Rubric

The relevance judgment relies on a detailed rubric stored in [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json) under the key `scoring_context_prompt`. This configuration defines five distinct relevance bands—**0-20** (irrelevant), **21-40** (poor), **41-60** (moderate), **61-80** (good), and **81-100** (excellent)—and explicitly instructs the model to output a single integer within this range. The processor retrieves this prompt at runtime using `config.get("scoring_context_prompt")` as shown on lines 93-94 of [`rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/rag_processor.py).

### 5. Score Aggregation for Final Ranking

After collecting LLM scores, the processor combines multiple signals to determine final relevance. It multiplies the **LLM relevance score** by the **cosine similarity** and **rerank score** to obtain a combined relevance value, as seen in the list comprehension on lines 57-60 of [`rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/rag_processor.py). The highest-scoring snippets from this weighted calculation become the final context passed to the answer generation model.

## Key Implementation Files

- **[`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py)** – Orchestrates end-to-end retrieval, embedding, reranking, and LLM relevance scoring.
- **[`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py)** – Wraps LangChain calls, selects the "simple" LLM, and parses structured output using Pydantic models.
- **[`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json)** – Stores the `scoring_context_prompt` that defines the relevance rubric used by the LLM.
- **[`llama_github/config/config.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.py)** – Provides singleton access to configuration values, including the scoring prompt.

## Practical Code Examples

The following examples demonstrate how to invoke the scoring logic directly or as part of the full RAG retrieval flow.

```python

# Example: Manually score a single GitHub snippet

from llama_github.rag_processing.rag_processor import RagProcessor

async def demo():
    processor = RagProcessor()
    query = "How can I list all open issues for a repo using PyGitHub?"
    snippet = """
    from github import Github
    g = Github("<TOKEN>")
    repo = g.get_repo("owner/repo")
    for issue in repo.get_issues(state="open"):
        print(issue.title)
    """
    # Returns an integer 0-100

    score = await processor.get_context_relevance_score(query, snippet)
    print(f"Relevance score: {score}")

# Run in an async environment

# asyncio.run(demo())

```

```python

# Example: Using the full pipeline (top-N contexts)

from llama_github.rag_processing.rag_processor import RagProcessor

async def top_contexts_demo():
    processor = RagProcessor()
    query = "How do I authenticate a GitHub API request with a token?"
    top_contexts = await processor.get_top_n_contexts(query, answer=None)
    for ctx in top_contexts:
        print(ctx["source_file"], "-> score", ctx["relevance_score"])

```

## Summary

- **Hybrid filtering** first reduces candidates via vector similarity and a Jina reranker before LLM evaluation.
- **Structured output** enforcement via Pydantic models ensures the LLM returns only valid integers between 0 and 100.
- **Configurable rubrics** allow customization of relevance criteria through the `scoring_context_prompt` in [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json).
- **Parallel execution** using `asyncio.gather` maintains pipeline efficiency despite LLM latency.
- **Multi-signal ranking** combines LLM scores with cosine similarity and rerank scores for robust final rankings.

## Frequently Asked Questions

### What LLM does Llama-GitHub use for context relevance scoring?

The system uses a lightweight model referred to as the "simple" LLM, accessed via `self.llm_manager.get_llm_simple()`. This is typically a smaller, faster model than the main generation LLM, optimizing for evaluation speed while maintaining sufficient judgment capability for relevance scoring.

### How does the scoring pipeline handle multiple candidates efficiently?

All relevance evaluations run asynchronously and in parallel using `asyncio.gather` within `RagProcessor.get_context_relevance_score`. This concurrent execution prevents sequential LLM calls from bottlenecking the retrieval pipeline when processing multiple code snippets.

### Can the relevance scoring rubric be customized?

Yes. The scoring criteria are defined in [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json) under the `scoring_context_prompt` key. Developers can modify this configuration to adjust the 0-100 scoring bands or refine the instructions given to the evaluating LLM without changing the source code.

### How is the final context ranking calculated?

The final ranking derives from a combined score calculation: the **LLM relevance score** (0-100) is multiplied by the **cosine similarity** score and the **reranker score** for each snippet. This multiplicative approach ensures that high LLM ratings amplify strong vector matches while suppressing false positives from any single signal.