How Llama-GitHub Implements Context Relevance Scoring Using LLM

Llama-GitHub evaluates code snippet relevance by invoking a lightweight LLM with a structured scoring rubric, then combines the output with vector similarity and rerank scores to produce a final 0-100 relevance ranking.

Llama-GitHub enhances its retrieval-augmented generation (RAG) pipeline by using a language model to judge how well candidate code snippets answer a developer's query. This context relevance scoring layer adds semantic evaluation beyond pure vector similarity, ensuring the final context contains genuinely useful code rather than just lexically similar text. The implementation spans the RagProcessor class, the LLMHandler module, and a configurable scoring prompt stored in config.json.

The Five-Step Scoring Pipeline

The context relevance scoring mechanism operates as a secondary filter after initial retrieval, using an LLM to apply nuanced judgment to the top candidate snippets.

1. Candidate Selection via Hybrid Retrieval

Before the LLM evaluates anything, the system narrows the candidate pool using traditional retrieval signals. The RagProcessor first fetches raw files from GitHub, then filters them through embedding similarity and a Jina reranker to surface the most promising snippets. This cosine-similarity loop reduces the volume of data sent to the expensive LLM evaluation step.

2. LLM Invocation with Structured Prompts

For each top-ranked candidate, RagProcessor.get_context_relevance_score constructs a system prompt from config["scoring_context_prompt"] and sends the user query alongside the candidate snippet to a lightweight LLM retrieved via self.llm_manager.get_llm_simple(). To maintain pipeline performance, these calls execute asynchronously and in parallel using asyncio.gather, as implemented in llama_github/rag_processing/rag_processor.py on lines 451-456.

3. Forcing Structured Integer Output

The processor enforces strict output formatting through a Pydantic model named _ContextRelevanceScore, which defines a single integer field score. The LLMHandler.ainvoke method receives output_structure=self._ContextRelevanceScore and wraps the model in with_structured_output, instructing LangChain to parse the response into this schema. The prompt explicitly directs the model to return only the integer value, eliminating parsing ambiguity. This logic resides in llama_github/llm_integration/llm_handler.py on lines 29-35 and 84-90, with prompt selection logic on lines 51-54.

4. The 0-100 Scoring Rubric

The relevance judgment relies on a detailed rubric stored in llama_github/config/config.json under the key scoring_context_prompt. This configuration defines five distinct relevance bands—0-20 (irrelevant), 21-40 (poor), 41-60 (moderate), 61-80 (good), and 81-100 (excellent)—and explicitly instructs the model to output a single integer within this range. The processor retrieves this prompt at runtime using config.get("scoring_context_prompt") as shown on lines 93-94 of rag_processor.py.

5. Score Aggregation for Final Ranking

After collecting LLM scores, the processor combines multiple signals to determine final relevance. It multiplies the LLM relevance score by the cosine similarity and rerank score to obtain a combined relevance value, as seen in the list comprehension on lines 57-60 of rag_processor.py. The highest-scoring snippets from this weighted calculation become the final context passed to the answer generation model.

Key Implementation Files

Practical Code Examples

The following examples demonstrate how to invoke the scoring logic directly or as part of the full RAG retrieval flow.


# Example: Manually score a single GitHub snippet

from llama_github.rag_processing.rag_processor import RagProcessor

async def demo():
    processor = RagProcessor()
    query = "How can I list all open issues for a repo using PyGitHub?"
    snippet = """
    from github import Github
    g = Github("<TOKEN>")
    repo = g.get_repo("owner/repo")
    for issue in repo.get_issues(state="open"):
        print(issue.title)
    """
    # Returns an integer 0-100

    score = await processor.get_context_relevance_score(query, snippet)
    print(f"Relevance score: {score}")

# Run in an async environment

# asyncio.run(demo())

# Example: Using the full pipeline (top-N contexts)

from llama_github.rag_processing.rag_processor import RagProcessor

async def top_contexts_demo():
    processor = RagProcessor()
    query = "How do I authenticate a GitHub API request with a token?"
    top_contexts = await processor.get_top_n_contexts(query, answer=None)
    for ctx in top_contexts:
        print(ctx["source_file"], "-> score", ctx["relevance_score"])

Summary

  • Hybrid filtering first reduces candidates via vector similarity and a Jina reranker before LLM evaluation.
  • Structured output enforcement via Pydantic models ensures the LLM returns only valid integers between 0 and 100.
  • Configurable rubrics allow customization of relevance criteria through the scoring_context_prompt in config.json.
  • Parallel execution using asyncio.gather maintains pipeline efficiency despite LLM latency.
  • Multi-signal ranking combines LLM scores with cosine similarity and rerank scores for robust final rankings.

Frequently Asked Questions

What LLM does Llama-GitHub use for context relevance scoring?

The system uses a lightweight model referred to as the "simple" LLM, accessed via self.llm_manager.get_llm_simple(). This is typically a smaller, faster model than the main generation LLM, optimizing for evaluation speed while maintaining sufficient judgment capability for relevance scoring.

How does the scoring pipeline handle multiple candidates efficiently?

All relevance evaluations run asynchronously and in parallel using asyncio.gather within RagProcessor.get_context_relevance_score. This concurrent execution prevents sequential LLM calls from bottlenecking the retrieval pipeline when processing multiple code snippets.

Can the relevance scoring rubric be customized?

Yes. The scoring criteria are defined in llama_github/config/config.json under the scoring_context_prompt key. Developers can modify this configuration to adjust the 0-100 scoring bands or refine the instructions given to the evaluating LLM without changing the source code.

How is the final context ranking calculated?

The final ranking derives from a combined score calculation: the LLM relevance score (0-100) is multiplied by the cosine similarity score and the reranker score for each snippet. This multiplicative approach ensures that high LLM ratings amplify strong vector matches while suppressing false positives from any single signal.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →