# How Code Search Criteria Generation Using an LLM Works in Llama-GitHub

> Learn how the RAGProcessor class generates GitHub code search queries with LLM. Extract key concepts from user questions and create structured search strings with language qualifiers.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: deep-dive
- Published: 2026-03-04

---

**The `RAGProcessor` class generates GitHub code search queries by prompting an LLM to extract key concepts from user questions and output structured search strings with language qualifiers.**

The llama-github repository implements an intelligent RAG (Retrieval-Augmented Generation) system that leverages Large Language Models to optimize GitHub code discovery. By automating code search criteria generation using an LLM, the system translates natural language questions into precise GitHub Search API queries. This architecture dramatically improves the relevance of retrieved code snippets for downstream question-answering tasks.

## The Code Search Criteria Generation Pipeline

The orchestration logic resides in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) within the `RAGProcessor` class. When processing a user query, the system executes a four-step workflow to produce executable GitHub search strings.

### Step 1: Prompt Retrieval from Configuration

The processor loads the specialized `code_search_criteria_prompt` from [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json). This prompt instructs the model to analyze the user's question and any optional draft answer, extracting key technical concepts and formatting them as GitHub search queries. The instructions explicitly require the inclusion of `language:` qualifiers to ensure search precision.

### Step 2: Structured LLM Invocation

The `get_code_search_criteria` method calls `self.llm_handler.ainvoke` with four critical parameters:

- `human_question`: The original user query
- `prompt`: The retrieved `code_search_criteria_prompt`
- `context`: An optional draft answer providing additional context
- `output_structure`: The Pydantic model `_GitHubCodeSearchCriteria` (defined at lines 71-79)

This structured approach forces the LLM to return a validated list of search strings rather than free-form text.

### Step 3: Response Parsing and Validation

The `LLMHandler` (implemented in [`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py)) parses the model's output into the `_GitHubCodeSearchCriteria` schema, specifically extracting the `search_criteria` field containing 1-2 optimized search strings. If parsing fails or the LLM returns invalid data, the system returns an empty list and logs the error, ensuring graceful degradation.

### Step 4: Downstream API Consumption

The generated criteria strings flow to `GitHubRAG` in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py), which passes them to `GitHubAPIHandler` ([`llama_github/data_retrieval/github_api.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/data_retrieval/github_api.py)). This component executes actual HTTP requests to the GitHub Search API, retrieving matching code snippets that feed back into the RAG context for final answer generation.

## Key Architectural Components

Several specialized classes collaborate to enable LLM-driven search criteria generation:

- **`RAGProcessor`**: Core orchestrator managing LLM calls and context arrangement in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py)
- **`_GitHubCodeSearchCriteria`**: Pydantic schema enforcing structured output of 1-2 search strings (lines 71-79)
- **`LLMHandler`**: Async wrapper around the LLM provider handling prompting and output parsing in [`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py)
- **`GitHubAPIHandler`**: Executes REST calls to GitHub's search endpoints using the generated criteria in [`llama_github/data_retrieval/github_api.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/data_retrieval/github_api.py)

## Implementation Example: Generating Search Criteria

Here is a practical implementation demonstrating the criteria generation workflow:

```python
from llama_github.rag_processing.rag_processor import RAGProcessor
from llama_github.data_retrieval.github_api import GitHubAPIHandler

# Initialise dependencies

github_api = GitHubAPIHandler()
rag = RAGProcessor(github_api_handler=github_api)

# Example user query

question = "How can I efficiently read a large CSV file with Pandas?"

# Optional draft answer that the LLM can use for extra context

draft = """You can use pandas.read_csv with the `chunksize` parameter to stream the file."""
    

# Generate search criteria (async)

criteria = await rag.get_code_search_criteria(question, draft_answer=draft)

print(criteria)

# Example output:

# [

#   "pandas read_csv chunksize language:python",

#   "large csv processing pandas language:python"

# ]

```

The returned list contains ready-to-use query strings compatible with the GitHub code search endpoint (e.g., `https://api.github.com/search/code?q=pandas+read_csv+chunksize+language:python`).

## Summary

- The `RAGProcessor` class in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) orchestrates the entire code search criteria generation pipeline
- Configuration-driven prompts from [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json) guide the LLM to extract technical concepts and enforce `language:` qualifiers
- Structured output is enforced using the `_GitHubCodeSearchCriteria` Pydantic model, ensuring the LLM returns 1-2 valid search strings
- The `LLMHandler.ainvoke` method handles async communication with the LLM provider and response validation
- Generated criteria flow through `GitHubRAG` to `GitHubAPIHandler`, which executes actual GitHub API searches to retrieve relevant code snippets

## Frequently Asked Questions

### What prompt does the LLM use to generate search criteria?

The system loads the `code_search_criteria_prompt` from [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json). This specialized prompt instructs the model to analyze the user's question, identify key technical concepts, and output GitHub search strings that always include a `language:` qualifier for precision.

### How does the system ensure the LLM returns structured search criteria?

The `get_code_search_criteria` method passes the `_GitHubCodeSearchCriteria` Pydantic model (defined at lines 71-79 of [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py)) as the `output_structure` parameter to `LLMHandler.ainvoke`. This forces the LLM to return a validated JSON object containing a list of search strings rather than unstructured text.

### Which component executes the actual GitHub search using the generated criteria?

The `GitHubAPIHandler` class in [`llama_github/data_retrieval/github_api.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/data_retrieval/github_api.py) receives the generated criteria strings and executes HTTP requests to the GitHub Search API. This class is invoked by `GitHubRAG` in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) as part of the broader RAG pipeline.

### Can the code search criteria generation work without a draft answer?

Yes. The `draft_answer` parameter in `get_code_search_criteria` is optional. When provided, it supplies additional context to help the LLM generate more specific search terms, but the system functions effectively using only the `human_question` parameter if no draft is available.