How LLM-Powered Question Analysis Generates Search Strategies in Llama-GitHub
Llama-GitHub transforms raw developer questions into executable GitHub search strategies through a two-stage LLM pipeline that abstracts queries into structured logic before refining them into concrete search criteria.
The jetxu-llm/llama-github repository implements an intelligent retrieval-augmented generation (RAG) system that eliminates hard-coded heuristics by using large language models to interpret user intent. This article explores the exact mechanism of LLM-powered question analysis and how it produces actionable search strategies for code and issue discovery.
The Two-Stage Pipeline Overview
The system processes every query through a strictly defined workflow that separates high-level reasoning from concrete implementation. This architecture ensures that the LLM first understands the conceptual nature of the question before attempting to formulate specific search parameters.
Stage 1: Abstraction and High-Level Strategy
The entry point for all analysis is the RAGProcessor.analyze_question method in llama_github/rag_processing/rag_processor.py. When invoked, this method executes the first stage of the pipeline:
-
System Prompt Loading: The processor loads the
always_answer_promptfromllama_github/config/config.json. This prompt instructs the LLM to condense the input into a single-sentence abstraction, provide a concise answer with sample code, and generate high-level logic describing how to search for relevant GitHub code and issues. -
Async LLM Invocation: The
LLMHandler.ainvokemethod (located inllama_github/llm_integration/llm_handler.py, lines 29-66) constructs a LangChain chat prompt using the loaded system prompt, injects the user query as aHumanMessage, and executes the LLM asynchronously. The underlying model is managed byLLMManager(fromllama_github/llm_integration/initial_load.py), which supports OpenAI, Mistral, or local HuggingFace models based on configuration. -
Structured Output Parsing: The response must conform to the
_LLMFirstGenenralAnswerPydantic model (defined inrag_processor.py, lines 31-49). This schema enforces four specific fields:question: The one-sentence abstraction of the original queryanswer: A concise explanation with optional code samplescode_search_logic: Plain English describing the approach to finding relevant codeissue_search_logic: Plain English describing the approach to finding relevant issues
The method returns these four elements as a list [question, answer, code_search_logic, issue_search_logic] (see implementation in rag_processor.py, lines 61-66).
Stage 2: Concrete Search Criteria Generation
The high-level logic strings generated in Stage 1 are not directly executed as searches. Instead, they feed into secondary prompt-driven generators:
get_code_search_criteriaprocesses thecode_search_logicusing thecode_search_criteria_promptfromconfig.jsonget_issue_search_criteriaprocesses theissue_search_logicusing theissue_search_criteria_prompt
These secondary stages translate the abstract reasoning into exact GitHub search strings, filters, and repository targeting parameters.
Key Implementation Details
The RAGProcessor.analyze_question Method
The core analysis logic resides in llama_github/rag_processing/rag_processor.py. The analyze_question method orchestrates the initial transformation without executing any external searches. It relies entirely on the LLM's reasoning capabilities guided by the system prompt configuration.
Structured Output Schema
The _LLMFirstGenenralAnswer model (lines 31-49 of rag_processor.py) acts as a strict contract between the LLM and the application logic. By requiring the model to output structured JSON matching this schema, the system ensures predictable parsing of the abstraction, answer, and search strategies. This eliminates the need for regex parsing or fragile string manipulation of raw LLM outputs.
LLMHandler and Prompt Management
The LLMHandler class in llama_github/llm_integration/llm_handler.py abstracts away provider-specific implementations. It handles:
- Prompt templating and message composition
- Structured output binding via LangChain's
.with_structured_output()method - Async execution through
ainvoke
The LLMManager singleton (from initial_load.py) handles model selection and API key management, allowing the question analysis to work with OpenAI GPT-4, Mistral models, or quantized local models without changing the analysis logic.
Practical Code Examples
Direct Analysis with RAGProcessor
For granular control over the question analysis phase, instantiate RAGProcessor directly and call analyze_question:
import asyncio
from llama_github.rag_processing.rag_processor import RAGProcessor
from llama_github.data_retrieval.github_api import GitHubAPIHandler
async def demo():
# Initialise the GitHub API client (uses unauthenticated public endpoints here)
gh = GitHubAPIHandler()
# Create the RAG processor – it will spin up the default LLM manager
processor = RAGProcessor(github_api_handler=gh)
# Raw developer question
raw_query = "How can I efficiently compute the cosine similarity between two vectors using NumPy?"
# Run the LLM-driven analysis
question, answer, code_logic, issue_logic = await processor.analyze_question(raw_query)
print("🧠 Question abstraction:", question)
print("\n✅ Answer preview:", answer)
print("\n🔎 Code-search logic:", code_logic)
print("\n🐞 Issue-search logic:", issue_logic)
# Execute the demo
asyncio.run(demo())
Under the hood, RAGProcessor.__init__ creates an LLMManager instance, and analyze_question pulls the always_answer prompt before invoking LLMHandler.ainvoke. The response is automatically validated against the _LLMFirstGenenralAnswer schema.
Using the GitHubRag Wrapper
For standard use cases, the GitHubRag façade handles the full pipeline including the analysis step:
import asyncio
from llama_github.github_rag import GitHubRag
async def query_github():
rag = GitHubRag() # Instantiates RAGProcessor internally
result = await rag.ask(
query="Explain how to paginate results when calling the GitHub REST API."
)
print(result)
asyncio.run(query_github())
The GitHubRag.ask method internally calls self.rag_processor.analyze_question as the first operation in its retrieval chain. The generated search strategies determine which repositories, files, and issues the system subsequently fetches for context augmentation.
Summary
- LLM-powered question analysis in Llama-GitHub operates as a two-stage pipeline that first abstracts queries into structured logic, then refines them into concrete search criteria.
- The
RAGProcessor.analyze_questionmethod inrag_processor.pydrives the initial stage using thealways_answer_promptand the_LLMFirstGenenralAnswerschema. LLMHandler.ainvokemanages async LLM execution with structured output parsing, supporting multiple providers throughLLMManager.- The system generates four distinct outputs: a question abstraction, a concise answer, code search logic, and issue search logic.
- Downstream methods convert the high-level logic into executable GitHub search strings using secondary prompt templates from
config.json.
Frequently Asked Questions
What role does the always_answer_prompt play in the analysis?
The always_answer_prompt stored in llama_github/config/config.json provides the system instructions that guide the LLM's reasoning. It explicitly directs the model to produce a single-sentence question abstraction, a concise answer with code samples, and high-level descriptions of how to search for code and issues. This prompt eliminates the need for hard-coded query classification logic by encoding the reasoning strategy directly in natural language instructions.
How does the _LLMFirstGenenralAnswer schema ensure reliable outputs?
The _LLMFirstGenenralAnswer Pydantic model (defined in rag_processor.py, lines 31-49) acts as a strict output contract. By binding this schema to the LLM invocation through LangChain's structured output features, the system forces the model to return valid JSON with exactly four fields: question, answer, code_search_logic, and issue_search_logic. This structured approach prevents hallucinated fields or inconsistent formatting that would break downstream processing.
Can I switch LLM providers without modifying the analysis logic?
Yes. The LLMManager class in llama_github/llm_integration/initial_load.py abstracts provider-specific implementations. You can configure OpenAI, Mistral, or local HuggingFace models by setting the appropriate API keys or model paths in LLMManager.__init__. The RAGProcessor.analyze_question method remains unchanged regardless of which backend serves the LLM calls, as LLMHandler.ainvoke standardizes the interface across providers.
How do high-level search strategies become concrete GitHub queries?
After analyze_question returns the code_search_logic and issue_search_logic strings, these descriptions feed into secondary generation methods (such as get_code_search_criteria and get_issue_search_criteria). These methods use additional system prompts from config.json—specifically code_search_criteria_prompt and issue_search_criteria_prompt—to instruct the LLM to translate the abstract reasoning into specific GitHub search syntax, repository filters, and file patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →