How LLM-Powered Question Analysis Generates Search Strategies in Llama-GitHub

Llama-GitHub transforms raw developer questions into executable GitHub search strategies through a two-stage LLM pipeline that abstracts queries into structured logic before refining them into concrete search criteria.

The jetxu-llm/llama-github repository implements an intelligent retrieval-augmented generation (RAG) system that eliminates hard-coded heuristics by using large language models to interpret user intent. This article explores the exact mechanism of LLM-powered question analysis and how it produces actionable search strategies for code and issue discovery.

The Two-Stage Pipeline Overview

The system processes every query through a strictly defined workflow that separates high-level reasoning from concrete implementation. This architecture ensures that the LLM first understands the conceptual nature of the question before attempting to formulate specific search parameters.

Stage 1: Abstraction and High-Level Strategy

The entry point for all analysis is the RAGProcessor.analyze_question method in llama_github/rag_processing/rag_processor.py. When invoked, this method executes the first stage of the pipeline:

  1. System Prompt Loading: The processor loads the always_answer_prompt from llama_github/config/config.json. This prompt instructs the LLM to condense the input into a single-sentence abstraction, provide a concise answer with sample code, and generate high-level logic describing how to search for relevant GitHub code and issues.

  2. Async LLM Invocation: The LLMHandler.ainvoke method (located in llama_github/llm_integration/llm_handler.py, lines 29-66) constructs a LangChain chat prompt using the loaded system prompt, injects the user query as a HumanMessage, and executes the LLM asynchronously. The underlying model is managed by LLMManager (from llama_github/llm_integration/initial_load.py), which supports OpenAI, Mistral, or local HuggingFace models based on configuration.

  3. Structured Output Parsing: The response must conform to the _LLMFirstGenenralAnswer Pydantic model (defined in rag_processor.py, lines 31-49). This schema enforces four specific fields:

    • question: The one-sentence abstraction of the original query
    • answer: A concise explanation with optional code samples
    • code_search_logic: Plain English describing the approach to finding relevant code
    • issue_search_logic: Plain English describing the approach to finding relevant issues

The method returns these four elements as a list [question, answer, code_search_logic, issue_search_logic] (see implementation in rag_processor.py, lines 61-66).

Stage 2: Concrete Search Criteria Generation

The high-level logic strings generated in Stage 1 are not directly executed as searches. Instead, they feed into secondary prompt-driven generators:

  • get_code_search_criteria processes the code_search_logic using the code_search_criteria_prompt from config.json
  • get_issue_search_criteria processes the issue_search_logic using the issue_search_criteria_prompt

These secondary stages translate the abstract reasoning into exact GitHub search strings, filters, and repository targeting parameters.

Key Implementation Details

The RAGProcessor.analyze_question Method

The core analysis logic resides in llama_github/rag_processing/rag_processor.py. The analyze_question method orchestrates the initial transformation without executing any external searches. It relies entirely on the LLM's reasoning capabilities guided by the system prompt configuration.

Structured Output Schema

The _LLMFirstGenenralAnswer model (lines 31-49 of rag_processor.py) acts as a strict contract between the LLM and the application logic. By requiring the model to output structured JSON matching this schema, the system ensures predictable parsing of the abstraction, answer, and search strategies. This eliminates the need for regex parsing or fragile string manipulation of raw LLM outputs.

LLMHandler and Prompt Management

The LLMHandler class in llama_github/llm_integration/llm_handler.py abstracts away provider-specific implementations. It handles:

  • Prompt templating and message composition
  • Structured output binding via LangChain's .with_structured_output() method
  • Async execution through ainvoke

The LLMManager singleton (from initial_load.py) handles model selection and API key management, allowing the question analysis to work with OpenAI GPT-4, Mistral models, or quantized local models without changing the analysis logic.

Practical Code Examples

Direct Analysis with RAGProcessor

For granular control over the question analysis phase, instantiate RAGProcessor directly and call analyze_question:

import asyncio
from llama_github.rag_processing.rag_processor import RAGProcessor
from llama_github.data_retrieval.github_api import GitHubAPIHandler

async def demo():
    # Initialise the GitHub API client (uses unauthenticated public endpoints here)

    gh = GitHubAPIHandler()
    # Create the RAG processor – it will spin up the default LLM manager

    processor = RAGProcessor(github_api_handler=gh)

    # Raw developer question

    raw_query = "How can I efficiently compute the cosine similarity between two vectors using NumPy?"
    
    # Run the LLM-driven analysis

    question, answer, code_logic, issue_logic = await processor.analyze_question(raw_query)

    print("🧠 Question abstraction:", question)
    print("\n✅ Answer preview:", answer)
    print("\n🔎 Code-search logic:", code_logic)
    print("\n🐞 Issue-search logic:", issue_logic)

# Execute the demo

asyncio.run(demo())

Under the hood, RAGProcessor.__init__ creates an LLMManager instance, and analyze_question pulls the always_answer prompt before invoking LLMHandler.ainvoke. The response is automatically validated against the _LLMFirstGenenralAnswer schema.

Using the GitHubRag Wrapper

For standard use cases, the GitHubRag façade handles the full pipeline including the analysis step:

import asyncio
from llama_github.github_rag import GitHubRag

async def query_github():
    rag = GitHubRag()               # Instantiates RAGProcessor internally

    result = await rag.ask(
        query="Explain how to paginate results when calling the GitHub REST API."
    )
    print(result)

asyncio.run(query_github())

The GitHubRag.ask method internally calls self.rag_processor.analyze_question as the first operation in its retrieval chain. The generated search strategies determine which repositories, files, and issues the system subsequently fetches for context augmentation.

Summary

  • LLM-powered question analysis in Llama-GitHub operates as a two-stage pipeline that first abstracts queries into structured logic, then refines them into concrete search criteria.
  • The RAGProcessor.analyze_question method in rag_processor.py drives the initial stage using the always_answer_prompt and the _LLMFirstGenenralAnswer schema.
  • LLMHandler.ainvoke manages async LLM execution with structured output parsing, supporting multiple providers through LLMManager.
  • The system generates four distinct outputs: a question abstraction, a concise answer, code search logic, and issue search logic.
  • Downstream methods convert the high-level logic into executable GitHub search strings using secondary prompt templates from config.json.

Frequently Asked Questions

What role does the always_answer_prompt play in the analysis?

The always_answer_prompt stored in llama_github/config/config.json provides the system instructions that guide the LLM's reasoning. It explicitly directs the model to produce a single-sentence question abstraction, a concise answer with code samples, and high-level descriptions of how to search for code and issues. This prompt eliminates the need for hard-coded query classification logic by encoding the reasoning strategy directly in natural language instructions.

How does the _LLMFirstGenenralAnswer schema ensure reliable outputs?

The _LLMFirstGenenralAnswer Pydantic model (defined in rag_processor.py, lines 31-49) acts as a strict output contract. By binding this schema to the LLM invocation through LangChain's structured output features, the system forces the model to return valid JSON with exactly four fields: question, answer, code_search_logic, and issue_search_logic. This structured approach prevents hallucinated fields or inconsistent formatting that would break downstream processing.

Can I switch LLM providers without modifying the analysis logic?

Yes. The LLMManager class in llama_github/llm_integration/initial_load.py abstracts provider-specific implementations. You can configure OpenAI, Mistral, or local HuggingFace models by setting the appropriate API keys or model paths in LLMManager.__init__. The RAGProcessor.analyze_question method remains unchanged regardless of which backend serves the LLM calls, as LLMHandler.ainvoke standardizes the interface across providers.

How do high-level search strategies become concrete GitHub queries?

After analyze_question returns the code_search_logic and issue_search_logic strings, these descriptions feed into secondary generation methods (such as get_code_search_criteria and get_issue_search_criteria). These methods use additional system prompts from config.json—specifically code_search_criteria_prompt and issue_search_criteria_prompt—to instruct the LLM to translate the abstract reasoning into specific GitHub search syntax, repository filters, and file patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →