Simple Mode vs Professional Mode in Llama-GitHub Context Retrieval: What's the Difference?

Simple mode executes a single Google search via Jina without LLM processing, while professional mode runs a full RAG pipeline with LLM-driven query analysis and parallel searches across Google, GitHub code, issues, and repositories.

The GithubRAG class in the jetxu-llm/llama-github repository provides two distinct retrieval strategies through the retrieve_context and async_retrieve_context methods. Choosing between simple_mode and professional mode in context retrieval depends on your query complexity, latency requirements, and token budget.

Core Differences Between Simple and Professional Mode

The two modes differ fundamentally in search scope, LLM utilization, and performance characteristics:

  • Search Scope: Simple mode queries only Google via the Jina API. Professional mode concurrently searches Google, GitHub code, GitHub issues, and GitHub repositories.
  • LLM Integration: Simple mode bypasses the LLM entirely. Professional mode uses the LLM to tokenize queries, generate analysis strategies via analyze_question, and rank retrieved snippets.
  • Query Handling: Simple mode suits short queries (approximately ≤20 words) and may truncate longer inputs. Professional mode handles complex, multi-sentence queries through LLM segmentation.
  • Latency: Simple mode completes in a single network call. Professional mode requires multiple async API calls to GitHub and LLM inference.
  • Configuration: In llama_github/github_rag.py (lines 20-31), the simple_mode flag defaults to False at the instance level but can be overridden per method call.

Simple Mode Implementation

When simple_mode=True is passed to retrieve_context, the method creates only a Google-search task, gathers the results, and returns the top-N contexts without embedding generation or reranking.

In llama_github/github_rag.py (lines 34-53), the implementation constructs a single search task:

from llama_github import GithubRAG

rag = GithubRAG(github_access_token="ghp_…", openai_api_key="sk-…")
query = "What is the license of the pandas library?"

# Simple mode – only Google search, no LLM invocation

context = rag.retrieve_context(query, simple_mode=True)
print(context)

This mode avoids all LLM tokenization and embedding overhead, making it ideal for straightforward factual lookups where latency matters more than comprehensive source coverage.

Professional Mode Implementation

Professional mode (the default when simple_mode=False) executes a full RAG pipeline. According to the source in llama_github/github_rag.py (lines 54-78), the method first calls self.rag_processor.analyze_question to generate a retrieval strategy, then launches four parallel async tasks:

  1. Google search via Jina
  2. GitHub code search
  3. GitHub issues search
  4. GitHub repository search

After aggregation, the LLM ranks the snippets for relevance:

from llama_github import GithubRAG

rag = GithubRAG(github_access_token="ghp_…", openai_api_key="sk-…")
query = """How does pandas implement its groupby operation?
Can I customize it for large datasets?"""

# Professional mode – full analysis and multi-source retrieval (default)

context = rag.retrieve_context(query)  # implicit simple_mode=False

print(context)

The LLM-driven analysis enables the system to understand intent across complex, multi-part questions and synthesize evidence from disparate GitHub sources.

Asynchronous Usage for Both Modes

Both modes support async execution via async_retrieve_context. The same simple_mode parameter controls the strategy in asynchronous contexts:

import asyncio
from llama_github import GithubRAG

async def demo():
    rag = GithubRAG(github_access_token="ghp_…", openai_api_key="sk-…")
    query = "Explain pandas' DataFrame merge algorithm."
    
    # Simple async retrieval

    simple = await rag.async_retrieve_context(query, simple_mode=True)
    
    # Professional async retrieval (default)

    pro = await rag.async_retrieve_context(query)
    
    print("Simple:", simple)
    print("Professional:", pro)

asyncio.run(demo())

Configuration Guidelines and Documentation

The simple_mode parameter can be set at the class level during instantiation or overridden per call. According to docs/usage.md (lines 34-42), simple mode is recommended for short queries where a quick Google search suffices. The docs/api_reference.md (lines 34-36) formally documents the boolean flag for both synchronous and asynchronous methods.

Use simple mode when:

  • Queries are brief and unambiguous
  • You need minimal latency
  • LLM token costs must be avoided

Use professional mode when:

  • Queries are complex or multi-part
  • You require comprehensive GitHub source coverage (code, issues, repos)
  • Answer quality and source diversity outweigh latency concerns

Summary

  • Simple mode performs a single Google search via Jina without LLM processing, storing no embeddings and executing no reranking (as implemented in llama_github/github_rag.py, lines 34-53).
  • Professional mode executes a full RAG pipeline including LLM query analysis, parallel searches across four GitHub/Google sources, and neural ranking (as implemented in llama_github/github_rag.py, lines 54-78).
  • The simple_mode flag defaults to False but can be toggled per call or at the instance level.
  • Simple mode suits short queries (≤20 words); professional mode handles arbitrary complexity via LLM tokenization.

Frequently Asked Questions

Can I use simple mode for complex, multi-sentence queries?

Simple mode is designed for short queries of approximately 20 words or fewer. Longer queries may be truncated because the LLM is not consulted to split or analyze the input. For complex questions, use professional mode to leverage the LLM's query analysis capabilities.

Do I need to recreate the GithubRAG instance to switch modes?

No. The simple_mode parameter can be overridden on a per-call basis in both retrieve_context and async_retrieve_context. While the instance stores a default value (self.simple_mode in lines 20-31 of llama_github/github_rag.py), you can pass simple_mode=True or simple_mode=False to any individual method call.

Which mode is better for finding specific code implementations?

Professional mode is superior for code-specific questions because it executes dedicated GitHub code searches alongside issue and repository searches. Simple mode only queries Google, which may miss recent commits or specific implementation details found directly in source files.

Does professional mode cost more in terms of API usage?

Yes. Professional mode consumes LLM tokens for the initial analyze_question call and for ranking retrieved snippets. It also makes multiple concurrent API calls to GitHub's search endpoints. Simple mode incurs only the single Google search via Jina and no LLM costs, making it cheaper and faster for basic lookups.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →