# Simple Mode vs Professional Mode in Llama-GitHub Context Retrieval: What's the Difference?

> Understand the difference between simple and professional modes in Llama-GitHub context retrieval. Simple mode uses Google search, while professional mode employs a full RAG pipeline for comprehensive data analysis.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: deep-dive
- Published: 2026-03-04

---

**Simple mode executes a single Google search via Jina without LLM processing, while professional mode runs a full RAG pipeline with LLM-driven query analysis and parallel searches across Google, GitHub code, issues, and repositories.**

The `GithubRAG` class in the `jetxu-llm/llama-github` repository provides two distinct retrieval strategies through the `retrieve_context` and `async_retrieve_context` methods. Choosing between **simple_mode and professional mode in context retrieval** depends on your query complexity, latency requirements, and token budget.

## Core Differences Between Simple and Professional Mode

The two modes differ fundamentally in search scope, LLM utilization, and performance characteristics:

- **Search Scope**: Simple mode queries only Google via the Jina API. Professional mode concurrently searches Google, GitHub code, GitHub issues, and GitHub repositories.
- **LLM Integration**: Simple mode bypasses the LLM entirely. Professional mode uses the LLM to tokenize queries, generate analysis strategies via `analyze_question`, and rank retrieved snippets.
- **Query Handling**: Simple mode suits short queries (approximately ≤20 words) and may truncate longer inputs. Professional mode handles complex, multi-sentence queries through LLM segmentation.
- **Latency**: Simple mode completes in a single network call. Professional mode requires multiple async API calls to GitHub and LLM inference.
- **Configuration**: In [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) (lines 20-31), the `simple_mode` flag defaults to `False` at the instance level but can be overridden per method call.

## Simple Mode Implementation

When `simple_mode=True` is passed to `retrieve_context`, the method creates only a Google-search task, gathers the results, and returns the top-N contexts without embedding generation or reranking.

In [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) (lines 34-53), the implementation constructs a single search task:

```python
from llama_github import GithubRAG

rag = GithubRAG(github_access_token="ghp_…", openai_api_key="sk-…")
query = "What is the license of the pandas library?"

# Simple mode – only Google search, no LLM invocation

context = rag.retrieve_context(query, simple_mode=True)
print(context)

```

This mode avoids all LLM tokenization and embedding overhead, making it ideal for straightforward factual lookups where latency matters more than comprehensive source coverage.

## Professional Mode Implementation

Professional mode (the default when `simple_mode=False`) executes a full RAG pipeline. According to the source in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) (lines 54-78), the method first calls `self.rag_processor.analyze_question` to generate a retrieval strategy, then launches four parallel async tasks:

1. Google search via Jina
2. GitHub code search
3. GitHub issues search  
4. GitHub repository search

After aggregation, the LLM ranks the snippets for relevance:

```python
from llama_github import GithubRAG

rag = GithubRAG(github_access_token="ghp_…", openai_api_key="sk-…")
query = """How does pandas implement its groupby operation?
Can I customize it for large datasets?"""

# Professional mode – full analysis and multi-source retrieval (default)

context = rag.retrieve_context(query)  # implicit simple_mode=False

print(context)

```

The LLM-driven analysis enables the system to understand intent across complex, multi-part questions and synthesize evidence from disparate GitHub sources.

## Asynchronous Usage for Both Modes

Both modes support async execution via `async_retrieve_context`. The same `simple_mode` parameter controls the strategy in asynchronous contexts:

```python
import asyncio
from llama_github import GithubRAG

async def demo():
    rag = GithubRAG(github_access_token="ghp_…", openai_api_key="sk-…")
    query = "Explain pandas' DataFrame merge algorithm."
    
    # Simple async retrieval

    simple = await rag.async_retrieve_context(query, simple_mode=True)
    
    # Professional async retrieval (default)

    pro = await rag.async_retrieve_context(query)
    
    print("Simple:", simple)
    print("Professional:", pro)

asyncio.run(demo())

```

## Configuration Guidelines and Documentation

The `simple_mode` parameter can be set at the class level during instantiation or overridden per call. According to [`docs/usage.md`](https://github.com/jetxu-llm/llama-github/blob/main/docs/usage.md) (lines 34-42), simple mode is recommended for short queries where a quick Google search suffices. The [`docs/api_reference.md`](https://github.com/jetxu-llm/llama-github/blob/main/docs/api_reference.md) (lines 34-36) formally documents the boolean flag for both synchronous and asynchronous methods.

Use **simple mode** when:
- Queries are brief and unambiguous
- You need minimal latency
- LLM token costs must be avoided

Use **professional mode** when:
- Queries are complex or multi-part
- You require comprehensive GitHub source coverage (code, issues, repos)
- Answer quality and source diversity outweigh latency concerns

## Summary

- **Simple mode** performs a single Google search via Jina without LLM processing, storing no embeddings and executing no reranking (as implemented in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py), lines 34-53).
- **Professional mode** executes a full RAG pipeline including LLM query analysis, parallel searches across four GitHub/Google sources, and neural ranking (as implemented in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py), lines 54-78).
- The `simple_mode` flag defaults to `False` but can be toggled per call or at the instance level.
- Simple mode suits short queries (≤20 words); professional mode handles arbitrary complexity via LLM tokenization.

## Frequently Asked Questions

### Can I use simple mode for complex, multi-sentence queries?

Simple mode is designed for short queries of approximately 20 words or fewer. Longer queries may be truncated because the LLM is not consulted to split or analyze the input. For complex questions, use professional mode to leverage the LLM's query analysis capabilities.

### Do I need to recreate the GithubRAG instance to switch modes?

No. The `simple_mode` parameter can be overridden on a per-call basis in both `retrieve_context` and `async_retrieve_context`. While the instance stores a default value (`self.simple_mode` in lines 20-31 of [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py)), you can pass `simple_mode=True` or `simple_mode=False` to any individual method call.

### Which mode is better for finding specific code implementations?

Professional mode is superior for code-specific questions because it executes dedicated GitHub code searches alongside issue and repository searches. Simple mode only queries Google, which may miss recent commits or specific implementation details found directly in source files.

### Does professional mode cost more in terms of API usage?

Yes. Professional mode consumes LLM tokens for the initial `analyze_question` call and for ranking retrieved snippets. It also makes multiple concurrent API calls to GitHub's search endpoints. Simple mode incurs only the single Google search via Jina and no LLM costs, making it cheaper and faster for basic lookups.