How Llama-GitHub Retrieves Repository Structure (First 3 Levels)
Llama-GitHub fetches the complete repository tree via the GitHub Trees API, converts the flat response into a nested dictionary hierarchy, and recursively truncates the result to the first three directory levels to optimize RAG context windows.
The jetxu-llm/llama-github library streamlines codebase analysis for Retrieval-Augmented Generation (RAG) by providing intelligent repository structure retrieval that balances comprehensive file hierarchy awareness with LLM token efficiency. The system implements a three-stage pipeline that transforms raw GitHub API responses into a condensed, JSON-serializable tree structure ideal for downstream AI processing.
Step 1: Fetching the Full Tree from GitHub
The retrieval process begins in llama_github/github_integration/github_auth_manager.py, where the ExtendedGithub.get_repo_structure method queries the GitHub Trees API endpoint:
/git/trees/{branch}?recursive=1
This returns a flat list where each entry contains a path (the full file or directory path), a type field indicating blob for files or tree for directories, and optional metadata such as size. According to the source code in lines 66–112, the method handles authentication and pagination to ensure the entire repository structure is captured in a single API response.
Step 2: Converting the Flat List to a Hierarchical Tree
Once the flat list is retrieved, the internal helper list_to_tree processes each item to build a nested dictionary structure. As implemented in github_auth_manager.py (lines 66–112), the function:
- Splits every
pathstring on the/delimiter to determine nesting depth - Creates directory nodes containing a
childrendictionary for further traversal - Stores file nodes as leaves containing their full
pathandsize - Strips the
typefield to produce a clean structure requiring no additional type checks during traversal
This transformation converts the GitHub API's linear response into a traversable tree that mirrors the actual filesystem hierarchy.
Step 3: Reducing to the First Three Levels
The truncation logic resides in llama_github/rag_processing/rag_processor.py (lines 60–78). The RAGProcessor.get_repo_simple_structure method first obtains the full tree via repo.get_structure(), then executes a recursive simplify_tree function that:
- Tracks the current depth during recursion
- Returns the full node if the depth is less than 3
- Replaces deeper branches with the placeholder
'...'when the depth reaches 3 - Returns a pretty-printed JSON string containing only top-level directories, their immediate subdirectories, and files directly under those first two tiers
This aggressive pruning ensures that RAG processors receive sufficient architectural context without consuming excessive tokens on deep dependency trees or node_modules directories.
Caching Strategy for Performance
The Repository class in llama_github/data_retrieval/github_entities.py implements a singleton caching pattern through the get_structure method (lines 32–48). On first invocation, the method stores the complete hierarchical tree in a private _structure attribute. Subsequent calls return the cached version immediately, eliminating redundant API requests and tree-building computations when the same repository is analyzed multiple times during a session.
Practical Implementation Examples
Direct API Access
To retrieve the complete repository structure without level limitations:
from llama_github.github_integration.github_auth_manager import ExtendedGithub
gh = ExtendedGithub(login_or_token="YOUR_TOKEN")
full_tree = gh.get_repo_structure("octocat/Hello-World", branch="main")
print(full_tree) # Nested dict with complete layout
Retrieving the Three-Level Summary
For RAG applications requiring the condensed view:
from llama_github.rag_processing.rag_processor import RAGProcessor
from llama_github.data_retrieval.github_api import GitHubAPIHandler
from llama_github.github_integration.github_auth_manager import RepositoryPool
api_handler = GitHubAPIHandler(token="YOUR_TOKEN")
rag = RAGProcessor(github_api_handler=api_handler)
pool = RepositoryPool(github_instance=api_handler.github_instance)
repo = pool.get_repository("octocat/Hello-World")
simple_structure_json = rag.get_repo_simple_structure(repo)
print(simple_structure_json)
Inspecting Both Structures
To compare the full tree against the simplified version:
full_tree = repo.get_structure()
print("Full tree node count:", len(full_tree))
simple_tree = rag.get_repo_simple_structure(repo)
import json, textwrap
print(textwrap.indent(json.dumps(json.loads(simple_tree), indent=2), " "))
Summary
- GitHub Trees API:
ExtendedGithub.get_repo_structurefetches the complete flat file list from/git/trees/{branch}?recursive=1as implemented ingithub_auth_manager.pylines 66–112. - Hierarchy Construction: The
list_to_treehelper converts flat paths into nested dictionaries withchildrennodes for directories and leaf nodes for files. - Depth Limitation:
RAGProcessor.get_repo_simple_structureapplies a recursivesimplify_treefunction that halts at level three, substituting deeper content with'...'(lines 60–78 inrag_processor.py). - Singleton Caching: The
Repositoryclass caches the full structure in_structureafter the firstget_structure()call, ensuring efficient reuse across multiple RAG operations.
Frequently Asked Questions
Why does Llama-GitHub limit repository structure retrieval to three levels?
The three-level restriction optimizes token consumption for Large Language Model prompts while preserving enough hierarchical context for RAG systems to understand codebase architecture. This prevents deep dependency directories like node_modules or build artifacts from overwhelming the context window, ensuring the LLM focuses on high-level project organization rather than granular file listings.
Which GitHub API endpoint powers the repository structure retrieval?
The library utilizes the GitHub Trees API endpoint /git/trees/{branch}?recursive=1 via the ExtendedGithub.get_repo_structure method in github_auth_manager.py. The recursive=1 parameter ensures the API returns every file and directory in the repository in a single request, regardless of nesting depth.
How does the library minimize API calls when analyzing the same repository repeatedly?
The Repository class in github_entities.py implements singleton-style caching through a private _structure attribute. When get_structure() is called for the first time, it fetches and stores the complete tree; subsequent invocations return the cached dictionary immediately without additional network requests or tree-building computations.
Can developers modify the depth limit for repository structure retrieval?
The current implementation hardcodes the three-level limit within the simplify_tree function inside RAGProcessor.get_repo_simple_structure (lines 60–78 of rag_processor.py). Developers requiring deeper traversal must modify the recursion depth check in the source code, as no configuration parameter currently exposes this threshold through the public API.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →