# How CodeWiki Fetches Repository Structures from GitHub: A Complete Technical Guide

> Learn how CodeWiki fetches repository structures from GitHub. It parses URLs, queries the GitHub API for file trees and readmes, and decodes content for a clear view.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: how-to-guide
- Published: 2026-02-16

---

**CodeWiki retrieves GitHub repository structures by parsing the repository URL to extract owner and name, then querying the GitHub REST API's git/trees endpoint for file trees and the /readme endpoint for documentation, filtering blob objects and decoding base64 content.**

CodeWiki is an open-source tool that automatically generates wiki documentation from code repositories. To accomplish this, it must first fetch repository structures from GitHub, retrieving both the complete file tree and the README content. This article examines the exact implementation details found in the CodeWiki source code, demonstrating how the `RepositoryStructureFetcher` class and supporting services interact with the GitHub API.

## The Two-Step Architecture for Fetching Repository Structures

CodeWiki implements a clean separation between URL parsing and API orchestration. The process begins in the frontend and flows through specialized utility modules before hitting GitHub's servers.

### Step 1: Parsing the Repository URL

When a user submits a repository URL (for example, `https://github.com/user/repo`), the system first validates and decomposes this input. The function `parse_repository_input()` located in [`utils/repository_parser.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_parser.py) extracts the **owner** and **repo** identifiers and marks the source type as *web*.

This normalization ensures that downstream components receive consistent data structures regardless of whether the input was a full URL, a shorthand string, or a local path.

### Step 2: Orchestrating the GitHub API Calls

With parsed credentials in hand, the `RepositoryStructureFetcher` class in [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py) takes control. Its primary method, `fetch_repository_structure()`, distinguishes between local and web repositories. For GitHub-hosted code, it delegates to specialized methods that construct authenticated HTTP requests against the GitHub REST API.

## How CodeWiki Uses the GitHub REST API to Fetch File Trees

The core mechanism for retrieving repository structures relies on GitHub's Git Data API, specifically the **git/trees** endpoint. This endpoint returns recursive file listings without downloading actual file contents, making it efficient for large repositories.

### Constructing the Tree API Endpoint

Inside `fetch_repository_structure()`, CodeWiki attempts to fetch the tree for both `main` and `master` branches to maximize compatibility. The URL construction follows this pattern:

```python
api_url = f"https://api.github.com/repos/{self.owner}/{self.repo}/git/trees/{branch}?recursive=1"

```

The `recursive=1` query parameter ensures that the API returns all nested directories and files in a single request, eliminating the need for multiple round trips to traverse deep folder hierarchies.

### Handling Authentication with GitHub Headers

To support both public and private repositories, CodeWiki implements a header generation strategy in `create_github_headers()` (lines 43-47 of [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)):

```python
def create_github_headers(self):
    headers = {
        "Accept": "application/vnd.github.v3+json",
        "Authorization": f"token {self.token}" if self.token else None
    }
    return {k: v for k, v in headers.items() if v is not None}

```

The `Accept` header specifies the GitHub API v3 media type, while the optional `Authorization` header carries a personal access token (PAT) when `self.token` is provided. This conditional inclusion ensures that public repositories work without authentication while private repositories remain accessible to authorized users.

### Processing the API Response and Filtering Files

Upon receiving the JSON response from the tree endpoint, CodeWiki extracts the `tree` list and filters for objects of type `"blob"` (representing files), discarding directories and submodules. This logic appears in lines 146-154 of [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py):

```python
response = requests.get(api_url, headers=headers)
data = response.json()

# Extract only file paths, ignoring directories

file_paths = [item["path"] for item in data.get("tree", []) if item["type"] == "blob"]
file_tree_data = "\n".join(file_paths)

```

The resulting `file_tree_data` string contains one file path per line, providing a clean, linear representation of the repository structure suitable for LLM processing.

## Fetching Repository README Content from GitHub

Beyond the file tree, CodeWiki retrieves the repository's README to provide context for documentation generation. The README fetch occurs immediately after the tree retrieval in `fetch_repository_structure()` (lines 155-164):

```python
readme_url = f"https://api.github.com/repos/{self.owner}/{self.repo}/readme"
readme_response = requests.get(readme_url, headers=headers)

if readme_response.status_code == 200:
    readme_data = readme_response.json()
    readme_content = base64.b64decode(readme_data["content"]).decode("utf-8")
else:
    readme_content = ""

```

The `/readme` endpoint returns the README file's metadata along with its content encoded in base64. CodeWiki decodes this content to obtain the raw Markdown text, which is then stored in `readme_content` for subsequent processing by the documentation generation pipeline.

## Reusable Components: The GithubService Class

To promote code reuse across different features (such as diagram generation), CodeWiki encapsulates the GitHub API logic in a dedicated `GithubService` class located in [`api/services/github_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/github_service.py). This service exposes the same tree and README retrieval methods used by the fetcher, but in a more modular form:

```python
from api.services.github_service import GithubService

# Initialize with optional authentication

service = GithubService(
    owner="quangdungluong",
    repo="codewiki",
    token=None  # Add your PAT here for private repos

)

# Retrieve file tree and default branch

file_tree, default_branch = service.get_tree_data()
print(f"Default branch: {default_branch}")
print(f"Total files: {len(file_tree.splitlines())}")

# Retrieve README content

readme_md = service.get_readme()
print(f"README length: {len(readme_md)} characters")

```

The `GithubService` implementation mirrors the logic found in `RepositoryStructureFetcher`, including the recursive tree endpoint usage, blob filtering, and base64 decoding for README content. This separation allows diagram generation endpoints in [`api/generate_diagram.py`](https://github.com/quangdungluong/codewiki/blob/main/api/generate_diagram.py) to leverage GitHub data without importing the heavier wiki-generation machinery.

## Complete Implementation Example

For developers looking to integrate similar functionality, here is the complete flow using CodeWiki's `RepositoryStructureFetcher`:

```python
import asyncio
from utils.repository_parser import parse_repository_input
from utils.repository_structure import RepositoryStructureFetcher

async def fetch_repo_structure(repo_url: str, token: str = None):
    """
    Fetch repository structure using CodeWiki's internal fetcher.
    """
    # Step 1: Parse the repository URL

    repo_info = parse_repository_input(repo_url)
    
    if repo_info["type"] != "web":
        raise ValueError("Only GitHub URLs are supported in this example")
    
    # Step 2: Initialize the fetcher

    fetcher = RepositoryStructureFetcher(
        repo_info=repo_info,
        repo_url=repo_url,
        owner=repo_info["owner"],
        repo=repo_info["repo"],
        token=token
    )
    
    # Step 3: Define status callback (optional)

    async def update_status(task_id, status, message, data=None):
        print(f"[{status}] {message}")
    
    # Step 4: Execute the fetch

    await fetcher.fetch_repository_structure(
        update_status, 
        task_id="example-task"
    )
    
    return {
        "file_tree": fetcher.file_tree_data,
        "readme": fetcher.readme_content,
        "wiki_structure": fetcher.wiki_structure
    }

# Usage example

if __name__ == "__main__":
    result = asyncio.run(fetch_repo_structure(
        "https://github.com/quangdungluong/codewiki"
    ))
    print(f"Fetched {len(result['file_tree'].splitlines())} files")

```

This example demonstrates the complete pipeline: URL parsing, fetcher initialization, asynchronous execution, and result extraction. The `file_tree_data` property contains the newline-separated file paths, while `readme_content` holds the decoded Markdown.

## Summary

CodeWiki fetches repository structures from GitHub through a well-architected pipeline that separates concerns between URL parsing, API communication, and data processing:

- **URL Parsing**: The `parse_repository_input()` function in [`utils/repository_parser.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_parser.py) extracts owner and repository names from GitHub URLs.
- **API Orchestration**: The `RepositoryStructureFetcher` class in [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py) manages the asynchronous workflow, handling both the recursive tree endpoint and the README endpoint.
- **Authentication**: The `create_github_headers()` method supports optional Personal Access Tokens via the `Authorization` header while maintaining compatibility with public repositories.
- **Data Processing**: The system filters GitHub's tree API response for `"blob"` types only, producing a clean newline-separated file list, and base64-decodes README content for immediate use.
- **Reusability**: The `GithubService` class in [`api/services/github_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/github_service.py) encapsulates the same logic for use by other features like diagram generation.

## Frequently Asked Questions

### What GitHub API endpoint does CodeWiki use to fetch repository structures?

CodeWiki uses the **GitHub Git Data API's tree endpoint** at `https://api.github.com/repos/{owner}/{repo}/git/trees/{branch}?recursive=1`. The `recursive=1` parameter ensures all nested directories are returned in a single request. This logic is implemented in the `fetch_repository_structure()` method of [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py) and the `get_tree_data()` method of [`api/services/github_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/github_service.py).

### How does CodeWiki handle private repositories when fetching structures?

CodeWiki supports private repositories through **Personal Access Token (PAT) authentication**. The `create_github_headers()` method in [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py) conditionally adds an `Authorization` header with the format `token {self.token}` when a token is provided. If no token is supplied, the request proceeds without authentication, which works for public repositories but returns 404 errors for private ones.

### What happens if a repository doesn't have a main or master branch?

CodeWiki implements a **fallback branch strategy** when fetching repository structures. The system first attempts to retrieve the tree using the `main` branch, and if that request fails, it automatically retries with the `master` branch. This logic appears in the error handling section of `fetch_repository_structure()` (lines 136-144 of [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)). If both branches fail, the system raises a detailed exception indicating that the repository structure could not be retrieved.

### Can I use CodeWiki's GitHub fetching logic independently of the wiki generation?

Yes, CodeWiki exposes a reusable **GithubService** class specifically for this purpose. Located in [`api/services/github_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/github_service.py), this class provides the `get_tree_data()` and `get_readme()` methods without requiring the full `RepositoryStructureFetcher` initialization. The diagram generation feature in [`api/generate_diagram.py`](https://github.com/quangdungluong/codewiki/blob/main/api/generate_diagram.py) demonstrates this independent usage, importing `GithubService` directly to cache repository data for architectural visualization without triggering the LLM-based wiki creation pipeline.