# How to Contribute to code-graph-rag: A Complete Developer Guide

> Learn how to contribute to code-graph-rag. Follow this guide to fork the repo, set up your environment, and submit pull requests for a seamless contribution experience.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-18

---

**You can contribute to code-graph-rag by forking the repository, setting up a Python virtual environment with pre-commit hooks, creating descriptive feature branches, and submitting pull requests that pass the pytest suite and automated Claude-code review workflow.**

The **code-graph-rag** repository is an open-source Python library that constructs graph representations of codebases to enable Retrieval-Augmented Generation (RAG). Whether you want to add support for new programming languages, optimize graph traversal queries, or enhance the CLI interface, following the established contribution guidelines ensures your changes integrate smoothly with the existing architecture.

## Understanding the code-graph-rag Architecture

Before contributing, familiarize yourself with the project structure. The codebase is organized into a **core package** (`codebase_rag/`) containing the graph builder, loaders, and query utilities, alongside support modules for configuration, CLI handling, and benchmarking.

### Core Components

- **[`codebase_rag/schema_builder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/schema_builder.py)** – Generates the GraphQL-style schema describing code entities (functions, classes, modules) and their relationships.
- **[`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)** – Loads persisted graphs (Neo4j, DuckDB, etc.) and provides traversal and Cypher query APIs.
- **[`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py)** – Dynamically discovers and registers language-specific parsers (AST-Grep, tree-sitter) that feed the graph builder.
- **[`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py)** – Wraps vector-store backends (e.g., OpenAI embeddings) and caches embeddings for semantic search.
- **[`codebase_rag/dead_code.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/dead_code.py)** – Detects dead code by analyzing call-graph reachability for the "dead-code" guide.
- **[`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py)** – Implements the command-line interface (`cgr`) orchestrating graph construction, updating, and querying.

## Setting Up Your Development Environment

Prepare your local environment to ensure consistent testing and code quality.

1. **Fork and clone** the repository from `vitali87/code-graph-rag`.

2. **Create a virtual environment** and install dependencies:

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

```

3. **Install pre-commit hooks** to enforce PEP 8 style and linting rules:

```bash
pre-commit install

```

## Step-by-Step Contribution Workflow

Follow this standardized process when you contribute to code-graph-rag:

1. **Create a dedicated branch** with a descriptive name (e.g., `feature/graph-loader-optimisation`).

2. **Run the existing test suite** to verify baseline functionality:

```bash
pytest -q

```

All tests must pass before you introduce changes. Tests reside in the `evals/` directory (e.g., [`evals/semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/semantic_search.py)).

3. **Make localized changes**:
   - Add new parsers to `codebase_rag/parsers/`
   - Fix graph-related issues in [`graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_loader.py) or [`schema_builder.py`](https://github.com/vitali87/code-graph-rag/blob/main/schema_builder.py)
   - Update documentation in `docs/` and ensure [`README.md`](https://github.com/vitali87/code-graph-rag/blob/main/README.md) references new features

4. **Add corresponding tests** under the appropriate `evals/` module. Use existing language-specific eval files ([`java_l1.py`](https://github.com/vitali87/code-graph-rag/blob/main/java_l1.py), [`go_l1.py`](https://github.com/vitali87/code-graph-rag/blob/main/go_l1.py), etc.) as templates for new language support.

5. **Run linters and formatters**:

```bash
pre-commit run --all-files

```

6. **Commit using Conventional Commits** style (e.g., `feat(parser): add Rust AST-Grep support`) and push to your fork.

7. **Open a Pull Request** targeting the upstream `main` branch:
   - Reference related issues (e.g., `Fixes #42`)
   - Include a short description, testing methodology, and any new dependencies per the PR template in [`docs/contributing.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/contributing.md)

8. **Address review feedback**:
   - CI runs tests, style checks, and the optional **Claude-code** review workflow (see [`.github/workflows/claude-code-review.yml`](https://github.com/vitali87/code-graph-rag/blob/main/.github/workflows/claude-code-review.yml))
   - Rebase if necessary to maintain a clean history

9. **Merge** once CI passes and maintainers approve your changes.

## Practical Contribution Examples

### Adding a New Language Parser

Create a parser class in [`codebase_rag/parsers/rust_parser.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/rust_parser.py):

```python
from .base import BaseParser

class RustParser(BaseParser):
    language = "rust"

    def parse(self, source_path: str):
        # implement AST-Grep or tree-sitter parsing for Rust files

        ...

```

Register the parser in [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py):

```python
from .rust_parser import RustParser
...
ParserRegistry.register(RustParser())

```

### Extending the Graph Schema

Modify [`codebase_rag/schema_builder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/schema_builder.py) to add new relationship types:

```python
def add_dead_code_edge(self, src_node, dst_node):
    """Create a DEAD_CODE edge if dst_node is unreachable from the entry point."""
    self.graph.add_edge(src_node, dst_node, type="DEAD_CODE")

```

### Testing Semantic Search

Verify RAG functionality using the CLI:

```bash
cgr query \
  --graph-path ./graph.db \
  --question "How does the caching layer work?" \
  --top-k 5

```

## Summary

- **Fork and branch** from `main` before making changes
- **Install pre-commit hooks** to automatically enforce PEP 8 style requirements
- **Localize changes**: place new parsers in `codebase_rag/parsers/` and graph fixes in [`schema_builder.py`](https://github.com/vitali87/code-graph-rag/blob/main/schema_builder.py) or [`graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_loader.py)
- **Test thoroughly** using `pytest -q` and add tests to the `evals/` directory
- **Follow Conventional Commits** and the PR template in [`docs/contributing.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/contributing.md)
- **Pass automated review** including the Claude-code workflow before merge

## Frequently Asked Questions

### What coding standards must I follow to contribute to code-graph-rag?

The project follows **PEP 8** style guidelines enforced through **pre-commit** hooks. You must run `pre-commit install` after setting up your virtual environment to ensure automatic linting and formatting before each commit.

### Where should I add tests for new features?

Add unit tests under the `evals/` directory. For language-specific contributions, reference existing evaluation files like [`evals/java_l1.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/java_l1.py) or [`evals/go_l1.py`](https://github.com/vitali87/code-graph-rag/blob/main/evals/go_l1.py) as templates to ensure consistent testing patterns.

### How do I register a new parser in the codebase?

Import your parser class in [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py) and call `ParserRegistry.register(YourParser())`. The dynamic discovery system in [`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py) handles the rest of the integration with the graph builder.

### What happens during the pull request review process?

CI automatically runs the full test suite, style checks, and an optional **Claude-code** review workflow defined in [`.github/workflows/claude-code-review.yml`](https://github.com/vitali87/code-graph-rag/blob/main/.github/workflows/claude-code-review.yml). Maintainers review for architectural alignment and test coverage before approving merges into the `main` branch.