# How to Add Support for a New Programming Language in Code-Graph-RAG: A Complete Guide

> Learn how to add support for a new programming language in Code-Graph-RAG. This guide covers FQNSpec, LanguageSpec, enum extension, and registry registration for seamless integration.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-19

---

**To add support for a new programming language in Code-Graph-RAG, you must define an `FQNSpec` for fully-qualified name extraction, create a `LanguageSpec` with tree-sitter node types, extend the `SupportedLanguage` enum, and register both specifications in the central registry dictionaries.**

Code-Graph-RAG (CG-RAG) is an open-source codebase analysis tool that builds language-aware abstract syntax graphs using tree-sitter parsers. Adding a new language requires implementing two core data structures—**`FQNSpec`** and **`LanguageSpec`**—within the `codebase_rag` package and registering them so the ingestion engine can automatically discover and parse files with the new extension.

## Architecture Overview

CG-RAG relies on tree-sitter to parse source files into AST nodes. The system extracts **fully-qualified names (FQNs)** for functions, classes, and modules by mapping tree-sitter node types to semantic concepts. Every supported language must provide:

- **Name extraction helpers**: Functions that extract identifiers from declaration nodes and convert file paths to module hierarchies.
- **Node type mappings**: Sets of tree-sitter node types that correspond to functions, classes, imports, and calls.
- **Registration entries**: Insertions into the global `LANGUAGE_FQN_SPECS` and `LANGUAGE_SPECS` dictionaries defined in [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py).

The ingestion pipeline (specifically in [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py), lines 291-304) automatically looks up the appropriate `LanguageSpec` for each file based on its extension, making the registration step critical for automatic discovery.

## Step 1: Define AST Constants (Optional)

Create a new file `codebase_rag/constants/ast_<lang>.py` to store tree-sitter node-type strings as constants. This follows the pattern established by [`ast_python.py`](https://github.com/vitali87/code-graph-rag/blob/main/ast_python.py) and keeps your specifications maintainable.

```python

# codebase_rag/constants/ast_kotlin.py

TS_KOTLIN_CLASS = "class_declaration"
TS_KOTLIN_FUNCTION = "function_declaration"
TS_KOTLIN_MODULE = "package_header"
TS_KOTLIN_CALL = "call_expression"
TS_KOTLIN_IMPORT = "import_header"

```

While optional, centralizing these constants prevents string duplication across your `FQNSpec` and `LanguageSpec` definitions.

## Step 2: Implement Name Extraction Helpers

Open [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py) and add two helper functions specific to your language. These mirror the existing `_python_get_name` and `_python_file_to_module` implementations (lines 13-19 and 22-29).

```python
def _kotlin_get_name(node: Node) -> str | None:
    """Extract the identifier from a declaration node."""
    name_node = node.child_by_field_name("name")
    return (
        name_node.text.decode("utf-8") 
        if name_node and name_node.text 
        else None
    )

def _kotlin_file_to_module(file_path: Path, repo_root: Path) -> list[str]:
    """Convert a file path to a dotted module name list."""
    rel = file_path.relative_to(repo_root)
    parts = list(rel.with_suffix("").parts)
    return parts

```

The `get_name` function must return the string identifier of declaration nodes (classes, functions), while `file_to_module_parts` returns a list of path components representing the module hierarchy.

## Step 3: Create an FQNSpec Configuration

Instantiate an **`FQNSpec`** in [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py) with your helper functions and relevant node type sets. See the Python reference at lines 23-28.

```python
KOTLIN_FQN_SPEC = FQNSpec(
    scope_node_types=frozenset({TS_KOTLIN_CLASS}),
    function_node_types=frozenset({TS_KOTLIN_FUNCTION}),
    get_name=_kotlin_get_name,
    file_to_module_parts=_kotlin_file_to_module,
)

```

The `scope_node_types` define which AST nodes create new naming scopes, while `function_node_types` identify callable definitions. The `get_name` and `file_to_module_parts` parameters wire your helpers into the FQN extraction engine.

## Step 4: Build a LanguageSpec Configuration

Create a **`LanguageSpec`** that aggregates file extensions, node type sets, and optional tree-sitter queries. Reference the Rust implementation at lines 42-51 of [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py).

```python
from codebase_rag.constants import SupportedLanguage

LANGUAGE_SPECS: dict[SupportedLanguage, LanguageSpec] = {
    # ... existing languages ...

    SupportedLanguage.KOTLIN: LanguageSpec(
        language=SupportedLanguage.KOTLIN,
        file_extensions={".kt", ".kts"},
        function_node_types={TS_KOTLIN_FUNCTION},
        class_node_types={TS_KOTLIN_CLASS},
        module_node_types={TS_KOTLIN_MODULE},
        call_node_types={TS_KOTLIN_CALL},
        import_node_types={TS_KOTLIN_IMPORT},
        import_from_node_types={TS_KOTLIN_IMPORT},  # Kotlin uses same type

        package_indicators={"build.gradle.kts", "settings.gradle.kts"},
        function_query="""(function_declaration name: (identifier) @function)""",
        class_query="""(class_declaration name: (identifier) @class)""",
        call_query="""(call_expression function: (identifier) @call)""",
    ),
}

```

The `LanguageSpec` tells the graph builder which node types correspond to semantic constructs. The optional query strings enable more precise AST extraction when the default node-type matching is insufficient.

## Step 5: Extend the SupportedLanguage Enum

Add a new member to the **`SupportedLanguage`** enum in [`codebase_rag/constants/__init__.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/__init__.py). This enables type-safe language handling throughout the codebase.

```python
class SupportedLanguage(Enum):
    PYTHON = auto()
    JS = auto()
    RUST = auto()
    DART = auto()
    KOTLIN = auto()  # New entry

```

## Step 6: Register the Specifications

Insert your `FQNSpec` into **`LANGUAGE_FQN_SPECS`** and your `LanguageSpec` into **`LANGUAGE_SPECS`** (lines 14-30 and 42-71 in [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py)). The `_EXTENSION_TO_SPEC` dictionary (lines 61-65) will automatically populate from your `LANGUAGE_SPECS` entry.

```python
LANGUAGE_FQN_SPECS: dict[SupportedLanguage, FQNSpec] = {
    SupportedLanguage.PYTHON: PYTHON_FQN_SPEC,
    SupportedLanguage.RUST: RUST_FQN_SPEC,
    SupportedLanguage.KOTLIN: KOTLIN_FQN_SPEC,  # Register FQN spec

}

# LANGUAGE_SPECS dictionary insertion shown in Step 4 above

```

Once registered, the `get_language_spec()` function will automatically return your configuration when processing `.kt` or `.kts` files.

## Step 7: Add Parser Support (If Needed)

If tree-sitter does not already include a grammar for your language in the repository, create a new module under `codebase_rag/parsers/<lang>/` with utility functions. Follow the structure of [`parsers/js_ts/utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/parsers/js_ts/utils.py).

Update [`requirements.txt`](https://github.com/vitali87/code-graph-rag/blob/main/requirements.txt) or [`pyproject.toml`](https://github.com/vitali87/code-graph-rag/blob/main/pyproject.toml) to include the `tree_sitter_<lang>` package from PyPI if available:

```toml
dependencies = [
    "tree-sitter>=0.20",
    "tree-sitter-python>=0.20",
    "tree-sitter-kotlin>=0.3",  # New dependency

]

```

## Step 8: Test Your Implementation

Run the test suite to ensure the new language does not break existing functionality:

```bash
pytest -q

```

Add specific tests in `tests/` that verify your parser produces correct fully-qualified names. A minimal test should parse a file containing a class and function, then assert the generated graph contains the expected FQN (e.g., `com.example.MyClass.myMethod`).

## Practical Example: Complete Kotlin Integration

Here is the complete implementation pattern for Kotlin support, consolidating all previous steps:

```python

# codebase_rag/constants/ast_kotlin.py

TS_KOTLIN_CLASS = "class_declaration"
TS_KOTLIN_FUNCTION = "function_declaration"
TS_KOTLIN_MODULE = "package_header"
TS_KOTLIN_CALL = "call_expression"
TS_KOTLIN_IMPORT = "import_header"

# codebase_rag/language_spec.py additions

def _kotlin_get_name(node: Node) -> str | None:
    name_node = node.child_by_field_name("name")
    return name_node.text.decode("utf-8") if name_node and name_node.text else None

def _kotlin_file_to_module(file_path: Path, repo_root: Path) -> list[str]:
    return list(file_path.relative_to(repo_root).with_suffix("").parts)

# Register FQNSpec

KOTLIN_FQN_SPEC = FQNSpec(
    scope_node_types=frozenset({TS_KOTLIN_CLASS}),
    function_node_types=frozenset({TS_KOTLIN_FUNCTION}),
    get_name=_kotlin_get_name,
    file_to_module_parts=_kotlin_file_to_module,
)

LANGUAGE_FQN_SPECS[SupportedLanguage.KOTLIN] = KOTLIN_FQN_SPEC

# Register LanguageSpec (as shown in Step 4)

```

After deployment, CG-RAG will automatically parse Kotlin files and generate graph nodes with fully-qualified names like [`src/main/com/example/MyClass.kt`](https://github.com/vitali87/code-graph-rag/blob/main/src/main/com/example/MyClass.kt) as `com.example.MyClass`.

## Summary

- **Define helpers**: Implement `_<lang>_get_name` and `_<lang>_file_to_module` in [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py) to extract identifiers and map file paths.
- **Create specifications**: Build an `FQNSpec` for name resolution and a `LanguageSpec` with node types and file extensions.
- **Extend the enum**: Add your language to `SupportedLanguage` in [`codebase_rag/constants/__init__.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/__init__.py).
- **Register specs**: Insert entries into `LANGUAGE_FQN_SPECS` and `LANGUAGE_SPECS` so the [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) ingestion pipeline discovers the language automatically.
- **Add parser**: Include the tree-sitter grammar under `parsers/` if not already present, and update dependencies.

Following these steps integrates your language into the CG-RAG ingestion pipeline, enabling automatic graph construction for any repository using that language.

## Frequently Asked Questions

### What is the minimum code required to add a new language to Code-Graph-RAG?

You must define two helper functions (`get_name` and `file_to_module`), create an `FQNSpec` and `LanguageSpec`, extend the `SupportedLanguage` enum, and register both specs in `LANGUAGE_FQN_SPECS` and `LANGUAGE_SPECS`. The system automatically handles the rest through the extension-to-spec mapping in [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py).

### Where does Code-Graph-RAG look up language specifications during ingestion?

The [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) module calls `get_language_spec(path.suffix)` (lines 291-304) for each file change. This function queries the `_EXTENSION_TO_SPEC` dictionary, which is automatically populated from your `LANGUAGE_SPECS` registration, returning the appropriate `LanguageSpec` for parsing.

### Do I need to modify the core ingestion logic to support a new language?

No. As long as you register your specifications in the dictionaries at the end of [`codebase_rag/language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py), the existing ingestion pipeline in [`realtime_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/realtime_updater.py) will automatically recognize and parse files with your new language's extensions without requiring changes to the core engine.

### How do I handle languages with different module resolution schemes (e.g., flat vs. hierarchical)?

Customize the `_<lang>_file_to_module` helper function in [`language_spec.py`](https://github.com/vitali87/code-graph-rag/blob/main/language_spec.py). For flat resolution, return a single-element list containing the filename without extension. For hierarchical resolution, return the relative path parts as shown in the examples for Python and Kotlin. The FQN builder uses this list to construct the dotted namespace.