# How to Add Support for a New Language Using the Pluggable AST-Grep Tier

> Easily add new language support to code-graph-rag with the pluggable AST-grep tier. Create a single YAML file, no Python needed, to extend its capabilities.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-20

---

**The pluggable AST-grep tier lets you add a new programming language to code-graph-rag by creating a single YAML configuration file—no Python code required.**

The `code-graph-rag` repository implements a **data-driven AST-grep tier** that automatically discovers language configurations from YAML files. This architecture separates language definitions from core parsing logic, enabling community-driven language support without modifying the Python source. The tier loads pattern configurations from `codebase_rag/parsers/ast_grep_patterns/` and matches file extensions to AST-grep grammars at runtime.

## Understanding the Pluggable AST-Grep Architecture

The pluggable AST-grep tier operates through three core mechanisms defined in [`codebase_rag/parsers/ast_grep_tier.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/ast_grep_tier.py):

**Configuration Loading (`load_pattern_configs`, lines 38-62)**

The `AstGrepTier` class discovers all `*.yaml` files in the patterns directory during initialization. Each YAML file maps one or more file extensions to:

- An `ast_grep_id` matching the grammar identifier in the AST-grep library
- Pattern lists for functions, classes, and imports

**File Processing (`process_file`, lines 99-108)**

When processing a source file, the tier:

1. Checks extension membership via `handles()` to determine applicability
2. Invokes `ast_grep_py.SgRoot` with the configured `ast_grep_id` to parse the AST
3. Executes YAML-defined patterns to extract structural nodes

**Metavariable Handling (`_strip_quotes`, lines 64-68)**

Import paths captured via `$PATH` undergo automatic quote stripping, accommodating languages with varied import syntax conventions.

## Step-by-Step: Adding a New Language

### Step 1: Create the YAML Configuration File

Create a new file in `codebase_rag/parsers/ast_grep_patterns/` named after your language (e.g., [`zig.yaml`](https://github.com/vitali87/code-graph-rag/blob/main/zig.yaml), [`kotlin.yaml`](https://github.com/vitali87/code-graph-rag/blob/main/kotlin.yaml)).

```bash
touch codebase_rag/parsers/ast_grep_patterns/mylang.yaml

```

### Step 2: Define Required Top-Level Keys

Every configuration file requires two mandatory fields:

- **`extensions`**: Comma-separated list of file extensions including the leading dot
- **`ast_grep_id`**: Valid grammar identifier from the [AST-grep language repository](https://ast-grep.github.io/)

### Step 3: Add AST-Grep Patterns

Include three optional pattern lists using AST-grep syntax:

| Pattern Type | Purpose | Required Metavariable |
|-------------|---------|----------------------|
| `functions` | Match function definitions | `$NAME` |
| `classes` | Match class/struct/type definitions | `$NAME` |
| `imports` | Match import/include/require statements | `$PATH` |

Pattern syntax follows AST-grep conventions:

- `$NAME` binds definition identifiers
- `$PATH` binds import module paths
- `$$$ARGS` and `$BODY` match multi-node fragments (ignored by tier, enables flexible matching)

### Step 4: Validate Metavariable Conventions

Critical constraint: Every pattern must bind the expected metavariable. The tier's `_emit_definition` method (lines 27-35) constructs fully-qualified names using `cs.SEPARATOR_DOT` concatenation based on `$NAME` captures.

### Step 5: Test Your Configuration

Run the existing test suite to verify integration:

```bash
pytest -q tests/test_ast_grep_tier.py

```

The test suite automatically loads new YAML files and validates `AstGrepTier.handles()` returns `True` for configured extensions.

### Step 6: Commit and Document

Add the YAML file to version control and update [`docs/roadmap.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/roadmap.md) to reflect expanded language coverage.

## Complete YAML Configuration Example

Below is a production-ready configuration for a hypothetical **MyLang** with `.ml` extension:

```yaml

# codebase_rag/parsers/ast_grep_patterns/mylang.yaml

extensions: ".ml"
ast_grep_id: mylang

functions:
  - "function $NAME($$$ARGS) {$BODY}"
  - "$NAME($$$ARGS) {$BODY}"

classes:
  - "class $NAME {$BODY}"

imports:
  - "import $PATH"
  - "use $PATH"

```

**Field explanations:**

- `extensions: ".ml"` — Routes all `.ml` files to this configuration
- `ast_grep_id: mylang` — Must exactly match the grammar identifier in `ast-grep` releases
- `functions` — Two patterns demonstrate fallback matching: explicit `function` keyword and implicit declaration styles
- `classes` — Captures class definitions via `$NAME`
- `imports` — Dual patterns accommodate `import` and `use` statement variants

## After Installation: Automatic Node Generation

With the optional dependency installed (`pip install code-graph-rag[ast-grep]`), the tier automatically produces three node types for matched files:

1. **MODULE** nodes — One per source file
2. **FUNCTION** and **CLASS** nodes — Extracted via pattern matches with fully-qualified naming
3. **EXTERNAL_MODULE** nodes — Derived from import statements using `_strip_quotes` processing

These nodes conform to `IngestorProtocol` expectations and integrate directly with the knowledge graph pipeline.

## Key Source Files Reference

| File | Responsibility |
|------|---------------|
| [`codebase_rag/parsers/ast_grep_tier.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/ast_grep_tier.py) | Core tier implementation; handles configuration loading, AST parsing, and node emission |
| `codebase_rag/parsers/ast_grep_patterns/` | YAML configuration directory—adding files here extends language support |
| [`codebase_rag/constants/structural.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/structural.py) | `NodeLabel` enum constants (`FUNCTION`, `CLASS`, `EXTERNAL_MODULE`) |
| [`tests/test_ast_grep_tier.py`](https://github.com/vitali87/code-graph-rag/blob/main/tests/test_ast_grep_tier.py) | Automated validation suite for tier behavior |
| [`docs/roadmap.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/roadmap.md) | Documentation of supported languages and extension mechanisms |

## Summary

- **Zero Python required**: Add languages via declarative YAML in `ast_grep_patterns/`
- **Three pattern categories**: Define `functions`, `classes`, and `imports` with AST-grep syntax
- **Mandatory metavariables**: Use `$NAME` for definitions, `$PATH` for imports
- **Automatic discovery**: The tier loads configurations at runtime without code changes
- **Test-driven validation**: Existing tests verify new language integration automatically

## Frequently Asked Questions

### What AST-grep language identifiers are valid?

Valid identifiers correspond to grammars shipped with the `ast-grep` library. Consult the [official AST-grep language support documentation](https://ast-grep.github.io/) for current identifiers. Common examples include `python`, `javascript`, `typescript`, `rust`, `go`, and `ruby`.

### Can I define multiple file extensions for one language?

Yes. Specify multiple extensions as a comma-separated list: `extensions: ".ml,.mli,.mlpack"`. The tier registers all listed extensions against the same configuration.

### What happens if my patterns miss the `$NAME` or `$PATH` metavariable?

The tier will fail to extract node identifiers. In `process_file` (lines 99-108), pattern matches without bound metavariables produce incomplete nodes that may cause downstream ingestor errors. Always validate patterns with `$$$NAME` or `$PATH` bindings.

### Do I need to restart the application after adding a YAML file?

Yes. The `AstGrepTier` loads configurations once during instantiation via `load_pattern_configs`. Restart the ingestion service to pick up new language definitions.