# How Many Tree-Sitter Grammars Does Codebase-Memory-MCP Support? Complete Technical Breakdown

> Discover how many Tree-sitter grammars Codebase-Memory-MCP supports. Explore the breakdown of 91 supported grammars including custom additions like Magma and Form. Learn more now.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: deep-dive
- Published: 2026-07-09

---

**codebase-memory-mcp supports 91 tree-sitter grammars** out of the box, comprising 89 languages defined in a central JSON registry and 2 custom grammars (Magma and Form) bundled directly in the repository.

codebase-memory-mcp is an open-source Model Context Protocol (MCP) server that leverages tree-sitter for high-performance source code analysis across multiple languages. Understanding exactly how many tree-sitter grammars the tool supports—and how they are organized within the codebase—enables developers to leverage its multi-language parsing capabilities effectively. The project uses a hybrid architecture that combines a dynamic language registry with statically linked custom parsers.

## The Complete Grammar Count: 91 Supported Languages

The **91 tree-sitter grammars** are split between a public language registry and two specialized internal grammars.

### Registry-Based Grammars (89 Languages)

The vast majority of supported languages are catalogued in [`scripts/new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/new-languages.json). This file contains **89 distinct entries**, each defining a language name, a tree-sitter function name (`ts_func`), and a GitHub repository pointer for the grammar source. Entries specify function signatures like `tree_sitter_rust` or `tree_sitter_go`, which correspond to the C functions generated by the tree-sitter compiler.

### Custom Bundled Grammars (2 Languages)

In addition to the registry, the repository bundles **2 custom grammars** not listed in the JSON file:

- **Magma**: Defined in [`tools/tree-sitter-magma/tree-sitter.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/tools/tree-sitter-magma/tree-sitter.json)
- **Form**: Defined in [`tools/tree-sitter-form/tree-sitter.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/tools/tree-sitter-form/tree-sitter.json)

These grammars are internal extensions used specifically by the MCP itself. Despite not appearing in [`new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/new-languages.json), they follow the same [`tree-sitter.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/tree-sitter.json) schema as public entries, allowing the system to discover and invoke them using the same unified API.

## Architecture: How Tree-Sitter Grammars Are Loaded

Understanding the loading mechanism clarifies how the 91 grammars are accessed at runtime.

### Language Registry and ts_func Mapping

At startup, the executable reads [`scripts/new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/new-languages.json) to build an in-memory language map. Each JSON entry supplies a `ts_func` string, such as `tree_sitter_python` or `tree_sitter_javascript`. This string represents the exact C-function symbol generated by the tree-sitter grammar’s parser code, following the standard signature `TSLanguage *tree_sitter_X(void)`.

### Dynamic Loading by File Extension

When processing a file, the system determines the language by inspecting file extensions or specific filenames (e.g., `.go` maps to `GOTEMPLATE`, `go.mod` maps to `GOMOD`). The function `resolve_language_from_path` returns a language enum, which `lookup_ts_func` then maps to the corresponding `ts_func` function pointer. This dynamic resolution works for all 89 registry grammars plus the 2 custom ones.

### Custom Grammar Integration

The Magma and Form grammars reside in the `tools/` directory and are compiled directly into the binary. They expose the same C API contract (`TSLanguage *tree_sitter_magma(void)` and `TSLanguage *tree_sitter_form(void)`), allowing the MCP to treat them identically to registry-based grammars. Adding a new custom grammar requires placing its descriptor in `tools/` and rebuilding; adding a standard grammar only requires appending to [`new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/new-languages.json).

## Practical Implementation: Parsing Code Examples

The following C-style pseudo-code demonstrates how codebase-memory-mcp selects and executes a tree-sitter parser for any of the 91 supported grammars:

```c
/* Resolve the language enum from a filename */
LanguageEnum lang = resolve_language_from_path("example.rs");

/* Look up the tree-sitter function for that language */
TSLanguage *(*ts_func)(void) = lookup_ts_func(lang);

/* Initialise the parser */
TSParser *parser = ts_parser_new();
ts_parser_set_language(parser, ts_func());

/* Parse the source */
TSTree *tree = ts_parser_parse_string(
    parser,
    NULL,
    source_code,          /* const char *source_code */
    source_len            /* uint32_t source_len */
);

/* Walk the syntax tree … */

```

This workflow relies on `resolve_language_from_path` to match file extensions against the registry and `lookup_ts_func` to return the correct `tree_sitter_X` function pointer—operations that function identically whether the grammar is one of the 89 JSON-defined languages or the 2 custom Magma/Form grammars.

## Key Configuration Files

The following files collectively enumerate the **91** tree-sitter grammars:

- **[`scripts/new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/new-languages.json)** – Contains the 89 public tree-sitter grammar definitions with their `ts_func` mappings and repository sources.
- **[`tools/tree-sitter-magma/tree-sitter.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/tools/tree-sitter-magma/tree-sitter.json)** – Descriptor for the custom Magma grammar (adds 1 to the total count).
- **[`tools/tree-sitter-form/tree-sitter.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/tools/tree-sitter-form/tree-sitter.json)** – Descriptor for the custom Form grammar (adds the final 1 to reach 91).

## Summary

- **codebase-memory-mcp supports 91 tree-sitter grammars**: 89 defined in [`scripts/new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/new-languages.json) and 2 custom grammars (Magma and Form) bundled in the `tools/` directory.
- The system uses a **unified C API** (`TSLanguage *tree_sitter_X(void)`) for all grammars, whether registry-based or custom.
- **Dynamic loading** relies on `resolve_language_from_path` and `lookup_ts_func` to map file extensions to the correct parser functions.
- Adding support for new languages requires either appending to [`new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/new-languages.json) (for standard grammars) or adding a custom descriptor in `tools/` (for specialized grammars).

## Frequently Asked Questions

### How do I add a new tree-sitter grammar to codebase-memory-mcp?

Append a JSON object to [`scripts/new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/new-languages.json) specifying the language name, file extensions, and the `ts_func` name (e.g., `tree_sitter_newlang`). Ensure the grammar library is compiled into the binary so the function symbol is available at runtime. The system will automatically recognize the new language on the next startup without modifying core logic.

### What are the Magma and Form grammars used for?

The **Magma** and **Form** grammars are internal extensions used by the MCP itself for parsing specialized syntax not covered by standard tree-sitter repositories. Located in `tools/tree-sitter-magma/` and `tools/tree-sitter-form/`, they follow the same descriptor schema as public grammars but remain separate from the [`new-languages.json`](https://github.com/DeusData/codebase-memory-mcp/blob/main/new-languages.json) registry to avoid conflicts with upstream language definitions.

### How does the system determine which grammar to use for a specific file?

The system maps file extensions and special filenames (like `go.mod` or [`Cargo.toml`](https://github.com/DeusData/codebase-memory-mcp/blob/main/Cargo.toml)) to language enums using `resolve_language_from_path`. It then calls `lookup_ts_func` to retrieve the appropriate `tree_sitter_X` function pointer from the registry or custom grammar set. This resolution supports all 91 grammars identically, whether standard or custom.

### Are all 91 grammars available in a standard binary build?

Yes. All **91 tree-sitter grammars** are compiled into the standard binary. The 89 registry grammars are linked as dependencies specified in the build configuration, while the 2 custom grammars (Magma and Form) are built from source located in the `tools/` directory. No external downloads are required at runtime.