How Many Tree-Sitter Grammars Does Codebase-Memory-MCP Support? Complete Technical Breakdown

codebase-memory-mcp supports 91 tree-sitter grammars out of the box, comprising 89 languages defined in a central JSON registry and 2 custom grammars (Magma and Form) bundled directly in the repository.

codebase-memory-mcp is an open-source Model Context Protocol (MCP) server that leverages tree-sitter for high-performance source code analysis across multiple languages. Understanding exactly how many tree-sitter grammars the tool supports—and how they are organized within the codebase—enables developers to leverage its multi-language parsing capabilities effectively. The project uses a hybrid architecture that combines a dynamic language registry with statically linked custom parsers.

The Complete Grammar Count: 91 Supported Languages

The 91 tree-sitter grammars are split between a public language registry and two specialized internal grammars.

Registry-Based Grammars (89 Languages)

The vast majority of supported languages are catalogued in scripts/new-languages.json. This file contains 89 distinct entries, each defining a language name, a tree-sitter function name (ts_func), and a GitHub repository pointer for the grammar source. Entries specify function signatures like tree_sitter_rust or tree_sitter_go, which correspond to the C functions generated by the tree-sitter compiler.

Custom Bundled Grammars (2 Languages)

In addition to the registry, the repository bundles 2 custom grammars not listed in the JSON file:

These grammars are internal extensions used specifically by the MCP itself. Despite not appearing in new-languages.json, they follow the same tree-sitter.json schema as public entries, allowing the system to discover and invoke them using the same unified API.

Architecture: How Tree-Sitter Grammars Are Loaded

Understanding the loading mechanism clarifies how the 91 grammars are accessed at runtime.

Language Registry and ts_func Mapping

At startup, the executable reads scripts/new-languages.json to build an in-memory language map. Each JSON entry supplies a ts_func string, such as tree_sitter_python or tree_sitter_javascript. This string represents the exact C-function symbol generated by the tree-sitter grammar’s parser code, following the standard signature TSLanguage *tree_sitter_X(void).

Dynamic Loading by File Extension

When processing a file, the system determines the language by inspecting file extensions or specific filenames (e.g., .go maps to GOTEMPLATE, go.mod maps to GOMOD). The function resolve_language_from_path returns a language enum, which lookup_ts_func then maps to the corresponding ts_func function pointer. This dynamic resolution works for all 89 registry grammars plus the 2 custom ones.

Custom Grammar Integration

The Magma and Form grammars reside in the tools/ directory and are compiled directly into the binary. They expose the same C API contract (TSLanguage *tree_sitter_magma(void) and TSLanguage *tree_sitter_form(void)), allowing the MCP to treat them identically to registry-based grammars. Adding a new custom grammar requires placing its descriptor in tools/ and rebuilding; adding a standard grammar only requires appending to new-languages.json.

Practical Implementation: Parsing Code Examples

The following C-style pseudo-code demonstrates how codebase-memory-mcp selects and executes a tree-sitter parser for any of the 91 supported grammars:

/* Resolve the language enum from a filename */
LanguageEnum lang = resolve_language_from_path("example.rs");

/* Look up the tree-sitter function for that language */
TSLanguage *(*ts_func)(void) = lookup_ts_func(lang);

/* Initialise the parser */
TSParser *parser = ts_parser_new();
ts_parser_set_language(parser, ts_func());

/* Parse the source */
TSTree *tree = ts_parser_parse_string(
    parser,
    NULL,
    source_code,          /* const char *source_code */
    source_len            /* uint32_t source_len */
);

/* Walk the syntax tree … */

This workflow relies on resolve_language_from_path to match file extensions against the registry and lookup_ts_func to return the correct tree_sitter_X function pointer—operations that function identically whether the grammar is one of the 89 JSON-defined languages or the 2 custom Magma/Form grammars.

Key Configuration Files

The following files collectively enumerate the 91 tree-sitter grammars:

Summary

  • codebase-memory-mcp supports 91 tree-sitter grammars: 89 defined in scripts/new-languages.json and 2 custom grammars (Magma and Form) bundled in the tools/ directory.
  • The system uses a unified C API (TSLanguage *tree_sitter_X(void)) for all grammars, whether registry-based or custom.
  • Dynamic loading relies on resolve_language_from_path and lookup_ts_func to map file extensions to the correct parser functions.
  • Adding support for new languages requires either appending to new-languages.json (for standard grammars) or adding a custom descriptor in tools/ (for specialized grammars).

Frequently Asked Questions

How do I add a new tree-sitter grammar to codebase-memory-mcp?

Append a JSON object to scripts/new-languages.json specifying the language name, file extensions, and the ts_func name (e.g., tree_sitter_newlang). Ensure the grammar library is compiled into the binary so the function symbol is available at runtime. The system will automatically recognize the new language on the next startup without modifying core logic.

What are the Magma and Form grammars used for?

The Magma and Form grammars are internal extensions used by the MCP itself for parsing specialized syntax not covered by standard tree-sitter repositories. Located in tools/tree-sitter-magma/ and tools/tree-sitter-form/, they follow the same descriptor schema as public grammars but remain separate from the new-languages.json registry to avoid conflicts with upstream language definitions.

How does the system determine which grammar to use for a specific file?

The system maps file extensions and special filenames (like go.mod or Cargo.toml) to language enums using resolve_language_from_path. It then calls lookup_ts_func to retrieve the appropriate tree_sitter_X function pointer from the registry or custom grammar set. This resolution supports all 91 grammars identically, whether standard or custom.

Are all 91 grammars available in a standard binary build?

Yes. All 91 tree-sitter grammars are compiled into the standard binary. The 89 registry grammars are linked as dependencies specified in the build configuration, while the 2 custom grammars (Magma and Form) are built from source located in the tools/ directory. No external downloads are required at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →