What Programming Languages Does Codebase-Memory-MCP Support? Full List of 158 Languages

Codebase-Memory-MCP supports 158 programming languages via vendored tree-sitter grammars compiled into the static binary, with advanced Hybrid LSP semantic analysis for 10 major languages including Python, TypeScript, Java, and Rust.

Codebase-Memory-MCP is a high-performance code indexing tool developed by DeusData that parses repositories into queryable graph structures. Understanding which programming languages Codebase-Memory-MCP supports is essential for teams evaluating the tool for polyglot codebases. The engine delivers broad syntactic coverage through tree-sitter while providing deep semantic analysis for the most widely-used languages.

Complete Language Support Overview

Universal Tree-sitter Coverage (158 Languages)

The core parsing engine supports 158 programming languages through vendored tree-sitter grammars that ship inside the static binary. These grammars live in the internal/cbm/vendored directory and are compiled directly into the executable at build-time, ensuring zero external dependencies at runtime. According to the source code in internal/cbm/vendored/ts_runtime/src/language.c, the runtime exposes a common C API (ts_language_* functions) that enables the parser to handle any supported language uniformly.

Hybrid LSP Semantic Analysis (10 Languages)

Beyond syntax trees, Codebase-Memory-MCP implements a Hybrid LSP layer that provides type-aware resolution for ten strategically selected languages. These languages receive full semantic analysis including import resolution, generics tracking, and inheritance mapping:

  • Python
  • TypeScript/JavaScript/JSX/TSX
  • PHP
  • C#
  • Go
  • C/C++
  • Java
  • Kotlin
  • Rust

For these languages, the engine runs a secondary analysis pass that augments the tree-sitter AST with semantic edges like CALLS and IMPORTS, enabling precise cross-file navigation and dependency tracing. The remaining 148 languages fall back to pure tree-sitter parsing with syntactic structure but without type-aware cross-references.

How Language Parsing Works Internally

The parsing pipeline follows a multi-phase architecture defined in src/discover/language.c and related files:

  1. Discovery Phase: The src/discover/language.c module scans the repository, applies ignore rules, and maps file extensions to language identifiers using compiled lookup tables.

  2. Tree-sitter Parsing: For each discovered file, the engine selects the appropriate vendored grammar from the 158 available options. The language.c file in the tree-sitter runtime (internal/cbm/vendored/ts_runtime/src/language.c) provides the low-level interface for tokenizing and building concrete syntax trees.

  3. Hybrid LSP Augmentation: When processing one of the ten supported languages, the engine performs semantic analysis to resolve types, imports, and symbol relationships, storing the results in an in-memory SQLite database.

The build process uses scripts/generate-lang-code.py to generate the C tables that embed all 158 grammars into the final binary, as documented in the repository's build configuration.

Querying Supported Languages via CLI

You can verify which languages Codebase-Memory-MCP detects in your repository using the command-line interface.

List detected languages in a project:

codebase-memory-mcp cli get_architecture '{"project":"myrepo"}' | jq '.languages'

Index with full language support:

codebase-memory-mcp cli index_repository '{"repo_path":"$(pwd)"}'

Trace Python functions with type resolution:

codebase-memory-mcp cli trace_path '{"function_name":"my_pkg.utils.process_data","direction":"both"}'

Python SDK example:

from codebase_memory_mcp import MCPClient

client = MCPClient()
arch = client.get_architecture(project="myrepo")
print("Languages parsed:", arch["languages"])

Summary

  • Codebase-Memory-MCP supports 158 programming languages through vendored tree-sitter grammars compiled into the static binary.
  • The Hybrid LSP layer provides deep semantic analysis for Python, TypeScript/JavaScript/JSX/TSX, PHP, C#, Go, C/C++, Java, Kotlin, and Rust.
  • Language detection and mapping logic resides in src/discover/language.c, while the tree-sitter runtime API is implemented in internal/cbm/vendored/ts_runtime/src/language.c.
  • Build-time code generation at scripts/generate-lang-code.py embeds all grammars into the executable.
  • All languages receive syntactic parsing, but only the ten Hybrid LSP languages get type-aware cross-reference resolution.

Frequently Asked Questions

Does Codebase-Memory-MCP require external language servers to parse code?

No. Codebase-Memory-MCP compiles all 158 tree-sitter grammars directly into the static binary, as managed in the internal/cbm/vendored directory. The tool operates with zero external dependencies for parsing, though the Hybrid LSP layer provides additional semantic analysis for supported languages without requiring separate language server processes.

How does Codebase-Memory-MCP handle unsupported or niche programming languages?

The engine falls back to generic tree-sitter parsing for any of the 148 languages outside the Hybrid LSP subset. While these languages receive full syntactic analysis and AST generation, they lack type-aware resolution for imports, generics, and inheritance. The concrete syntax tree still enables structural search and navigation, but cross-file semantic queries may be limited.

Where is the language support configuration defined in the source code?

Language support is hard-coded at build-time via scripts/generate-lang-code.py, which generates C tables embedding the grammars. Runtime language detection logic lives in src/discover/language.c, while the tree-sitter runtime interface is defined in internal/cbm/vendored/ts_runtime/src/language.c. The README documents the full tier list and Hybrid LSP coverage.

Can I add support for a new programming language to Codebase-Memory-MCP?

Adding support requires vendoring a new tree-sitter grammar into the internal/cbm/vendored directory and regenerating the language tables using the build scripts. The architecture in src/discover/language.c supports arbitrary language registration as long as a compatible grammar exists that implements the ts_language_* API defined in the runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →