# What Programming Languages Does Codebase-Memory-MCP Support? Full List of 158 Languages

> Discover the 158 programming languages supported by Codebase-Memory-MCP for parsing. Learn about its advanced LSP semantic analysis capabilities.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: getting-started
- Published: 2026-07-04

---

**Codebase-Memory-MCP supports 158 programming languages via vendored tree-sitter grammars compiled into the static binary, with advanced Hybrid LSP semantic analysis for 10 major languages including Python, TypeScript, Java, and Rust.**

Codebase-Memory-MCP is a high-performance code indexing tool developed by DeusData that parses repositories into queryable graph structures. Understanding which programming languages Codebase-Memory-MCP supports is essential for teams evaluating the tool for polyglot codebases. The engine delivers broad syntactic coverage through tree-sitter while providing deep semantic analysis for the most widely-used languages.

## Complete Language Support Overview

### Universal Tree-sitter Coverage (158 Languages)

The core parsing engine supports **158 programming languages** through vendored tree-sitter grammars that ship inside the static binary. These grammars live in the `internal/cbm/vendored` directory and are compiled directly into the executable at build-time, ensuring zero external dependencies at runtime. According to the source code in [`internal/cbm/vendored/ts_runtime/src/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/vendored/ts_runtime/src/language.c), the runtime exposes a common C API (`ts_language_*` functions) that enables the parser to handle any supported language uniformly.

### Hybrid LSP Semantic Analysis (10 Languages)

Beyond syntax trees, Codebase-Memory-MCP implements a **Hybrid LSP layer** that provides type-aware resolution for ten strategically selected languages. These languages receive full semantic analysis including import resolution, generics tracking, and inheritance mapping:

- **Python**
- **TypeScript/JavaScript/JSX/TSX**
- **PHP**
- **C#**
- **Go**
- **C/C++**
- **Java**
- **Kotlin**
- **Rust**

For these languages, the engine runs a secondary analysis pass that augments the tree-sitter AST with semantic edges like `CALLS` and `IMPORTS`, enabling precise cross-file navigation and dependency tracing. The remaining 148 languages fall back to pure tree-sitter parsing with syntactic structure but without type-aware cross-references.

## How Language Parsing Works Internally

The parsing pipeline follows a multi-phase architecture defined in [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c) and related files:

1. **Discovery Phase**: The [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c) module scans the repository, applies ignore rules, and maps file extensions to language identifiers using compiled lookup tables.

2. **Tree-sitter Parsing**: For each discovered file, the engine selects the appropriate vendored grammar from the 158 available options. The [`language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/language.c) file in the tree-sitter runtime ([`internal/cbm/vendored/ts_runtime/src/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/vendored/ts_runtime/src/language.c)) provides the low-level interface for tokenizing and building concrete syntax trees.

3. **Hybrid LSP Augmentation**: When processing one of the ten supported languages, the engine performs semantic analysis to resolve types, imports, and symbol relationships, storing the results in an in-memory SQLite database.

The build process uses [`scripts/generate-lang-code.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/generate-lang-code.py) to generate the C tables that embed all 158 grammars into the final binary, as documented in the repository's build configuration.

## Querying Supported Languages via CLI

You can verify which languages Codebase-Memory-MCP detects in your repository using the command-line interface.

**List detected languages in a project:**

```bash
codebase-memory-mcp cli get_architecture '{"project":"myrepo"}' | jq '.languages'

```

**Index with full language support:**

```bash
codebase-memory-mcp cli index_repository '{"repo_path":"$(pwd)"}'

```

**Trace Python functions with type resolution:**

```bash
codebase-memory-mcp cli trace_path '{"function_name":"my_pkg.utils.process_data","direction":"both"}'

```

**Python SDK example:**

```python
from codebase_memory_mcp import MCPClient

client = MCPClient()
arch = client.get_architecture(project="myrepo")
print("Languages parsed:", arch["languages"])

```

## Summary

- Codebase-Memory-MCP supports **158 programming languages** through vendored tree-sitter grammars compiled into the static binary.
- The **Hybrid LSP layer** provides deep semantic analysis for Python, TypeScript/JavaScript/JSX/TSX, PHP, C#, Go, C/C++, Java, Kotlin, and Rust.
- Language detection and mapping logic resides in [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c), while the tree-sitter runtime API is implemented in [`internal/cbm/vendored/ts_runtime/src/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/vendored/ts_runtime/src/language.c).
- Build-time code generation at [`scripts/generate-lang-code.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/generate-lang-code.py) embeds all grammars into the executable.
- All languages receive syntactic parsing, but only the ten Hybrid LSP languages get type-aware cross-reference resolution.

## Frequently Asked Questions

### Does Codebase-Memory-MCP require external language servers to parse code?

No. Codebase-Memory-MCP compiles all 158 tree-sitter grammars directly into the static binary, as managed in the `internal/cbm/vendored` directory. The tool operates with zero external dependencies for parsing, though the Hybrid LSP layer provides additional semantic analysis for supported languages without requiring separate language server processes.

### How does Codebase-Memory-MCP handle unsupported or niche programming languages?

The engine falls back to generic tree-sitter parsing for any of the 148 languages outside the Hybrid LSP subset. While these languages receive full syntactic analysis and AST generation, they lack type-aware resolution for imports, generics, and inheritance. The concrete syntax tree still enables structural search and navigation, but cross-file semantic queries may be limited.

### Where is the language support configuration defined in the source code?

Language support is hard-coded at build-time via [`scripts/generate-lang-code.py`](https://github.com/DeusData/codebase-memory-mcp/blob/main/scripts/generate-lang-code.py), which generates C tables embedding the grammars. Runtime language detection logic lives in [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c), while the tree-sitter runtime interface is defined in [`internal/cbm/vendored/ts_runtime/src/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/vendored/ts_runtime/src/language.c). The README documents the full tier list and Hybrid LSP coverage.

### Can I add support for a new programming language to Codebase-Memory-MCP?

Adding support requires vendoring a new tree-sitter grammar into the `internal/cbm/vendored` directory and regenerating the language tables using the build scripts. The architecture in [`src/discover/language.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/discover/language.c) supports arbitrary language registration as long as a compatible grammar exists that implements the `ts_language_*` API defined in the runtime.