Code-Graph-RAG Supported Languages for Code Optimization: Complete 2024 Guide
Code-Graph-RAG supports 14 programming languages for code optimization, including Python, JavaScript, TypeScript, C, C++, Java, Go, Rust, PHP, C#, Dart, Lua, Scala (in development), and SQL (in development), with language-specific parsing and call-graph construction managed through the central SupportedLanguage enumeration.
The vitali87/code-graph-rag repository provides a multi-language optimization engine that ingests source code, constructs precise call graphs, and applies performance profiling across diverse technology stacks. Understanding which languages supported for code optimization with Code-Graph-RAG are available—and how the system handles each—is essential for leveraging its full static analysis capabilities.
Complete List of Supported Languages
Code-Graph-RAG defines its language support centrally in codebase_rag/constants/languages.py【L54-L70】, utilizing a comprehensive enumeration that tracks implementation status and capabilities. The following table outlines the current support matrix:
| Language Enum | Display Name | Status | Key Optimization Capabilities |
|---|---|---|---|
PYTHON |
Python | Fully Supported | Type inference, decorators, nested functions |
JS |
JavaScript | Fully Supported | ES6 modules, CommonJS, prototype and arrow functions |
TS |
TypeScript | Fully Supported | Interfaces, type aliases, enums, namespaces, ES6/CommonJS modules |
TSX |
TypeScript (TSX) | Fully Supported | All TypeScript features plus JSX elements & components |
C |
C | Fully Supported | Functions, structs, unions, enums, pre-processor includes |
CPP |
C++ | Fully Supported | Constructors, destructors, operator overloading, templates, C++20 modules, namespaces, macros |
LUA |
Lua | Fully Supported | Local/global functions, metatables, closures, coroutines |
RUST |
Rust | Fully Supported | impl blocks, associated functions, macro_rules! macros |
JAVA |
Java | Fully Supported | Generics, annotations, records, sealed classes, concurrency, reflection |
GO |
Go | Fully Supported | Receiver methods, structs, interfaces, type declarations, function-local types |
SCALA |
Scala | In Development | Case classes, objects |
PHP |
PHP | Fully Supported | Classes, interfaces, traits, enums, namespaces, PHP 8 attributes |
CSHARP |
C# | Fully Supported | Namespaces (block & file-scoped), classes/structs/records/interfaces/enums, generics, inheritance, overload resolution |
| DART | Dart | Fully Supported | Classes, mixins, extensions, enhanced enums, factory/named constructors, Flutter widgets |
| SQL | SQL (PostgreSQL) | In Development | Stored functions, schema-qualified names (full PL/pgSQL support pending) |
Detailed metadata for these languages, including parser configurations and display properties, is defined in the same constants file【L88-L122】.
How Language Support Is Implemented
The optimization pipeline relies on two core components: the enumeration definition and the language specification registry.
The SupportedLanguage Enumeration
At the heart of the system lies the SupportedLanguage enum in codebase_rag/constants/languages.py. This class acts as the single source of truth for all language-specific logic, determining which Tree-Sitter grammars to load and which optimization passes to apply. Each variant maps to a LanguageSpec object that describes node types for functions, classes, modules, calls, and imports.
Language Detection and Parser Selection
When analyzing a repository, Code-Graph-RAG performs automatic language detection through the LANGUAGE_SPECS registry defined in codebase_rag/language_spec.py【L30-L46】. The system maps file extensions to SupportedLanguage values and initializes the appropriate TreeSitterModule for parsing.
The detection workflow follows these steps:
- Extension Mapping – File extensions (e.g.,
.py,.js,.cpp) are matched against the registry viaget_language_spec() - Parser Initialization – The corresponding Tree-Sitter grammar is loaded based on the
LanguageSpecdefinition - Symbol Extraction – Language-specific node types enable precise extraction of functions, classes, and imports
- Call-Graph Construction – The engine builds a language-aware call graph for dependency analysis
Optimization Passes
Optimization scripts iterate over LANGUAGE_SPECS.values() to apply analyses uniformly across all supported languages. For example, optimize/profile_io.py【L375-L378】and optimize/memory_profile.py implement profiling passes that respect language-specific semantics while measuring I/O operations and memory usage.
Language-Specific Optimization Capabilities
Each supported language receives tailored treatment in the optimization pipeline:
- Python – Handles decorators and nested functions for accurate scope analysis
- JavaScript/TypeScript – Distinguishes between ES6 modules and CommonJS, supporting both prototype-based and modern arrow function patterns
- C++ – Processes templates, operator overloading, and C++20 modules alongside traditional macros
- Rust – Parses
implblocks andmacro_rules!macros for systems-level optimization - Java – Supports modern features including records, sealed classes, and generics
- Go – Handles receiver methods and interface satisfaction for concurrent code analysis
- PHP – Recognizes PHP 8 attributes, traits, and namespace hierarchies
- C# – Distinguishes between block-scoped and file-scoped namespaces, supporting advanced generics and overload resolution
Languages marked "In Development" (Scala and SQL) have partial support with active implementation of call-graph construction and symbol extraction.
Working with Language Support in Code
Developers can interact with the language system programmatically using the provided API.
Retrieving Language Specifications
Use get_language_spec() to access metadata for specific file extensions:
from codebase_rag.language_spec import get_language_spec
# Get the spec for a Python file
python_spec = get_language_spec(".py")
print(python_spec.language) # → SupportedLanguage.PYTHON
print(python_spec.file_extensions) # → ('.py',)
Running Multi-Language Optimization
Execute the optimization pipeline across all detected languages in a repository:
# Assuming the repository root is ./my_repo
python -m codegraph_rag.main \
--repo ./my_repo \
--optimise # triggers the optimisation pipeline for all detected languages
The engine automatically detects present languages, loads appropriate parsers, constructs call graphs, and applies language-specific optimization passes including I/O profiling and memory analysis.
Summary
- Code-Graph-RAG supports 14 languages ranging from systems languages (C, C++, Rust) to web technologies (JavaScript, TypeScript, PHP) and mobile frameworks (Dart/Flutter)
- Centralized configuration lives in
codebase_rag/constants/languages.pyvia theSupportedLanguageenumeration - Automatic detection occurs through
LANGUAGE_SPECSincodebase_rag/language_spec.py, mapping file extensions to parser configurations - Optimization passes in
optimize/profile_io.pyandoptimize/memory_profile.pyiterate across all supported languages - 12 languages are fully supported, while Scala and SQL (PostgreSQL) remain in active development
Frequently Asked Questions
How does Code-Graph-RAG detect which programming language a file uses?
The system uses the get_language_spec() function in codebase_rag/language_spec.py to map file extensions to SupportedLanguage enum values. When scanning a repository, each file extension is checked against the LANGUAGE_SPECS registry, which then determines which Tree-Sitter grammar and optimization passes to apply.
Can Code-Graph-RAG optimize codebases containing multiple programming languages?
Yes. The optimization engine iterates over LANGUAGE_SPECS.values() and processes each detected language individually. Scripts like optimize/profile_io.py explicitly loop through all supported languages, applying I/O and memory profiling across heterogeneous repositories containing, for example, both Python and JavaScript or Java and Go.
Which modern language features are supported in the C++ and Java implementations?
For C++, the engine supports C++20 modules, templates, operator overloading, constructors/destructors, and namespaces. For Java, it handles generics, annotations, records (Java 14+), sealed classes, and concurrency primitives. These capabilities enable accurate call-graph construction even for codebases using cutting-edge language standards.
Are dynamically typed languages like JavaScript and Python analyzed differently than statically typed ones?
Yes. While both receive full support, the analysis adapts to language paradigms. JavaScript optimization handles ES6 modules, CommonJS, and prototype chains, while Python focuses on decorators, type inference hints, and nested function scopes. The LanguageSpec objects in language_spec.py define language-specific node types that guide the parser in extracting meaningful symbols regardless of typing discipline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →