How Graphify Parses Code Without LLMs: Tree-Sitter Static Analysis Explained
Graphify v8 extracts structural information from source files using deterministic Tree-Sitter parsers instead of probabilistic large language models, ensuring fast and language-accurate code analysis.
Graphify is an open-source code analysis tool that builds graph representations of source code without relying on AI inference. Unlike LLM-based approaches that require expensive GPU resources and prompt engineering, Graphify leverages Tree-Sitter to parse code without LLMs. This architecture provides deterministic results while supporting multiple programming languages through modular parser bindings.
Language-Specific Tree-Sitter Modules
Graphify dynamically imports Tree-Sitter grammar modules at runtime based on the source file extension. For each supported language, the system loads the corresponding tree_sitter_* package (e.g., tree_sitter_python, tree_sitter_javascript, tree_sitter_rust) through dedicated extractor modules located under graphify/extractors/.
The Python-specific logic resides in graphify/extractors/python.py, while Rust and SQL implementations live in graphify/extractors/rust.py and graphify/extractors/sql.py respectively. If a required parser is missing, Graphify returns a graceful "not installed" error rather than attempting fallback inference. This behavior is validated in the test suite within tests/test_extract.py and tests/test_security.py.
Parser Creation and AST Traversal
Each extractor instantiates a tree_sitter.Parser and loads a compiled language object via tree_sitter.Language. The language grammars are compiled from upstream Tree-Sitter sources and packaged as shared libraries (e.g., tree_sitter_python.so).
After parsing, Graphify walks the Concrete Syntax Tree (CST) to identify language-specific constructs such as classes, functions, imports, and calls. The traversal logic is implemented in extractor functions like extract_python defined in graphify/extract.py. These functions normalize parsed nodes and edges into Graphify’s generic graph model, ensuring consistent representation across languages.
# Example: parsing a Python file with Graphify
from pathlib import Path
from graphify.extract import extract_python
source_path = Path("example.py")
graph = extract_python(source_path)
print("Nodes:", len(graph["nodes"]))
print("Edges:", len(graph["edges"]))
Rationale Nodes and Static Analysis
For Python specifically, Graphify performs an additional analysis pass through _extract_python_rationale to generate rationale nodes. These nodes document why particular edges exist in the graph—for example, marking an import-based reference between modules.
This analysis is pure static analysis, examining the CST structure without consulting probabilistic language models. The deterministic nature of Tree-Sitter ensures that identical source code always produces identical graph structures, eliminating the non-determinism inherent in LLM-based parsing approaches.
Extending Graphify to New Languages
Adding support for new languages requires only a Tree-Sitter grammar and a thin Python wrapper. The dispatch table in graphify/extract.py maps file extensions and language identifiers to their respective extractors (e.g., ".py" → extract_python, "python" → extract_python).
New extractors are placed in graphify/extractors/<lang>.py and must implement the standard interface for AST traversal and graph normalization.
# Example: adding a new language (e.g., Go) – only a thin wrapper is needed
# graphify/extractors/go.py
from tree_sitter import Language, Parser
ts_go = Language("build/my-languages.so", "go")
parser = Parser()
parser.set_language(ts_go)
def extract_go(path: Path) -> dict:
tree = parser.parse(path.read_bytes())
# Walk the tree and return Graphify‑compatible dict
...
# Build the shared library containing all needed grammars
tree-sitter build-wasm ./tree-sitter-python ./tree-sitter-go
# Then install the Python bindings
pip install tree-sitter
Summary
- Tree-Sitter dependency: Graphify relies entirely on Tree-Sitter parsers rather than LLMs for code analysis, as implemented in the Graphify-Labs/graphify repository.
- Deterministic parsing: The use of compiled grammars and concrete syntax trees ensures consistent, reproducible results across identical source files.
- Modular architecture: Language support is extended through isolated extractor modules in
graphify/extractors/with minimal boilerplate code. - Static analysis only: Features like rationale nodes in Python are derived from AST structure without probabilistic inference.
Frequently Asked Questions
Does Graphify use any machine learning models for parsing?
No. Graphify v8 uses only Tree-Sitter parsers to analyze source code. The parsing pipeline in graphify/extract.py and its associated extractor modules perform deterministic static analysis without invoking neural networks or language models.
What happens if a Tree-Sitter parser is not installed?
Graphify implements graceful error handling when a required tree_sitter_* package is missing. Rather than attempting to parse with incomplete tools or falling back to LLM-based inference, the system returns a clear "not installed" message. This behavior is rigorously tested in tests/test_extract.py and tests/test_security.py.
How does Graphify handle multiple programming languages?
The dispatch table in graphify/extract.py maps file extensions to language-specific extractor functions. Each extractor imports its corresponding Tree-Sitter grammar module and implements standardized AST traversal logic. This architecture allows Graphify to support Python, Rust, JavaScript, SQL, and other languages through independent, composable modules.
Is Tree-Sitter parsing faster than LLM-based code analysis?
Yes. Because Tree-Sitter produces deterministic parse trees through compiled grammars rather than autoregressive token generation, Graphify’s parsing is significantly faster and more resource-efficient than LLM-based approaches. The process requires no GPU acceleration and executes in milliseconds for typical source files, while maintaining complete language accuracy according to formal grammar specifications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →