How to Perform a Basic Extraction with Graphify: CLI and Python API Guide
To perform a basic extraction with Graphify, run graphify extract <path> from the CLI after installing with uv tool install graphifyy, or use the Python API from graphify.ingest import ingest followed by ingest(path="your_project/") to generate a structured knowledge graph from source code and documentation.
Graphify is an open-source knowledge graph generator that transforms source code repositories into queryable structured data. The Graphify-Labs/graphify repository provides both a command-line interface and a Python API for extracting entities and relationships from code files and non-code assets like markdown and PDFs.
Install Graphify
Before performing a basic extraction, install the tool using either uv or pipx:
uv tool install graphifyy
# or
pipx install graphifyy
Method 1: Command-Line Extraction
The fastest way to perform a basic extraction with Graphify is through the CLI. Navigate to your project directory and run:
graphify extract .
This command initiates the file discovery process, walking the directory to build a manifest of files to process according to the logic in graphify/cli.py.
Understanding the Generated Output
After extraction completes, Graphify writes results to the graphify-out/ directory:
graph.jsoncontains the raw graph data with nodes and edgesGRAPH_REPORT.mdpresents a human-readable report including god nodes, community clusters, and confidence tags (EXTRACTEDvs.INFERRED)
View the interactive graph visualization:
graphify-open graphify-out/graph.html
Inspect the markdown report:
cat graphify-out/GRAPH_REPORT.md
Method 2: Programmatic Extraction with Python
For integration into existing workflows, perform a basic extraction using the Python API. Import the ingest function from graphify.ingest:
from graphify.ingest import ingest
# Ingest a single file or entire directory
graph = ingest(path="my_project/")
graph.save("graphify-out/graph.json")
print(graph.summary()) # High-level overview of extracted entities
The ingest function handles the complete pipeline: file discovery, parsing, and graph construction.
Method 3: Direct Language-Specific Extractors
For fine-grained control over how specific programming languages are parsed, invoke the dedicated extractors directly from graphify/extract.py:
from graphify.extract import extract_python
from pathlib import Path
result = extract_python(Path("example.py"))
print(result["functions"]) # List of discovered functions
print(result["classes"]) # List of discovered classes
Available extractors include extract_python, extract_js, extract_java, and others, each implementing tree-sitter based AST parsing.
The Extraction Pipeline Architecture
Understanding how Graphify processes files helps optimize your extraction workflows.
File Discovery and Classification
Graphify begins by walking the target directory to identify source files. Files matching recognized code extensions are routed to tree-sitter parsers, while markdown, PDFs, and images are queued for LLM-based processing.
Language-Specific Parsing
In graphify/extract.py, each language-specific function parses the file's AST via tree-sitter. The extractor walks the AST to identify nodes (classes, functions, modules) and edges (calls, imports, inheritance), emitting dictionaries that describe discovered entities.
Non-Code Asset Extraction
Files without tree-sitter parsers are handled by extract_files_direct in graphify/llm.py. This function uses language model inference to extract headings, code blocks, and referenced assets from documentation and binary files.
Graph Construction and Deduplication
The graphify/build.py module merges individual extraction results into a unified knowledge graph. The deduplication logic in graphify/dedup.py ensures unique entities before persisting the final graph to graphify-out/graph.json.
Querying Your Knowledge Graph
After you perform a basic extraction with Graphify, explore relationships without grepping source files:
# Ask natural language questions
graphify query "How does the authentication module work?"
# Find paths between entities
graphify path module_A module_B
Summary
- Install Graphify using
uv tool install graphifyyorpipx install graphifyyto get the CLI and Python library - Run basic extraction with
graphify extract <path>from the command line oringest(path="<path>")from Python - Access outputs in
graphify-out/includinggraph.jsonfor raw data andGRAPH_REPORT.mdfor human-readable summaries - Leverage language-specific extractors like
extract_pythoningraphify/extract.pyfor custom parsing workflows - Query results using
graphify queryorgraphify pathcommands to explore codebase relationships
Frequently Asked Questions
What file types does Graphify support for extraction?
Graphify supports programming languages through tree-sitter grammars (Python, JavaScript, Java, etc.) via graphify/extract.py, and non-code assets (markdown, PDFs, images) through LLM-based extraction in graphify/llm.py. The tool automatically classifies files during the discovery phase and routes them to the appropriate processor.
How do I extract only specific file types?
Use the Python API to filter paths before ingestion, or invoke language-specific extractors directly from graphify/extract.py (e.g., extract_python, extract_js) for targeted processing. The CLI currently processes all recognized files in the provided directory.
What is the difference between EXTRACTED and INFERRED tags in GRAPH_REPORT.md?
Tags marked EXTRACTED indicate relationships derived directly from static code analysis via tree-sitter parsing in graphify/extract.py. INFERRED tags represent relationships generated by the LLM extractor (extract_files_direct in graphify/llm.py) for non-code assets or ambiguous code patterns.
Can I perform extraction without installing the CLI tool?
Yes. Import the extraction functions directly in Python: from graphify.ingest import ingest for full pipeline processing, or from graphify.extract import extract_python (and similar functions) for specific languages. This approach requires the graphify package installed via pip rather than as a standalone tool.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →