How to Perform a Basic Extraction with Graphify: CLI and Python API Guide

To perform a basic extraction with Graphify, run graphify extract <path> from the CLI after installing with uv tool install graphifyy, or use the Python API from graphify.ingest import ingest followed by ingest(path="your_project/") to generate a structured knowledge graph from source code and documentation.

Graphify is an open-source knowledge graph generator that transforms source code repositories into queryable structured data. The Graphify-Labs/graphify repository provides both a command-line interface and a Python API for extracting entities and relationships from code files and non-code assets like markdown and PDFs.

Install Graphify

Before performing a basic extraction, install the tool using either uv or pipx:

uv tool install graphifyy

# or

pipx install graphifyy

Method 1: Command-Line Extraction

The fastest way to perform a basic extraction with Graphify is through the CLI. Navigate to your project directory and run:

graphify extract .

This command initiates the file discovery process, walking the directory to build a manifest of files to process according to the logic in graphify/cli.py.

Understanding the Generated Output

After extraction completes, Graphify writes results to the graphify-out/ directory:

  • graph.json contains the raw graph data with nodes and edges
  • GRAPH_REPORT.md presents a human-readable report including god nodes, community clusters, and confidence tags (EXTRACTED vs. INFERRED)

View the interactive graph visualization:

graphify-open graphify-out/graph.html

Inspect the markdown report:

cat graphify-out/GRAPH_REPORT.md

Method 2: Programmatic Extraction with Python

For integration into existing workflows, perform a basic extraction using the Python API. Import the ingest function from graphify.ingest:

from graphify.ingest import ingest

# Ingest a single file or entire directory

graph = ingest(path="my_project/")
graph.save("graphify-out/graph.json")
print(graph.summary())  # High-level overview of extracted entities

The ingest function handles the complete pipeline: file discovery, parsing, and graph construction.

Method 3: Direct Language-Specific Extractors

For fine-grained control over how specific programming languages are parsed, invoke the dedicated extractors directly from graphify/extract.py:

from graphify.extract import extract_python
from pathlib import Path

result = extract_python(Path("example.py"))
print(result["functions"])   # List of discovered functions

print(result["classes"])       # List of discovered classes

Available extractors include extract_python, extract_js, extract_java, and others, each implementing tree-sitter based AST parsing.

The Extraction Pipeline Architecture

Understanding how Graphify processes files helps optimize your extraction workflows.

File Discovery and Classification

Graphify begins by walking the target directory to identify source files. Files matching recognized code extensions are routed to tree-sitter parsers, while markdown, PDFs, and images are queued for LLM-based processing.

Language-Specific Parsing

In graphify/extract.py, each language-specific function parses the file's AST via tree-sitter. The extractor walks the AST to identify nodes (classes, functions, modules) and edges (calls, imports, inheritance), emitting dictionaries that describe discovered entities.

Non-Code Asset Extraction

Files without tree-sitter parsers are handled by extract_files_direct in graphify/llm.py. This function uses language model inference to extract headings, code blocks, and referenced assets from documentation and binary files.

Graph Construction and Deduplication

The graphify/build.py module merges individual extraction results into a unified knowledge graph. The deduplication logic in graphify/dedup.py ensures unique entities before persisting the final graph to graphify-out/graph.json.

Querying Your Knowledge Graph

After you perform a basic extraction with Graphify, explore relationships without grepping source files:


# Ask natural language questions

graphify query "How does the authentication module work?"

# Find paths between entities

graphify path module_A module_B

Summary

  • Install Graphify using uv tool install graphifyy or pipx install graphifyy to get the CLI and Python library
  • Run basic extraction with graphify extract <path> from the command line or ingest(path="<path>") from Python
  • Access outputs in graphify-out/ including graph.json for raw data and GRAPH_REPORT.md for human-readable summaries
  • Leverage language-specific extractors like extract_python in graphify/extract.py for custom parsing workflows
  • Query results using graphify query or graphify path commands to explore codebase relationships

Frequently Asked Questions

What file types does Graphify support for extraction?

Graphify supports programming languages through tree-sitter grammars (Python, JavaScript, Java, etc.) via graphify/extract.py, and non-code assets (markdown, PDFs, images) through LLM-based extraction in graphify/llm.py. The tool automatically classifies files during the discovery phase and routes them to the appropriate processor.

How do I extract only specific file types?

Use the Python API to filter paths before ingestion, or invoke language-specific extractors directly from graphify/extract.py (e.g., extract_python, extract_js) for targeted processing. The CLI currently processes all recognized files in the provided directory.

What is the difference between EXTRACTED and INFERRED tags in GRAPH_REPORT.md?

Tags marked EXTRACTED indicate relationships derived directly from static code analysis via tree-sitter parsing in graphify/extract.py. INFERRED tags represent relationships generated by the LLM extractor (extract_files_direct in graphify/llm.py) for non-code assets or ambiguous code patterns.

Can I perform extraction without installing the CLI tool?

Yes. Import the extraction functions directly in Python: from graphify.ingest import ingest for full pipeline processing, or from graphify.extract import extract_python (and similar functions) for specific languages. This approach requires the graphify package installed via pip rather than as a standalone tool.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →