# How to Perform a Basic Extraction with Graphify: CLI and Python API Guide

> Learn to perform a basic extraction with Graphify using its CLI or Python API. Generate structured knowledge graphs from your code effortlessly with this comprehensive guide.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: how-to-guide
- Published: 2026-07-16

---

**To perform a basic extraction with Graphify, run `graphify extract <path>` from the CLI after installing with `uv tool install graphifyy`, or use the Python API `from graphify.ingest import ingest` followed by `ingest(path="your_project/")` to generate a structured knowledge graph from source code and documentation.**

Graphify is an open-source knowledge graph generator that transforms source code repositories into queryable structured data. The Graphify-Labs/graphify repository provides both a command-line interface and a Python API for extracting entities and relationships from code files and non-code assets like markdown and PDFs.

## Install Graphify

Before performing a basic extraction, install the tool using either `uv` or `pipx`:

```bash
uv tool install graphifyy

# or

pipx install graphifyy

```

## Method 1: Command-Line Extraction

The fastest way to perform a basic extraction with Graphify is through the CLI. Navigate to your project directory and run:

```bash
graphify extract .

```

This command initiates the file discovery process, walking the directory to build a manifest of files to process according to the logic in [`graphify/cli.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/cli.py).

## Understanding the Generated Output

After extraction completes, Graphify writes results to the `graphify-out/` directory:

- [`graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/graph.json) contains the raw graph data with nodes and edges
- [`GRAPH_REPORT.md`](https://github.com/Graphify-Labs/graphify/blob/main/GRAPH_REPORT.md) presents a human-readable report including god nodes, community clusters, and confidence tags (`EXTRACTED` vs. `INFERRED`)

View the interactive graph visualization:

```bash
graphify-open graphify-out/graph.html

```

Inspect the markdown report:

```bash
cat graphify-out/GRAPH_REPORT.md

```

## Method 2: Programmatic Extraction with Python

For integration into existing workflows, perform a basic extraction using the Python API. Import the `ingest` function from `graphify.ingest`:

```python
from graphify.ingest import ingest

# Ingest a single file or entire directory

graph = ingest(path="my_project/")
graph.save("graphify-out/graph.json")
print(graph.summary())  # High-level overview of extracted entities

```

The `ingest` function handles the complete pipeline: file discovery, parsing, and graph construction.

## Method 3: Direct Language-Specific Extractors

For fine-grained control over how specific programming languages are parsed, invoke the dedicated extractors directly from [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py):

```python
from graphify.extract import extract_python
from pathlib import Path

result = extract_python(Path("example.py"))
print(result["functions"])   # List of discovered functions

print(result["classes"])       # List of discovered classes

```

Available extractors include `extract_python`, `extract_js`, `extract_java`, and others, each implementing tree-sitter based AST parsing.

## The Extraction Pipeline Architecture

Understanding how Graphify processes files helps optimize your extraction workflows.

### File Discovery and Classification

Graphify begins by walking the target directory to identify source files. Files matching recognized code extensions are routed to tree-sitter parsers, while markdown, PDFs, and images are queued for LLM-based processing.

### Language-Specific Parsing

In [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py), each language-specific function parses the file's AST via tree-sitter. The extractor walks the AST to identify **nodes** (classes, functions, modules) and **edges** (calls, imports, inheritance), emitting dictionaries that describe discovered entities.

### Non-Code Asset Extraction

Files without tree-sitter parsers are handled by `extract_files_direct` in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py). This function uses language model inference to extract headings, code blocks, and referenced assets from documentation and binary files.

### Graph Construction and Deduplication

The [`graphify/build.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/build.py) module merges individual extraction results into a unified knowledge graph. The deduplication logic in [`graphify/dedup.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/dedup.py) ensures unique entities before persisting the final graph to [`graphify-out/graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/graphify-out/graph.json).

## Querying Your Knowledge Graph

After you perform a basic extraction with Graphify, explore relationships without grepping source files:

```bash

# Ask natural language questions

graphify query "How does the authentication module work?"

# Find paths between entities

graphify path module_A module_B

```

## Summary

- **Install Graphify** using `uv tool install graphifyy` or `pipx install graphifyy` to get the CLI and Python library
- **Run basic extraction** with `graphify extract <path>` from the command line or `ingest(path="<path>")` from Python
- **Access outputs** in `graphify-out/` including [`graph.json`](https://github.com/Graphify-Labs/graphify/blob/main/graph.json) for raw data and [`GRAPH_REPORT.md`](https://github.com/Graphify-Labs/graphify/blob/main/GRAPH_REPORT.md) for human-readable summaries
- **Leverage language-specific extractors** like `extract_python` in [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py) for custom parsing workflows
- **Query results** using `graphify query` or `graphify path` commands to explore codebase relationships

## Frequently Asked Questions

### What file types does Graphify support for extraction?

Graphify supports programming languages through tree-sitter grammars (Python, JavaScript, Java, etc.) via [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py), and non-code assets (markdown, PDFs, images) through LLM-based extraction in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py). The tool automatically classifies files during the discovery phase and routes them to the appropriate processor.

### How do I extract only specific file types?

Use the Python API to filter paths before ingestion, or invoke language-specific extractors directly from [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py) (e.g., `extract_python`, `extract_js`) for targeted processing. The CLI currently processes all recognized files in the provided directory.

### What is the difference between EXTRACTED and INFERRED tags in GRAPH_REPORT.md?

Tags marked **EXTRACTED** indicate relationships derived directly from static code analysis via tree-sitter parsing in [`graphify/extract.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/extract.py). **INFERRED** tags represent relationships generated by the LLM extractor (`extract_files_direct` in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py)) for non-code assets or ambiguous code patterns.

### Can I perform extraction without installing the CLI tool?

Yes. Import the extraction functions directly in Python: `from graphify.ingest import ingest` for full pipeline processing, or `from graphify.extract import extract_python` (and similar functions) for specific languages. This approach requires the graphify package installed via pip rather than as a standalone tool.