What Kind of Output Does code-graph-rag Produce? A Complete Guide to Graphs, Vectors, and RAG Results
code-graph-rag produces four primary output types: a directed code-knowledge graph representing entities and relationships, dense embedding vectors stored in FAISS-compatible indices, natural language answers with source code citations generated via RAG retrieval, and optional diagnostic logs for pipeline debugging.
code-graph-rag is an open-source Python library that analyzes source code repositories to enable semantic search and automated reasoning. The vitali87/code-graph-rag project transforms raw codebases into structured, queryable data formats that power downstream retrieval-augmented generation (RAG) applications. Understanding the specific code-graph-rag output formats is essential for integrating the tool into documentation workflows, IDE plugins, or code review automation systems.
Overview of code-graph-rag Output Types
When you process a repository through the pipeline, the system generates distinct artifacts at each stage:
- Graph Data: A directed, typed graph with nodes for functions, classes, modules, and variables, connected by edges representing calls, imports, inheritance, and references.
- Embedding Vectors: High-dimensional vector representations of each graph node (and optionally source files) optimized for similarity search.
- RAG Results: Textual answers to natural language queries, accompanied by citations linking back to specific nodes and code snippets.
- Diagnostic Logs: Human-readable execution traces showing retrieval decisions, cache hits, and pipeline timing when running in verbose mode.
All outputs are generated programmatically as Python objects, with optional serialization to JSON, GraphML, DOT, or FAISS indices.
Graph Data: NetworkX Objects and Serialization
The core code-graph-rag output is a rich knowledge graph built by parsing the entire codebase. In codebase_rag/cg_graph.py, the CodeGraph class orchestrates this construction.
The graph captures:
- Nodes: Code entities with metadata including name, type, file location, and docstrings.
- Edges: Typed relationships such as
calls,imports,inherits,defines, andreferences.
You can access this output programmatically as a NetworkX DiGraph object or export it to standard formats.
from codebase_rag.cg_graph import CodeGraph
from pathlib import Path
# Initialize and build the graph
cg = CodeGraph(root_dir=Path("my_project"))
cg.build() # Parses the codebase
# Access as NetworkX object
graph = cg.to_networkx() # Returns a NetworkX DiGraph
# Serialize to JSON for downstream tools
cg.export_json(Path("graph.json")) # See examples/graph_export_example.py
The export_json() method creates a portable representation suitable for visualization tools or external graph databases. For GraphViz visualization, the system also supports DOT format exports, demonstrated in examples/graph_export_example.py.
Embedding Vectors: FAISS Indices and NumPy Arrays
To enable semantic search, code-graph-rag generates dense vector embeddings for each node. The codebase_rag/embedder.py module handles this transformation, using either OpenAI models or local embedding engines.
The embedding output consists of:
- NumPy arrays containing high-dimensional vectors for each node.
- FAISS-compatible indices that enable fast approximate nearest neighbor search.
The codebase_rag/vector_store.py module manages persistence and retrieval of these vectors.
from codebase_rag.embedder import Embedder
from codebase_rag.vector_store import VectorStore
from pathlib import Path
# Generate embeddings for all graph nodes
embedder = Embedder()
vectors = embedder.embed_nodes(cg.nodes) # Returns np.ndarray of shape (n_nodes, dim)
# Store in FAISS index
store = VectorStore()
store.add_vectors(vectors, cg.nodes)
store.save(Path("faiss.index")) # Persisted binary index for later retrieval
This vector store output allows the system to perform similarity searches across the codebase using natural language queries, mapping questions to relevant code entities without keyword matching.
RAG Answers: Natural Language with Source Citations
The primary consumer-facing code-graph-rag output is the RAG result: a synthesized answer to a natural language question, backed by specific code citations. The codebase_rag/cli.py module (exposed through main.py) orchestrates this retrieval and generation pipeline.
When you submit a query, the system:
- Loads the FAISS index and graph structure.
- Retrieves the most relevant nodes via vector similarity.
- Assembles a contextual answer citing specific functions, classes, or files.
from codebase_rag.cli import query
from pathlib import Path
# Query the codebase
answer, citations = query(
query="How does the user authentication flow work?",
index_path=Path("faiss.index"),
graph_path=Path("graph.json")
)
print("Answer:", answer)
print("Citations:", citations) # List of relevant nodes with source locations
The answer return value is a plain string containing the generated response. The citations return value provides a structured list of referenced nodes, enabling users to verify answers against the actual source code.
Diagnostic Logging and Verbose Output
For debugging and optimization, code-graph-rag optionally produces diagnostic logs that expose the internal retrieval decisions. When you invoke the CLI with the --verbose flag, the system writes human-readable log lines to stdout and stderr.
These logs reveal:
- Which graph nodes were retrieved for a given query and why.
- Cache hit/miss statistics for embedding operations.
- Timing metrics for graph construction and vector search phases.
This output is invaluable when tuning retrieval parameters or verifying that the RAG pipeline correctly identifies relevant code paths.
Summary
- Graph Data: Directed knowledge graphs built in
codebase_rag/cg_graph.py, exportable as JSON, GraphML, or DOT via methods liketo_networkx()andexport_json(). - Embeddings: NumPy arrays and FAISS indices generated by
embedder.embed_nodes()and managed throughcodebase_rag/vector_store.py. - RAG Results: Natural language answers with precise citations sourced from the graph, produced by the
query()function incodebase_rag/cli.py. - Logs: Optional verbose logging for pipeline inspection and performance tuning.
All outputs are available as Python objects for programmatic use, with explicit serialization methods for filesystem persistence.
Frequently Asked Questions
What file formats does code-graph-rag support for graph export?
According to the source code in codebase_rag/cg_graph.py and demonstrated in examples/graph_export_example.py, the system supports exporting the knowledge graph as JSON for data interchange, GraphML for graph database import, and DOT for GraphViz visualization. The export_json() method provides the most common integration path for web-based visualizations.
How are the embedding vectors stored and retrieved?
The codebase_rag/vector_store.py implementation persists vectors as FAISS-compatible binary indices, which enable millisecond-scale similarity search across thousands of code entities. These indices store NumPy arrays alongside node metadata, allowing the system to map vector search results back to specific graph nodes and source file locations.
Can I use code-graph-rag outputs without the RAG question-answering feature?
Yes. The library is modular: you can generate the graph output and embedding vectors independently of the retrieval pipeline. The CodeGraph class in codebase_rag/cg_graph.py and the Embedder class in codebase_rag/embedder.py function as standalone components, allowing you to use the knowledge graph for static analysis, dependency mapping, or custom search implementations without invoking the natural language generation features.
Where does the RAG pipeline store its final answer output?
The query() function in codebase_rag/cli.py returns results as a Python tuple (answer, citations) rather than writing to disk. The answer is a string containing the generated response, while citations is a structured list of referenced graph nodes. If you need to persist these results, you must handle file I/O in your calling code, as the library does not auto-save RAG outputs to the filesystem.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →