How to Use the code-graph-rag CLI: Setup, Commands, and Configuration
The code-graph-rag CLI (cgr) enables you to ingest Git repositories, build Neo4j-compatible code graphs, and execute RAG queries via command-line flags such as --repo, --embedder, and --vector-store, with configuration managed through environment variables defined in codebase_rag/workspaces/constants.py.
The vitali87/code-graph-rag repository provides a Python-based framework for converting source code into a queryable graph structure. The CLI serves as the primary interface for orchestrating the ingestion pipeline, handling everything from repository cloning to vector store population.
Prerequisites and Installation
Before running the CLI, install the required dependencies and verify the entry point.
-
Clone the repository and install dependencies:
git clone https://github.com/vitali87/code-graph-rag.git cd code-graph-rag pip install -r requirements.txt -
Verify the CLI entry point is available. The tool exposes the
cgrcommand, implemented incodebase_rag/workspaces/cli.py, which parses arguments and delegates to the main pipeline incodebase_rag/main.py.
Environment Variable Configuration
The CLI relies on external services for embeddings and vector storage. Required secrets and endpoints are declared in codebase_rag/workspaces/constants.py and loaded at runtime.
Set the following variables in your shell or .env file:
OPENAI_API_KEY– Authentication token for OpenAI embedding models.VECTOR_STORE_URL– Endpoint for your vector database (e.g., Pinecone, Weaviate, or local Qdrant).VECTOR_STORE_API_KEY– Optional API key for authenticated vector stores.
The CGRConfig dataclass in codebase_rag/workspaces/models.py aggregates these environment variables with CLI flags to initialize the pipeline.
Core CLI Commands and Flags
The cgr command accepts several flags to control the ingestion and processing pipeline. The argument parser is implemented in codebase_rag/workspaces/cli.py.
Repository Ingestion
To process a codebase, use the --repo flag to specify the target repository URL:
cgr \
--repo https://github.com/your/project.git \
--embedder openai \
--vector-store qdrant \
--output-dir ./cgr_workspace
This command triggers the following sequence:
- Clones the repository.
- Parses the AST to build a graph structure compatible with Neo4j.
- Generates embeddings using the specified embedder.
- Stores vectors in the configured vector store.
Advanced Configuration Options
Fine-tune the pipeline behavior with these additional flags:
--embedder-model– Select specific embedding models (e.g.,text-embedding-ada-002).--chunk-size– Define the number of lines per code chunk for embedding.--prune-threshold– Set the minimum edge score for graph pruning to maintain compactness.
These parameters are passed to the graph builder and embedding pipeline defined in codebase_rag/main.py.
Programmatic Usage via Python
For integration into existing applications, instantiate the RAGEngine class directly instead of using the CLI.
from codebase_rag.rag_engine import RAGEngine
from codebase_rag.workspaces.models import CGRConfig
cfg = CGRConfig(
repo_url="https://github.com/your/project.git",
embedder="openai",
embedder_model="text-embedding-ada-002",
vector_store="qdrant",
output_dir="./cgr_workspace",
chunk_size=200,
)
engine = RAGEngine(cfg)
answer = engine.query("Explain the caching strategy used in the project.")
print(answer)
This approach bypasses the CLI argument parsing in codebase_rag/workspaces/cli.py and uses the same configuration validation defined in codebase_rag/workspaces/models.py.
Step-by-Step Workflow
Follow this sequence to index a repository and query the resulting knowledge graph:
-
Configure environment variables as described in the environment section to authenticate with OpenAI and your vector store.
-
Create a workspace directory to store intermediate files, embeddings, and logs. The default is
./cgr_workspace, configurable via--output-dir. -
Execute the ingestion command:
cgr \ --repo https://github.com/your/project.git \ --embedder openai \ --vector-store qdrant \ --output-dir ./cgr_workspace \ --chunk-size 300 \ --prune-threshold 0.05 -
Query the graph using the REPL mode or programmatic API. For CLI-based querying:
python -m codebase_rag.main --query "How does the authentication flow work?"This invokes the query handler in
codebase_rag/main.pyagainst the populated vector store.
Summary
- The code-graph-rag CLI (
cgr) provides the primary interface for repository ingestion and graph construction. - Environment variables in
codebase_rag/workspaces/constants.pycontrol external service authentication. - CLI flags defined in
codebase_rag/workspaces/cli.pyoverride defaults and specify repository sources, embedders, and vector stores. - The
CGRConfigdataclass incodebase_rag/workspaces/models.pycentralizes configuration for both CLI and programmatic usage. - The
RAGEngineclass incodebase_rag/rag_engine.pyenables direct Python integration for embedded applications.
Frequently Asked Questions
What environment variables are required to run the code-graph-rag CLI?
The CLI requires OPENAI_API_KEY for embedding generation and VECTOR_STORE_URL for connecting to your vector database. Optional variables like VECTOR_STORE_API_KEY are defined in codebase_rag/workspaces/constants.py depending on your provider's authentication requirements.
How do I specify a different embedding model when using the CLI?
Use the --embedder-model flag to override the default model. For example, append --embedder-model text-embedding-ada-002 to your cgr command. This value is processed by the argument parser in codebase_rag/workspaces/cli.py and passed to the embedder initialization logic.
Can I use code-graph-rag as a library instead of a CLI tool?
Yes. Import RAGEngine from codebase_rag/rag_engine and instantiate it with a CGRConfig object from codebase_rag/workspaces/models. This programmatic approach bypasses the CLI parsing in codebase_rag/workspaces/cli.py while maintaining full access to the graph building and RAG pipeline defined in codebase_rag/main.py.
Where is the CLI entry point defined in the source code?
The command-line interface is implemented in codebase_rag/workspaces/cli.py, which defines the cgr command using argparse. This module parses flags like --repo and --embedder, then constructs a CGRConfig instance to launch the pipeline orchestrated in codebase_rag/main.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →