How Configuration is Managed in Code-Graph-RAG: A Complete Guide
Code-Graph-RAG centralizes all runtime configuration in a single Pydantic-based AppConfig class inside codebase_rag/config.py, using environment variables, dotenv files, and runtime singleton access to manage Memgraph connections, LLM providers, and security policies.
The vitali87/code-graph-rag repository implements a robust, environment-driven configuration system that separates secrets from code while providing type-safe validation through Pydantic. Understanding how this system works is essential for customizing vector stores, switching between LLM providers like OpenAI and Anthropic, and securing shell command execution.
Centralized Configuration Module
All configuration logic lives in codebase_rag/config.py, which exports a single settings object consumed throughout the application. This module handles everything from API key validation to ignore-pattern parsing, ensuring consistent behavior across the RAG pipeline.
Environment Variable Loading
At import time, the module automatically loads environment variables from a .env file in the current working directory using python-dotenv:
from dotenv import load_dotenv
from pathlib import Path
load_dotenv(dotenv_path=Path.cwd() / ".env")
This call executes immediately when codebase_rag/config.py is imported, making .env variables available before Pydantic settings resolution begins. Variables defined here take precedence over defaults but yield to explicit keyword arguments passed during AppConfig initialization.
The AppConfig Settings Model
The core of the system is AppConfig, a Pydantic BaseSettings subclass that declares every configurable option as type-annotated class attributes. This includes Memgraph connection parameters (MEMGRAPH_HOST, MEMGRAPH_PORT), vector store backend selection (CGR_VECTOR_STORE_BACKEND), and LLM provider credentials.
BaseSettings resolves values in this strict priority order:
- Explicit kwargs passed to the constructor
- Environment variables (including those loaded from
.env) - Default values defined in the class schema
This hierarchy allows local development overrides via .env files while preserving production flexibility through actual environment variables.
Configuration Architecture
The system employs several patterns to ensure thread-safe, predictable access to configuration values across the async RAG runtime.
Singleton Pattern for Global Access
At the bottom of codebase_rag/config.py, the module instantiates and exports a single settings object:
settings = AppConfig()
All downstream code imports this singleton rather than constructing new instances:
from codebase_rag.config import settings
# Access Memgraph host
host = settings.MEMGRAPH_HOST
This pattern prevents configuration drift and ensures that runtime changes made via setter methods propagate immediately to all consumers.
Model-Specific Configuration
LLM and embedding providers are described by the ModelConfig dataclass, which stores provider names, model IDs, API keys, endpoints, and region settings. Rather than manipulating raw strings, AppConfig exposes helper methods that construct ModelConfig instances:
_get_default_config()– builds base configuration from provider-specific environment variables_get_default_orchestrator_config()– constructs config for the query orchestration layer usingORCHESTRATOR_PROVIDERandORCHESTRATOR_MODEL_get_default_cypher_config()– creates config for Cypher generation usingCYPHER_PROVIDERandCYPHER_MODEL
These methods map environment variable prefixes to structured ModelConfig objects that the provider factory layer consumes.
Dynamic Runtime Activation
AppConfig tracks currently active models via private attributes _active_orchestrator and _active_cypher. The public properties active_orchestrator_config and active_cypher_config return either the user-override set via set_orchestrator() or set_cypher(), or fall back to the default configurations generated by the helper methods above.
This allows runtime model switching without restarting the application:
from codebase_rag.config import settings
# Switch to Anthropic Claude for query orchestration
settings.set_orchestrator(
provider="anthropic",
model="claude-2.0",
api_key="sk-ant-xxxxxxxxxxxx",
)
File-Based Configuration
Beyond environment variables, Code-Graph-RAG supports repository-local configuration through special dotfiles that influence indexing behavior.
Ignore Patterns with .cgrignore
The system respects a .cgrignore file (similar to .gitignore) that defines file and directory patterns excluded from code indexing. The load_cgrignore_patterns() function and load_ignore_patterns() property parse this file using gitignore-style syntax:
# Example .cgrignore
build/
dist/
*.egg-info
!dist/keep.py # Exception to re-include specific files
These patterns prevent build artifacts and dependencies from polluting the graph database during repository ingestion.
Custom Instructions with .cgr.md
Optional markdown instructions loaded from .cgr.md files provide context-specific guidance to the LLM. The load_cgr_instructions() function concatenates global and per-repository instruction files, making the combined text available to orchestration prompts for better query understanding.
Security and Sandboxing
Configuration management in Code-Graph-RAG includes critical safety mechanisms for shell command execution and API credential validation.
API Key Validation
The ModelConfig.validate_api_key() method checks that required credentials exist either in the environment (using provider-specific variable names like OPENAI_API_KEY) or are supplied directly to the config. Missing keys trigger formatted error messages generated by format_missing_api_key_errors(), preventing cryptic authentication failures during graph queries.
Command Allowlists
The AppConfig class enumerates shell command restrictions via attributes like SHELL_COMMAND_ALLOWLIST and SHELL_READ_ONLY_COMMANDS. These lists gate which system commands the RAG engine may execute during repository indexing or tool use, preventing arbitrary code execution in sandboxed environments.
Practical Configuration Examples
Basic Environment Setup
Create a .env file in your repository root:
# LLM provider credentials
OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxx
ANTHROPIC_API_KEY=sk-ant-xxxxxxxxxxxx
# Graph database connection
MEMGRAPH_HOST=localhost
MEMGRAPH_PORT=7687
# Vector store backend selection
CGR_VECTOR_STORE_BACKEND=qdrant
Runtime Configuration Access
Access configuration values in your application code:
from codebase_rag.config import settings
# Read simple connection parameters
host = settings.MEMGRAPH_HOST
port = settings.MEMGRAPH_PORT
# Get structured model configuration
orchestrator_cfg = settings.active_orchestrator_config
print(f"Using {orchestrator_cfg.provider} model {orchestrator_cfg.model_id}")
Switching Providers Dynamically
Override default models without modifying environment variables:
from codebase_rag.config import settings
# Activate OpenAI GPT-4 for Cypher generation
settings.set_cypher(
provider="openai",
model="gpt-4",
api_key="sk-xxxxxxxxxxxxxxxx",
)
Summary
- Centralized module: All configuration lives in
codebase_rag/config.pyand exports a singletonsettingsobject - Environment-driven: Values resolve from kwargs → environment variables (including
.env) → class defaults using PydanticBaseSettings - Structured providers: The
ModelConfigdataclass encapsulates LLM/embedding settings with validation logic - Runtime flexibility:
set_orchestrator()andset_cypher()methods enable dynamic model switching without restarts - File-based rules:
.cgrignoreand.cgr.mdfiles provide repository-specific exclusion patterns and LLM instructions - Security boundaries: Command allowlists and API key validation prevent unsafe operations in production environments
Frequently Asked Questions
How do I override configuration without modifying the source code?
Create a .env file in your working directory. The load_dotenv() call in codebase_rag/config.py automatically loads these variables at import time, and Pydantic's BaseSettings uses them to populate the AppConfig instance. For temporary overrides, pass values directly to settings.set_orchestrator() or settings.set_cypher() at runtime.
Can I use different LLM providers for query orchestration and Cypher generation?
Yes. The AppConfig class maintains separate active configurations for the orchestrator (natural language understanding) and Cypher generator (graph query construction). Use settings.set_orchestrator(provider="anthropic", ...) and settings.set_cypher(provider="openai", ...) independently to mix providers optimized for different tasks.
What happens if a required API key is missing?
The ModelConfig.validate_api_key() method raises a descriptive error before any network request occurs. The error message, formatted by format_missing_api_key_errors(), specifies exactly which environment variable (e.g., ANTHROPIC_API_KEY) must be set for the selected provider, preventing silent authentication failures.
How do I prevent specific files from being indexed in the knowledge graph?
Add patterns to a .cgrignore file in your repository root. The load_cgrignore_patterns() function parses these rules using standard gitignore syntax, excluding matches from the graph construction process while still allowing repository-specific exceptions with ! negation patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →