How Configuration is Managed in Code-Graph-RAG: A Complete Guide

Code-Graph-RAG centralizes all runtime configuration in a single Pydantic-based AppConfig class inside codebase_rag/config.py, using environment variables, dotenv files, and runtime singleton access to manage Memgraph connections, LLM providers, and security policies.

The vitali87/code-graph-rag repository implements a robust, environment-driven configuration system that separates secrets from code while providing type-safe validation through Pydantic. Understanding how this system works is essential for customizing vector stores, switching between LLM providers like OpenAI and Anthropic, and securing shell command execution.

Centralized Configuration Module

All configuration logic lives in codebase_rag/config.py, which exports a single settings object consumed throughout the application. This module handles everything from API key validation to ignore-pattern parsing, ensuring consistent behavior across the RAG pipeline.

Environment Variable Loading

At import time, the module automatically loads environment variables from a .env file in the current working directory using python-dotenv:

from dotenv import load_dotenv
from pathlib import Path

load_dotenv(dotenv_path=Path.cwd() / ".env")

This call executes immediately when codebase_rag/config.py is imported, making .env variables available before Pydantic settings resolution begins. Variables defined here take precedence over defaults but yield to explicit keyword arguments passed during AppConfig initialization.

The AppConfig Settings Model

The core of the system is AppConfig, a Pydantic BaseSettings subclass that declares every configurable option as type-annotated class attributes. This includes Memgraph connection parameters (MEMGRAPH_HOST, MEMGRAPH_PORT), vector store backend selection (CGR_VECTOR_STORE_BACKEND), and LLM provider credentials.

BaseSettings resolves values in this strict priority order:

  1. Explicit kwargs passed to the constructor
  2. Environment variables (including those loaded from .env)
  3. Default values defined in the class schema

This hierarchy allows local development overrides via .env files while preserving production flexibility through actual environment variables.

Configuration Architecture

The system employs several patterns to ensure thread-safe, predictable access to configuration values across the async RAG runtime.

Singleton Pattern for Global Access

At the bottom of codebase_rag/config.py, the module instantiates and exports a single settings object:

settings = AppConfig()

All downstream code imports this singleton rather than constructing new instances:

from codebase_rag.config import settings

# Access Memgraph host

host = settings.MEMGRAPH_HOST

This pattern prevents configuration drift and ensures that runtime changes made via setter methods propagate immediately to all consumers.

Model-Specific Configuration

LLM and embedding providers are described by the ModelConfig dataclass, which stores provider names, model IDs, API keys, endpoints, and region settings. Rather than manipulating raw strings, AppConfig exposes helper methods that construct ModelConfig instances:

  • _get_default_config() – builds base configuration from provider-specific environment variables
  • _get_default_orchestrator_config() – constructs config for the query orchestration layer using ORCHESTRATOR_PROVIDER and ORCHESTRATOR_MODEL
  • _get_default_cypher_config() – creates config for Cypher generation using CYPHER_PROVIDER and CYPHER_MODEL

These methods map environment variable prefixes to structured ModelConfig objects that the provider factory layer consumes.

Dynamic Runtime Activation

AppConfig tracks currently active models via private attributes _active_orchestrator and _active_cypher. The public properties active_orchestrator_config and active_cypher_config return either the user-override set via set_orchestrator() or set_cypher(), or fall back to the default configurations generated by the helper methods above.

This allows runtime model switching without restarting the application:

from codebase_rag.config import settings

# Switch to Anthropic Claude for query orchestration

settings.set_orchestrator(
    provider="anthropic",
    model="claude-2.0",
    api_key="sk-ant-xxxxxxxxxxxx",
)

File-Based Configuration

Beyond environment variables, Code-Graph-RAG supports repository-local configuration through special dotfiles that influence indexing behavior.

Ignore Patterns with .cgrignore

The system respects a .cgrignore file (similar to .gitignore) that defines file and directory patterns excluded from code indexing. The load_cgrignore_patterns() function and load_ignore_patterns() property parse this file using gitignore-style syntax:


# Example .cgrignore

build/
dist/
*.egg-info
!dist/keep.py  # Exception to re-include specific files

These patterns prevent build artifacts and dependencies from polluting the graph database during repository ingestion.

Custom Instructions with .cgr.md

Optional markdown instructions loaded from .cgr.md files provide context-specific guidance to the LLM. The load_cgr_instructions() function concatenates global and per-repository instruction files, making the combined text available to orchestration prompts for better query understanding.

Security and Sandboxing

Configuration management in Code-Graph-RAG includes critical safety mechanisms for shell command execution and API credential validation.

API Key Validation

The ModelConfig.validate_api_key() method checks that required credentials exist either in the environment (using provider-specific variable names like OPENAI_API_KEY) or are supplied directly to the config. Missing keys trigger formatted error messages generated by format_missing_api_key_errors(), preventing cryptic authentication failures during graph queries.

Command Allowlists

The AppConfig class enumerates shell command restrictions via attributes like SHELL_COMMAND_ALLOWLIST and SHELL_READ_ONLY_COMMANDS. These lists gate which system commands the RAG engine may execute during repository indexing or tool use, preventing arbitrary code execution in sandboxed environments.

Practical Configuration Examples

Basic Environment Setup

Create a .env file in your repository root:


# LLM provider credentials

OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxx
ANTHROPIC_API_KEY=sk-ant-xxxxxxxxxxxx

# Graph database connection

MEMGRAPH_HOST=localhost
MEMGRAPH_PORT=7687

# Vector store backend selection

CGR_VECTOR_STORE_BACKEND=qdrant

Runtime Configuration Access

Access configuration values in your application code:

from codebase_rag.config import settings

# Read simple connection parameters

host = settings.MEMGRAPH_HOST
port = settings.MEMGRAPH_PORT

# Get structured model configuration

orchestrator_cfg = settings.active_orchestrator_config
print(f"Using {orchestrator_cfg.provider} model {orchestrator_cfg.model_id}")

Switching Providers Dynamically

Override default models without modifying environment variables:

from codebase_rag.config import settings

# Activate OpenAI GPT-4 for Cypher generation

settings.set_cypher(
    provider="openai",
    model="gpt-4",
    api_key="sk-xxxxxxxxxxxxxxxx",
)

Summary

  • Centralized module: All configuration lives in codebase_rag/config.py and exports a singleton settings object
  • Environment-driven: Values resolve from kwargs → environment variables (including .env) → class defaults using Pydantic BaseSettings
  • Structured providers: The ModelConfig dataclass encapsulates LLM/embedding settings with validation logic
  • Runtime flexibility: set_orchestrator() and set_cypher() methods enable dynamic model switching without restarts
  • File-based rules: .cgrignore and .cgr.md files provide repository-specific exclusion patterns and LLM instructions
  • Security boundaries: Command allowlists and API key validation prevent unsafe operations in production environments

Frequently Asked Questions

How do I override configuration without modifying the source code?

Create a .env file in your working directory. The load_dotenv() call in codebase_rag/config.py automatically loads these variables at import time, and Pydantic's BaseSettings uses them to populate the AppConfig instance. For temporary overrides, pass values directly to settings.set_orchestrator() or settings.set_cypher() at runtime.

Can I use different LLM providers for query orchestration and Cypher generation?

Yes. The AppConfig class maintains separate active configurations for the orchestrator (natural language understanding) and Cypher generator (graph query construction). Use settings.set_orchestrator(provider="anthropic", ...) and settings.set_cypher(provider="openai", ...) independently to mix providers optimized for different tasks.

What happens if a required API key is missing?

The ModelConfig.validate_api_key() method raises a descriptive error before any network request occurs. The error message, formatted by format_missing_api_key_errors(), specifies exactly which environment variable (e.g., ANTHROPIC_API_KEY) must be set for the selected provider, preventing silent authentication failures.

How do I prevent specific files from being indexed in the knowledge graph?

Add patterns to a .cgrignore file in your repository root. The load_cgrignore_patterns() function parses these rules using standard gitignore syntax, excluding matches from the graph construction process while still allowing repository-specific exceptions with ! negation patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →