# How Configuration is Managed in Code-Graph-RAG: A Complete Guide

> Discover how Code-Graph-RAG manages configuration centrally using Pydantic AppConfig, environment variables, and dotenv files for Memgraph, LLMs, and security.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Code-Graph-RAG centralizes all runtime configuration in a single Pydantic-based `AppConfig` class inside [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py), using environment variables, dotenv files, and runtime singleton access to manage Memgraph connections, LLM providers, and security policies.**

The `vitali87/code-graph-rag` repository implements a robust, environment-driven configuration system that separates secrets from code while providing type-safe validation through Pydantic. Understanding how this system works is essential for customizing vector stores, switching between LLM providers like OpenAI and Anthropic, and securing shell command execution.

## Centralized Configuration Module

All configuration logic lives in **[`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py)**, which exports a single `settings` object consumed throughout the application. This module handles everything from API key validation to ignore-pattern parsing, ensuring consistent behavior across the RAG pipeline.

### Environment Variable Loading

At import time, the module automatically loads environment variables from a `.env` file in the current working directory using `python-dotenv`:

```python
from dotenv import load_dotenv
from pathlib import Path

load_dotenv(dotenv_path=Path.cwd() / ".env")

```

This call executes immediately when [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) is imported, making `.env` variables available before Pydantic settings resolution begins. Variables defined here take precedence over defaults but yield to explicit keyword arguments passed during `AppConfig` initialization.

### The AppConfig Settings Model

The core of the system is **`AppConfig`**, a Pydantic `BaseSettings` subclass that declares every configurable option as type-annotated class attributes. This includes Memgraph connection parameters (`MEMGRAPH_HOST`, `MEMGRAPH_PORT`), vector store backend selection (`CGR_VECTOR_STORE_BACKEND`), and LLM provider credentials.

`BaseSettings` resolves values in this strict priority order:

1. Explicit kwargs passed to the constructor
2. Environment variables (including those loaded from `.env`)
3. Default values defined in the class schema

This hierarchy allows local development overrides via `.env` files while preserving production flexibility through actual environment variables.

## Configuration Architecture

The system employs several patterns to ensure thread-safe, predictable access to configuration values across the async RAG runtime.

### Singleton Pattern for Global Access

At the bottom of [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py), the module instantiates and exports a single `settings` object:

```python
settings = AppConfig()

```

All downstream code imports this singleton rather than constructing new instances:

```python
from codebase_rag.config import settings

# Access Memgraph host

host = settings.MEMGRAPH_HOST

```

This pattern prevents configuration drift and ensures that runtime changes made via setter methods propagate immediately to all consumers.

### Model-Specific Configuration

LLM and embedding providers are described by the **`ModelConfig`** dataclass, which stores provider names, model IDs, API keys, endpoints, and region settings. Rather than manipulating raw strings, `AppConfig` exposes helper methods that construct `ModelConfig` instances:

- `_get_default_config()` – builds base configuration from provider-specific environment variables
- `_get_default_orchestrator_config()` – constructs config for the query orchestration layer using `ORCHESTRATOR_PROVIDER` and `ORCHESTRATOR_MODEL`
- `_get_default_cypher_config()` – creates config for Cypher generation using `CYPHER_PROVIDER` and `CYPHER_MODEL`

These methods map environment variable prefixes to structured `ModelConfig` objects that the provider factory layer consumes.

### Dynamic Runtime Activation

`AppConfig` tracks currently active models via private attributes **`_active_orchestrator`** and **`_active_cypher`**. The public properties `active_orchestrator_config` and `active_cypher_config` return either the user-override set via `set_orchestrator()` or `set_cypher()`, or fall back to the default configurations generated by the helper methods above.

This allows runtime model switching without restarting the application:

```python
from codebase_rag.config import settings

# Switch to Anthropic Claude for query orchestration

settings.set_orchestrator(
    provider="anthropic",
    model="claude-2.0",
    api_key="sk-ant-xxxxxxxxxxxx",
)

```

## File-Based Configuration

Beyond environment variables, Code-Graph-RAG supports repository-local configuration through special dotfiles that influence indexing behavior.

### Ignore Patterns with .cgrignore

The system respects a **`.cgrignore`** file (similar to `.gitignore`) that defines file and directory patterns excluded from code indexing. The `load_cgrignore_patterns()` function and `load_ignore_patterns()` property parse this file using gitignore-style syntax:

```text

# Example .cgrignore

build/
dist/
*.egg-info
!dist/keep.py  # Exception to re-include specific files

```

These patterns prevent build artifacts and dependencies from polluting the graph database during repository ingestion.

### Custom Instructions with .cgr.md

Optional markdown instructions loaded from [`.cgr.md`](https://github.com/vitali87/code-graph-rag/blob/main/.cgr.md) files provide context-specific guidance to the LLM. The `load_cgr_instructions()` function concatenates global and per-repository instruction files, making the combined text available to orchestration prompts for better query understanding.

## Security and Sandboxing

Configuration management in Code-Graph-RAG includes critical safety mechanisms for shell command execution and API credential validation.

### API Key Validation

The `ModelConfig.validate_api_key()` method checks that required credentials exist either in the environment (using provider-specific variable names like `OPENAI_API_KEY`) or are supplied directly to the config. Missing keys trigger formatted error messages generated by `format_missing_api_key_errors()`, preventing cryptic authentication failures during graph queries.

### Command Allowlists

The `AppConfig` class enumerates shell command restrictions via attributes like **`SHELL_COMMAND_ALLOWLIST`** and **`SHELL_READ_ONLY_COMMANDS`**. These lists gate which system commands the RAG engine may execute during repository indexing or tool use, preventing arbitrary code execution in sandboxed environments.

## Practical Configuration Examples

### Basic Environment Setup

Create a `.env` file in your repository root:

```dotenv

# LLM provider credentials

OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxx
ANTHROPIC_API_KEY=sk-ant-xxxxxxxxxxxx

# Graph database connection

MEMGRAPH_HOST=localhost
MEMGRAPH_PORT=7687

# Vector store backend selection

CGR_VECTOR_STORE_BACKEND=qdrant

```

### Runtime Configuration Access

Access configuration values in your application code:

```python
from codebase_rag.config import settings

# Read simple connection parameters

host = settings.MEMGRAPH_HOST
port = settings.MEMGRAPH_PORT

# Get structured model configuration

orchestrator_cfg = settings.active_orchestrator_config
print(f"Using {orchestrator_cfg.provider} model {orchestrator_cfg.model_id}")

```

### Switching Providers Dynamically

Override default models without modifying environment variables:

```python
from codebase_rag.config import settings

# Activate OpenAI GPT-4 for Cypher generation

settings.set_cypher(
    provider="openai",
    model="gpt-4",
    api_key="sk-xxxxxxxxxxxxxxxx",
)

```

## Summary

- **Centralized module**: All configuration lives in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) and exports a singleton `settings` object
- **Environment-driven**: Values resolve from kwargs → environment variables (including `.env`) → class defaults using Pydantic `BaseSettings`
- **Structured providers**: The `ModelConfig` dataclass encapsulates LLM/embedding settings with validation logic
- **Runtime flexibility**: `set_orchestrator()` and `set_cypher()` methods enable dynamic model switching without restarts
- **File-based rules**: `.cgrignore` and [`.cgr.md`](https://github.com/vitali87/code-graph-rag/blob/main/.cgr.md) files provide repository-specific exclusion patterns and LLM instructions
- **Security boundaries**: Command allowlists and API key validation prevent unsafe operations in production environments

## Frequently Asked Questions

### How do I override configuration without modifying the source code?

Create a `.env` file in your working directory. The `load_dotenv()` call in [`codebase_rag/config.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/config.py) automatically loads these variables at import time, and Pydantic's `BaseSettings` uses them to populate the `AppConfig` instance. For temporary overrides, pass values directly to `settings.set_orchestrator()` or `settings.set_cypher()` at runtime.

### Can I use different LLM providers for query orchestration and Cypher generation?

Yes. The `AppConfig` class maintains separate active configurations for the orchestrator (natural language understanding) and Cypher generator (graph query construction). Use `settings.set_orchestrator(provider="anthropic", ...)` and `settings.set_cypher(provider="openai", ...)` independently to mix providers optimized for different tasks.

### What happens if a required API key is missing?

The `ModelConfig.validate_api_key()` method raises a descriptive error before any network request occurs. The error message, formatted by `format_missing_api_key_errors()`, specifies exactly which environment variable (e.g., `ANTHROPIC_API_KEY`) must be set for the selected provider, preventing silent authentication failures.

### How do I prevent specific files from being indexed in the knowledge graph?

Add patterns to a `.cgrignore` file in your repository root. The `load_cgrignore_patterns()` function parses these rules using standard gitignore syntax, excluding matches from the graph construction process while still allowing repository-specific exceptions with `!` negation patterns.