How to Configure API Keys and Model Settings for code-graph-rag AI Components

Configure API keys via environment variables or .env file, then use settings.set_orchestrator() to select providers and models at runtime.

All AI configuration in code-graph-rag lives in codebase_rag/config.py, with provider-specific logic handled by codebase_rag/providers/base.py. This guide shows you how to supply credentials, override defaults, and instantiate live LLM clients.


Setting Up API Keys

The repository expects credentials as environment variables. It automatically loads a .env file at startup via load_dotenv (config.py:24-25).

Supported Providers and Required Environment Variables

Provider Environment Variable Notes
OpenAI OPENAI_API_KEY Required for GPT models
Anthropic ANTHROPIC_API_KEY Required for Claude models
Google GOOGLE_API_KEY Required for Gemini/Vertex models
Azure AZURE_API_KEY Required for Azure OpenAI
MiniMax MINIMAX_API_KEY Required for MiniMax models
Ollama (none) Local deployment needs no key

This mapping is defined in API_KEY_INFO at config.py:33-59.

Sample .env File


# .env in project root

OPENAI_API_KEY=sk-your-openai-key-here
ANTHROPIC_API_KEY=sk-ant-your-anthropic-key
GOOGLE_API_KEY=AIza-your-google-key
AZURE_API_KEY=your-azure-key
MINIMAX_API_KEY=your-minimax-key

If a required key is missing, ModelConfig.validate_api_key triggers format_missing_api_key_errors (config.py:62-100) to raise a detailed error message.


Model Configuration Structure

The ModelConfig Dataclass

The core data structure for model settings is ModelConfig (config.py:107-118):

Field Purpose
provider Provider name: openai, anthropic, google, azure, minimax, ollama
model_id Provider-specific identifier (e.g., gpt-4o, claude-3-sonnet-20240229)
api_key Optional override; falls back to environment variable
endpoint Custom API base URL (for Azure, proxies, or self-hosted solutions)
project_id, region, provider_type Google-specific configuration
thinking_budget, service_account_file Advanced options for specific providers

Default Fallback Behavior

Without explicit configuration, AppConfig._get_default_config (config.py:33-57) creates an Ollama fallback pointing to llama3.2 on localhost.


Runtime Configuration Methods

Changing the Orchestrator LLM

The orchestrator handles query answering. Override it with settings.set_orchestrator() (config.py:75-80):

from codebase_rag.config import settings

settings.set_orchestrator(
    provider="openai",
    model="gpt-4o",
    # api_key="sk-...",  # optional: uses OPENAI_API_KEY env var if omitted

    # endpoint="https://api.openai.com/v1",  # optional: custom endpoint

)

Changing the Cypher Generation LLM

For Neo4j query generation, use settings.set_cypher() with identical parameters:

settings.set_cypher(
    provider="anthropic",
    model="claude-3-sonnet-20240229",
)

Embedding Model Configuration

Set the embedding provider via settings.EMBEDDING_PROVIDER or environment variables. The same ModelConfig pattern applies.


Instantiating Live Model Clients

Configuration objects become usable models through the provider factory. The function get_provider_from_config (providers/base.py:29-39) performs this conversion:

  1. Looks up the provider class in PROVIDER_REGISTRY (providers/base.py:95-102)
  2. Calls _resolve_api_key (providers/base.py:48-52) to fetch or validate credentials
  3. Invokes create_model(model_id) to return a pydantic_ai model instance

Each provider implements validate_config (e.g., OpenAIProvider.validate_config at providers/base.py:42-45) to verify required keys before instantiation.


Complete Configuration Example

from codebase_rag.config import settings
from codebase_rag.providers.base import get_provider_from_config

# 1. Configuration loads automatically on import from .env or environment

# 2. Set custom orchestrator (Claude via Anthropic)

settings.set_orchestrator(
    provider="anthropic",
    model="claude-3-sonnet-20240229",
)

# 3. Retrieve active configuration

orchestrator_cfg = settings.active_orchestrator_config

# 4. Build concrete model client

orchestrator = get_provider_from_config(orchestrator_cfg)

# 5. Use for inference

model = orchestrator.create_model(orchestrator_cfg.model_id)
response = model.chat([{"role": "user", "content": "Explain this codebase"}])

Key Implementation Files

File Responsibility
codebase_rag/config.py AppConfig singleton, ModelConfig dataclass, .env loading, set_orchestrator(), set_cypher()
codebase_rag/providers/base.py PROVIDER_REGISTRY, get_provider_from_config(), OpenAIProvider, AnthropicProvider, OllamaProvider, etc.
codebase_rag/constants/providers.py Provider enums, environment variable names, default constants

Summary

  • Store secrets in .env — code-graph-rag loads these automatically via load_dotenv in config.py
  • Import settings singleton — from codebase_rag.config import settings provides access to all configuration
  • Use set_orchestrator() or set_cypher() — runtime methods to switch providers and models without code changes
  • Call get_provider_from_config() — factory function that converts ModelConfig into live pydantic_ai models
  • Provider registry handles validation — each provider checks required keys via validate_config() before instantiation

Frequently Asked Questions

How do I use a local Ollama model without API keys?

Ollama requires no API key. Ensure Ollama is running locally, then call settings.set_orchestrator(provider="ollama", model="llama3.2"). The internal placeholder key in API_KEY_INFO handles authentication transparently.

Can I override the API key at runtime instead of using environment variables?

Yes. Pass api_key="your-key" directly to set_orchestrator() or set_cypher(). The _resolve_api_key helper in providers/base.py:48-52 prefers explicit keys over environment variables.

What happens if I configure a provider but the API key is missing?

ModelConfig.validate_api_key() raises an error via format_missing_api_key_errors (config.py:62-100), listing which environment variable is required for your chosen provider.

How do I connect to Azure OpenAI or a custom LiteLLM proxy?

Use the endpoint parameter in set_orchestrator(). For Azure, also set AZURE_API_KEY and specify the full Azure deployment URL as the endpoint.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →