How to Configure code-graph-rag with OpenAI: Environment Variables and Programmatic Setup

Configure code-graph-rag with OpenAI by setting ORCHESTRATOR_PROVIDER=openai in your .env file or calling settings.set_orchestrator("openai", model="gpt-4o", api_key="sk-...") in Python, which instantiates the OpenAIProvider class from codebase_rag/providers/base.py to route LLM calls to OpenAI's API.

code-graph-rag is an open-source RAG framework that supports multiple LLM providers through a unified provider architecture. To configure code-graph-rag with OpenAI for graph-based code analysis, you can use environment variables for static configuration or the programmatic Python API for dynamic provider switching. The OpenAIProvider class implements the common interface required by the orchestrator and cypher components, handling authentication and model instantiation automatically via the provider registry.

Provider Architecture and OpenAI Integration

The framework uses a pluggable provider pattern where each LLM backend implements a standardized interface. The OpenAIProvider class resides in codebase_rag/providers/base.py and is registered in the global PROVIDER_REGISTRY under the key "openai" [source].

When the orchestrator or cypher components request a model, the registry resolves the provider name and delegates instantiation to the provider's create_model() method. This method validates the API key (from environment or explicit arguments) and constructs an OpenAIResponsesModel or OpenAIChatModel that communicates with OpenAI's HTTP endpoints [source].

Method 1: Configure via Environment Variables

The simplest way to configure code-graph-rag with OpenAI is through a .env file. The framework loads these variables at startup via the settings module, automatically instantiating the correct provider based on the ORCHESTRATOR_PROVIDER and CYPHER_PROVIDER values.

Required environment variables include:

  • ORCHESTRATOR_PROVIDER=openai — Sets the LLM provider for the main orchestrator
  • ORCHESTRATOR_MODEL — Specifies the model ID (e.g., gpt-4o, gpt-5.6-terra)
  • ORCHESTRATOR_API_KEY — Your OpenAI API key
  • ORCHESTRATOR_ENDPOINT — Optional custom endpoint (defaults to https://api.openai.com/v1)

For Cypher query generation (graph database queries), use the CYPHER_* prefix:


# .env

ORCHESTRATOR_PROVIDER=openai
ORCHESTRATOR_MODEL=gpt-5.6-terra
ORCHESTRATOR_API_KEY=sk-your-openai-key

# Optional: Azure OpenAI or custom endpoint

# ORCHESTRATOR_ENDPOINT=https://myresource.openai.azure.com/v1

CYPHER_PROVIDER=openai
CYPHER_MODEL=gpt-5.6-luna
CYPHER_API_KEY=sk-your-openai-key

See the .env.example file in the repository root for a complete template of supported variables [source].

Method 2: Programmatic Configuration

For dynamic configuration or multi-tenant deployments, use the settings API to configure code-graph-rag with OpenAI at runtime. The set_orchestrator() and set_cypher() methods in codebase_rag/config.py update the internal Config singleton and reinitialize the provider without restarting the process.

from cgr import settings

# Configure the orchestrator to use OpenAI

settings.set_orchestrator(
    provider="openai",
    model="gpt-5.6-terra",
    api_key="sk-XXXXXXXXXXXXXXXX"
)

# Configure Cypher query generation with different OpenAI model

settings.set_cypher(
    provider="openai",
    model="gpt-5.6-luna",
    api_key="sk-XXXXXXXXXXXXXXXX"
)

These methods validate the provider string against the registry, instantiate the OpenAIProvider, and update the global configuration object so subsequent LLM calls route to OpenAI's API [source].

Advanced: Direct Model Instantiation

For custom workflows that bypass the global settings, instantiate the OpenAI provider directly using the configuration utilities. The parse_model_string() function extracts the provider name and model ID from a URI-style string like "openai:gpt-4o".

from cgr import config

# Parse provider and model from string

provider_name, model_id = config.parse_model_string("openai:gpt-4o")

# Get provider instance with explicit API key

provider = config.get_provider(provider_name, api_key="sk-...")

# Create the concrete model

model = provider.create_model(model_id)

# Execute completion

response = model.completions(prompt="Explain GraphQL in one sentence.")
print(response.text)

This approach is useful when you need multiple concurrent OpenAI configurations or when integrating code-graph-rag into existing applications with their own configuration management [source].

Verifying Your Configuration

After configuration, verify that the orchestrator is using the OpenAI provider by inspecting the runtime settings:

from cgr import settings

# Verify provider is correctly set

assert settings.orchestrator.provider_name == "openai"
print(f"Active model: {settings.orchestrator.model_id}")
print(f"Provider class: {type(settings.orchestrator).__module__}")

This assertion confirms that the provider registry correctly resolved the "openai" string to the OpenAIProvider class and that API calls will route to OpenAI's endpoints rather than other providers like Anthropic or Google.

Summary

  • Provider Registry: The OpenAIProvider class in codebase_rag/providers/base.py registers under the key "openai" and handles all OpenAI-specific authentication and model creation.
  • Environment Configuration: Set ORCHESTRATOR_PROVIDER=openai, ORCHESTRATOR_MODEL, and ORCHESTRATOR_API_KEY in your .env file for static configuration.
  • Programmatic Control: Use settings.set_orchestrator("openai", model="...", api_key="...") from codebase_rag/config.py to switch providers at runtime without restarting.
  • Model String Format: Use the "openai:model-name" format with config.parse_model_string() for direct instantiation when you need fine-grained control over the provider lifecycle.
  • Endpoint Flexibility: Override the default OpenAI endpoint using the ORCHESTRATOR_ENDPOINT variable or endpoint parameter for Azure OpenAI or proxy configurations.

Frequently Asked Questions

Can I use Azure OpenAI instead of the standard OpenAI API?

Yes. When you configure code-graph-rag with OpenAI, set the ORCHESTRATOR_ENDPOINT environment variable to your Azure OpenAI endpoint URL (e.g., https://myresource.openai.azure.com/v1), or pass the endpoint parameter to settings.set_orchestrator(). The OpenAIProvider class uses this endpoint when constructing the underlying HTTP client, allowing you to target Azure deployments while using the same provider interface.

How do I switch between OpenAI and another provider without restarting my application?

Call settings.set_orchestrator() with the new provider name and credentials. This method updates the global configuration singleton and reinitializes the orchestrator with the specified provider. For example, switching from Anthropic to OpenAI requires only a single function call: settings.set_orchestrator(provider="openai", model="gpt-4o", api_key="sk-...").

Where should I store my OpenAI API key when using environment variables?

Store sensitive credentials in a .env file at your project root, never in source code. The framework loads these via python-dotenv at startup. The .env.example file documents the required variables including ORCHESTRATOR_API_KEY and CYPHER_API_KEY, which OpenAIProvider reads during initialization to authenticate with OpenAI's API.

What happens if I don't specify an API key when configuring the OpenAI provider?

If no API key is provided via the api_key parameter or ORCHESTRATOR_API_KEY environment variable, the OpenAIProvider constructor will raise a configuration error during model instantiation. The provider validates that either an explicit key or the standard OPENAI_API_KEY environment variable is present before attempting to create the model instance, ensuring clear error messages rather than runtime authentication failures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →