How to Configure an LLM Provider in Cognee: Environment Setup Guide
Cognee centralizes LLM configuration through environment variables read by the LLMConfig class, allowing you to switch between providers like OpenAI, Anthropic, or Ollama by setting LLM_PROVIDER, LLM_MODEL, and LLM_API_KEY without changing application code.
Cognee (topoteretes/cognee) unifies LLM configuration in a single settings layer that propagates to every graph generation, entity extraction, and RAG component. Understanding how to configure an LLM provider in Cognee ensures your pipelines use the correct model for embedding and inference tasks across the entire framework.
How Cognee's LLM Configuration Works
The Configuration Layer (LLMConfig)
The configuration entry point is cognee.infrastructure.llm.config.LLMConfig defined in cognee/infrastructure/llm/config.py. This class inherits from Pydantic's BaseSettings and reads environment variables at startup to validate and expose LLM settings.
The following variables control provider selection:
- LLM_PROVIDER – The provider name (
openai,anthropic,gemini,ollama,mistral,bedrock, etc.) - LLM_MODEL – The specific model identifier (e.g.,
gpt-4o-mini,claude-3-5-sonnet) - LLM_API_KEY – The authentication secret required by the provider
- LLM_ENDPOINT – Optional URL for self-hosted services (Ollama, local vLLM)
If required variables are missing, LLMConfig raises a clear validation error indicating exactly which environment variable is absent.
The Provider Factory
Downstream components do not read environment variables directly. Instead, they import LLMProvider from cognee/tasks/translation/providers/llm_provider.py. This factory class instantiates a concrete client using the Litellm-Instructor adapter based on the values stored in LLMConfig.
Because all LLM-driven tasks (entity extraction, graph generation, search) consume this centralized provider, switching from OpenAI to a local Ollama instance requires only changing environment variables—no code modifications needed.
Observability Integration
Every LLM call emits OpenTelemetry trace attributes defined in cognee/infrastructure/llm/structured_output_framework/litellm_instructor/llm/generic_llm_api/adapter.py. Specifically, the framework adds:
cognee.llm.provider– The active provider namecognee.llm.model– The specific model identifier
This gives you trace-level visibility into which provider and model handled each request across distributed pipelines.
Setting Up Environment Variables
Using a .env File
The recommended approach is creating a .env file in your project root (copy from the repository's /.env.example). Cognee automatically loads these values via BaseSettings.
# .env configuration
LLM_PROVIDER=openai
LLM_MODEL=gpt-4o-mini
LLM_API_KEY=sk-XXXXXXXXXXXXXXXXXXXXXXXXXXXX
# Required only for self-hosted providers like Ollama
LLM_ENDPOINT=http://localhost:11434/v1
Provider-Specific Configuration Examples
OpenAI
LLM_PROVIDER=openaiLLM_MODEL=gpt-4o-miniLLM_API_KEY=sk-...
Anthropic
LLM_PROVIDER=anthropicLLM_MODEL=claude-3-5-sonnetLLM_API_KEY=sk-ant-...
Ollama (Local)
LLM_PROVIDER=ollamaLLM_MODEL=llama3.1:8bLLM_ENDPOINT=http://localhost:11434/v1LLM_API_KEYcan be set to any non-empty placeholder if required
Programmatic and CLI Configuration
Runtime Overrides in Python
You can override configuration programmatically before initializing components. This is useful for testing or dynamic provider selection.
import os
from cognee.infrastructure.llm.config import LLMConfig
# Set environment variables before instantiating config
os.environ["LLM_PROVIDER"] = "ollama"
os.environ["LLM_MODEL"] = "llama3.1:8b"
os.environ["LLM_ENDPOINT"] = "http://localhost:11434/v1"
# Validate and load configuration
config = LLMConfig()
print(f"Provider: {config.LLM_PROVIDER}, Model: {config.LLM_MODEL}")
One-Off CLI Execution
For single command execution without permanent configuration files, prefix the command with environment variables:
LLM_PROVIDER=anthropic LLM_MODEL=claude-3-5-sonnet LLM_API_KEY=your_key \
cognee-cli add "Your data chunk to process"
Accessing the Configured LLM in Your Code
Once environment variables are set, retrieve the client through the provider factory or use the high-level gateway for structured output:
from cognee.tasks.translation.providers.llm_provider import LLMProvider
from cognee.infrastructure.llm.LLMGateway import LLMGateway
from pydantic import BaseModel
class SummaryOutput(BaseModel):
summary: str
key_points: list[str]
# Factory provides the configured client
provider = LLMProvider()
client = provider.get_client()
# Generate structured output using the configured provider
response = await LLMGateway.acreate_structured_output(
user_prompt="Summarize the following text",
system_prompt="You are a concise technical summarizer.",
output_schema=SummaryOutput,
text="Long document content here..."
)
print(response)
The LLMGateway class in cognee/infrastructure/llm/LLMGateway.py provides async methods like acreate_structured_output and arun that automatically respect the LLMConfig settings.
API and MCP Server Integration
The same environment variable configuration propagates to Cognee's public interfaces:
- REST API endpoints (
/v1/add,/v1/cognify,/v1/search) documented incognee/api/v1/add/add.pyandcognee/api/v1/search/search.pyexposeLLM_PROVIDERsettings in their OpenAPI schemas - MCP Server (
cognee-mcp/src/server.py) forwards these LLM settings to the Model Context Protocol server, ensuring consistent provider usage across tool integrations
Summary
- Centralized config:
LLMConfigincognee/infrastructure/llm/config.pyreadsLLM_PROVIDER,LLM_MODEL,LLM_API_KEY, and optionalLLM_ENDPOINTfrom environment variables - Factory pattern:
LLMProviderincognee/tasks/translation/providers/llm_provider.pycreates the concrete Litellm client used by all tasks - Observability: Traces include
cognee.llm.providerandcognee.llm.modelattributes for debugging - Zero-code switching: Change providers by updating environment variables; no code changes required in pipelines or API routes
Frequently Asked Questions
What environment variables are required to configure an LLM provider in Cognee?
You must set LLM_PROVIDER (e.g., openai, anthropic), LLM_MODEL (e.g., gpt-4o-mini), and LLM_API_KEY. For self-hosted providers like Ollama, you must also specify LLM_ENDPOINT with the local API URL.
Can I switch LLM providers without restarting my Cognee application?
LLMConfig reads environment variables at instantiation time using Pydantic BaseSettings. To switch providers in a running process, you must create a new LLMConfig instance after updating os.environ, or restart the application to pick up new .env values.
Does Cognee support local or self-hosted LLMs?
Yes. Set LLM_PROVIDER=ollama (or another local-compatible provider) and configure LLM_ENDPOINT to point to your local inference server (e.g., http://localhost:11434/v1 for Ollama). The Litellm adapter handles the translation between Cognee's structured output requirements and local API formats.
How can I verify which LLM provider is currently active in my traces?
Check the OpenTelemetry trace attributes cognee.llm.provider and cognee.llm.model emitted by the adapter in cognee/infrastructure/llm/structured_output_framework/litellm_instructor/llm/generic_llm_api/adapter.py. Alternatively, print an LLMConfig instance to inspect the loaded LLM_PROVIDER and LLM_MODEL values at runtime.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →