How to Configure PrivateGPT with Different LLM Providers: Ollama, OpenAI, Azure, SageMaker, and Gemini
Configure PrivateGPT with different LLM providers by setting the llm.mode field in a YAML profile and launching with the PGPT_PROFILES environment variable.
PrivateGPT supports multiple large language model backends through a flexible, profile-based configuration system. Whether you want to run models locally with Ollama, use cloud APIs like OpenAI or Google Gemini, or deploy on AWS SageMaker, the repository's settings architecture allows you to switch providers without modifying application code.
Understanding PrivateGPT's Configuration Architecture
PrivateGPT uses a hierarchical settings system defined in private_gpt/settings/settings_loader.py. The SettingsLoader class merges multiple YAML configuration files into a single settings object at runtime.
The loading process follows this sequence:
-
Profile Discovery – The loader builds an
active_profileslist starting withdefault, then appends any profiles specified in thePGPT_PROFILESenvironment variable (ortestduring test execution). -
File Loading – For each profile,
load_settings_from_profile()reads eithersettings.yaml(for the default profile) orsettings-<profile>.yaml(for named profiles). -
Deep Merge – The
merge_settings()function recursively updates dictionaries, allowing later profiles to override values from earlier ones. -
Schema Validation – The merged dictionary instantiates the
SettingsPydantic model defined inprivate_gpt/settings/settings.py, which exposes thellm.modefield and provider-specific subsections likeollama,openai,azopenai,sagemaker, andgemini.
How LLM Mode Selection Works
The LLMComponent class in private_gpt/components/llm/llm_component.py acts as a factory that instantiates the correct LLM client based on the settings.llm.mode value.
The component uses a Python match statement to route to the appropriate implementation:
match settings.llm.mode:
case "ollama":
# Instantiates Ollama client with settings.ollama parameters
case "openai":
# Instantiates OpenAI client with settings.openai parameters
case "azopenai":
# Instantiates AzureOpenAI client with settings.azopenai parameters
case "sagemaker":
# Instantiates SageMaker client with settings.sagemaker parameters
case "gemini":
# Instantiates Gemini client with settings.gemini parameters
case "mock":
# Uses MockLLM for testing
Each case reads its respective configuration subsection to configure API keys, endpoints, model names, and inference parameters.
Configuring Ollama for Local Inference
Ollama allows you to run open-source models locally. To configure PrivateGPT with Ollama, create or modify settings-ollama.yaml:
server:
env_name: ${APP_ENV:ollama}
llm:
mode: ollama
max_new_tokens: 512
temperature: 0.1
embedding:
mode: ollama
ollama:
llm_model: llama3.1
api_base: http://localhost:11434
keep_alive: 5m
autopull_models: true
tfs_z: 1.0
top_k: 40
top_p: 0.9
The autopull_models: true setting triggers automatic model downloading via the helper in private_gpt/utils/ollama.py, which displays a progress bar during the pull operation.
Launch with:
PGPT_PROFILES=ollama python -m private_gpt
Configuring OpenAI
For OpenAI's GPT models, use settings-openai.yaml:
server:
env_name: ${APP_ENV:openai}
llm:
mode: openai
embedding:
mode: openai
openai:
api_key: ${OPENAI_API_KEY:}
model: gpt-4o
api_base: https://api.openai.com/v1
request_timeout: 30.0
You can override the API key via environment variable:
export OPENAI_API_KEY=sk-...
PGPT_PROFILES=openai python -m private_gpt
Configuring Azure OpenAI
Azure OpenAI requires additional deployment-specific parameters in settings-azopenai.yaml:
server:
env_name: ${APP_ENV:azopenai}
llm:
mode: azopenai
embedding:
mode: azopenai
azopenai:
api_key: ${AZ_OPENAI_API_KEY:}
azure_endpoint: https://my-resource.openai.azure.com/
llm_deployment_name: gpt-35-turbo-deployment
embedding_deployment_name: ada-embedding-deployment
api_version: "2023-05-15"
llm_model: gpt-35-turbo
embedding_model: text-embedding-ada-002
Launch with:
export AZ_OPENAI_API_KEY=...
PGPT_PROFILES=azopenai python -m private_gpt
Configuring Amazon SageMaker
For AWS SageMaker endpoints, create settings-sagemaker.yaml:
server:
env_name: ${APP_ENV:sagemaker}
port: ${PORT:8001}
llm:
mode: sagemaker
embedding:
mode: sagemaker
sagemaker:
llm_endpoint_name: my-llm-endpoint
embedding_endpoint_name: my-embed-endpoint
Ensure your AWS credentials are configured via standard environment variables or the ~/.aws/credentials file, then run:
PGPT_PROFILES=sagemaker python -m private_gpt
Configuring Google Gemini
For Google's Gemini models, use settings-gemini.yaml:
llm:
mode: gemini
embedding:
mode: gemini
gemini:
api_key: ${GOOGLE_API_KEY:}
model: models/gemini-pro
embedding_model: models/embedding-001
Set your API key and launch:
export GOOGLE_API_KEY=...
PGPT_PROFILES=gemini python -m private_gpt
Switching Between Providers Using Profiles
The PGPT_PROFILES environment variable controls which configuration files are loaded. You can specify multiple profiles separated by commas for inheritance:
# Load settings.yaml, then settings-dev.yaml, then settings-ollama.yaml
PGPT_PROFILES=dev,ollama python -m private_gpt
If you do not set a profile, PrivateGPT uses the default settings.yaml. In this case, you must provide all required configuration via environment variables such as OPENAI_API_KEY, AZ_OPENAI_API_KEY, or GOOGLE_API_KEY.
Summary
- Profile-based configuration: PrivateGPT uses YAML profiles loaded by
SettingsLoaderinprivate_gpt/settings/settings_loader.pyand merged according to thePGPT_PROFILESenvironment variable. - Mode selection: The
llm.modefield in your YAML determines which LLM client is instantiated byLLMComponentinprivate_gpt/components/llm/llm_component.py. - Supported providers: Ollama (
ollama), OpenAI (openai), Azure OpenAI (azopenai), AWS SageMaker (sagemaker), and Google Gemini (gemini). - Environment variables: API keys and secrets should use the
${VAR:default}syntax in YAML or be exported directly before runningPGPT_PROFILES=<profile> python -m private_gpt.
Frequently Asked Questions
How do I switch between Ollama and OpenAI without editing code?
Set the PGPT_PROFILES environment variable to the desired provider name before launching. For example, PGPT_PROFILES=ollama loads settings-ollama.yaml, while PGPT_PROFILES=openai loads settings-openai.yaml. The LLMComponent automatically instantiates the correct client based on the llm.mode value in the loaded configuration.
Can I use environment variables instead of YAML files for API keys?
Yes. PrivateGPT supports variable interpolation using the ${VAR:default} syntax in YAML files. You can define api_key: ${OPENAI_API_KEY:} in your YAML and export the actual key in your shell: export OPENAI_API_KEY=sk-.... If you prefer not to use profiles, you can also set all configuration via environment variables and run with the default settings.yaml.
What is the difference between Azure OpenAI and standard OpenAI configuration?
Azure OpenAI requires additional deployment-specific parameters including azure_endpoint, llm_deployment_name, embedding_deployment_name, and api_version. While standard OpenAI uses api_base: https://api.openai.com/v1, Azure OpenAI uses your custom endpoint URL (e.g., https://my-resource.openai.azure.com/). The mode value is also different: use azopenai instead of openai.
Does PrivateGPT support running multiple LLM providers simultaneously?
No, PrivateGPT operates with a single active LLM mode at a time as determined by the llm.mode configuration field. The LLMComponent uses a match statement to instantiate exactly one client implementation. However, you can quickly switch between providers by stopping the server, changing the PGPT_PROFILES environment variable, and restarting, or by using different profiles for different deployment environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →