How to Configure PrivateGPT with Different LLM Providers: Ollama, OpenAI, Azure, SageMaker, and Gemini

Configure PrivateGPT with different LLM providers by setting the llm.mode field in a YAML profile and launching with the PGPT_PROFILES environment variable.

PrivateGPT supports multiple large language model backends through a flexible, profile-based configuration system. Whether you want to run models locally with Ollama, use cloud APIs like OpenAI or Google Gemini, or deploy on AWS SageMaker, the repository's settings architecture allows you to switch providers without modifying application code.

Understanding PrivateGPT's Configuration Architecture

PrivateGPT uses a hierarchical settings system defined in private_gpt/settings/settings_loader.py. The SettingsLoader class merges multiple YAML configuration files into a single settings object at runtime.

The loading process follows this sequence:

  1. Profile Discovery – The loader builds an active_profiles list starting with default, then appends any profiles specified in the PGPT_PROFILES environment variable (or test during test execution).

  2. File Loading – For each profile, load_settings_from_profile() reads either settings.yaml (for the default profile) or settings-<profile>.yaml (for named profiles).

  3. Deep Merge – The merge_settings() function recursively updates dictionaries, allowing later profiles to override values from earlier ones.

  4. Schema Validation – The merged dictionary instantiates the Settings Pydantic model defined in private_gpt/settings/settings.py, which exposes the llm.mode field and provider-specific subsections like ollama, openai, azopenai, sagemaker, and gemini.

How LLM Mode Selection Works

The LLMComponent class in private_gpt/components/llm/llm_component.py acts as a factory that instantiates the correct LLM client based on the settings.llm.mode value.

The component uses a Python match statement to route to the appropriate implementation:

match settings.llm.mode:
    case "ollama":
        # Instantiates Ollama client with settings.ollama parameters

    case "openai":
        # Instantiates OpenAI client with settings.openai parameters

    case "azopenai":
        # Instantiates AzureOpenAI client with settings.azopenai parameters

    case "sagemaker":
        # Instantiates SageMaker client with settings.sagemaker parameters

    case "gemini":
        # Instantiates Gemini client with settings.gemini parameters

    case "mock":
        # Uses MockLLM for testing

Each case reads its respective configuration subsection to configure API keys, endpoints, model names, and inference parameters.

Configuring Ollama for Local Inference

Ollama allows you to run open-source models locally. To configure PrivateGPT with Ollama, create or modify settings-ollama.yaml:

server:
  env_name: ${APP_ENV:ollama}

llm:
  mode: ollama
  max_new_tokens: 512
  temperature: 0.1

embedding:
  mode: ollama

ollama:
  llm_model: llama3.1
  api_base: http://localhost:11434
  keep_alive: 5m
  autopull_models: true
  tfs_z: 1.0
  top_k: 40
  top_p: 0.9

The autopull_models: true setting triggers automatic model downloading via the helper in private_gpt/utils/ollama.py, which displays a progress bar during the pull operation.

Launch with:

PGPT_PROFILES=ollama python -m private_gpt

Configuring OpenAI

For OpenAI's GPT models, use settings-openai.yaml:

server:
  env_name: ${APP_ENV:openai}

llm:
  mode: openai

embedding:
  mode: openai

openai:
  api_key: ${OPENAI_API_KEY:}
  model: gpt-4o
  api_base: https://api.openai.com/v1
  request_timeout: 30.0

You can override the API key via environment variable:

export OPENAI_API_KEY=sk-...
PGPT_PROFILES=openai python -m private_gpt

Configuring Azure OpenAI

Azure OpenAI requires additional deployment-specific parameters in settings-azopenai.yaml:

server:
  env_name: ${APP_ENV:azopenai}

llm:
  mode: azopenai

embedding:
  mode: azopenai

azopenai:
  api_key: ${AZ_OPENAI_API_KEY:}
  azure_endpoint: https://my-resource.openai.azure.com/
  llm_deployment_name: gpt-35-turbo-deployment
  embedding_deployment_name: ada-embedding-deployment
  api_version: "2023-05-15"
  llm_model: gpt-35-turbo
  embedding_model: text-embedding-ada-002

Launch with:

export AZ_OPENAI_API_KEY=...
PGPT_PROFILES=azopenai python -m private_gpt

Configuring Amazon SageMaker

For AWS SageMaker endpoints, create settings-sagemaker.yaml:

server:
  env_name: ${APP_ENV:sagemaker}
  port: ${PORT:8001}

llm:
  mode: sagemaker

embedding:
  mode: sagemaker

sagemaker:
  llm_endpoint_name: my-llm-endpoint
  embedding_endpoint_name: my-embed-endpoint

Ensure your AWS credentials are configured via standard environment variables or the ~/.aws/credentials file, then run:

PGPT_PROFILES=sagemaker python -m private_gpt

Configuring Google Gemini

For Google's Gemini models, use settings-gemini.yaml:

llm:
  mode: gemini

embedding:
  mode: gemini

gemini:
  api_key: ${GOOGLE_API_KEY:}
  model: models/gemini-pro
  embedding_model: models/embedding-001

Set your API key and launch:

export GOOGLE_API_KEY=...
PGPT_PROFILES=gemini python -m private_gpt

Switching Between Providers Using Profiles

The PGPT_PROFILES environment variable controls which configuration files are loaded. You can specify multiple profiles separated by commas for inheritance:


# Load settings.yaml, then settings-dev.yaml, then settings-ollama.yaml

PGPT_PROFILES=dev,ollama python -m private_gpt

If you do not set a profile, PrivateGPT uses the default settings.yaml. In this case, you must provide all required configuration via environment variables such as OPENAI_API_KEY, AZ_OPENAI_API_KEY, or GOOGLE_API_KEY.

Summary

  • Profile-based configuration: PrivateGPT uses YAML profiles loaded by SettingsLoader in private_gpt/settings/settings_loader.py and merged according to the PGPT_PROFILES environment variable.
  • Mode selection: The llm.mode field in your YAML determines which LLM client is instantiated by LLMComponent in private_gpt/components/llm/llm_component.py.
  • Supported providers: Ollama (ollama), OpenAI (openai), Azure OpenAI (azopenai), AWS SageMaker (sagemaker), and Google Gemini (gemini).
  • Environment variables: API keys and secrets should use the ${VAR:default} syntax in YAML or be exported directly before running PGPT_PROFILES=<profile> python -m private_gpt.

Frequently Asked Questions

How do I switch between Ollama and OpenAI without editing code?

Set the PGPT_PROFILES environment variable to the desired provider name before launching. For example, PGPT_PROFILES=ollama loads settings-ollama.yaml, while PGPT_PROFILES=openai loads settings-openai.yaml. The LLMComponent automatically instantiates the correct client based on the llm.mode value in the loaded configuration.

Can I use environment variables instead of YAML files for API keys?

Yes. PrivateGPT supports variable interpolation using the ${VAR:default} syntax in YAML files. You can define api_key: ${OPENAI_API_KEY:} in your YAML and export the actual key in your shell: export OPENAI_API_KEY=sk-.... If you prefer not to use profiles, you can also set all configuration via environment variables and run with the default settings.yaml.

What is the difference between Azure OpenAI and standard OpenAI configuration?

Azure OpenAI requires additional deployment-specific parameters including azure_endpoint, llm_deployment_name, embedding_deployment_name, and api_version. While standard OpenAI uses api_base: https://api.openai.com/v1, Azure OpenAI uses your custom endpoint URL (e.g., https://my-resource.openai.azure.com/). The mode value is also different: use azopenai instead of openai.

Does PrivateGPT support running multiple LLM providers simultaneously?

No, PrivateGPT operates with a single active LLM mode at a time as determined by the llm.mode configuration field. The LLMComponent uses a match statement to instantiate exactly one client implementation. However, you can quickly switch between providers by stopping the server, changing the PGPT_PROFILES environment variable, and restarting, or by using different profiles for different deployment environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →