Integrating nGPT with Custom OpenAI-Compatible API Endpoints: Patterns and Best Practices
nGPT integrates with any OpenAI-compatible endpoint through a configuration-first architecture that injects custom base URLs, API keys, and models via JSON configs, environment variables, or CLI flags without requiring code changes.
nGPT is an extensible CLI tool designed for seamless interaction with large language model APIs. When integrating nGPT with custom OpenAI-compatible API endpoints—such as local Ollama servers, Azure OpenAI deployments, or self-hosted inference services—the tool's modular architecture separates configuration from implementation, allowing you to switch providers instantly without modifying source code.
Configuration-First Architecture
The foundation of nGPT's flexibility lies in ngpt/core/config.py, which implements a hierarchical configuration system. This design allows custom endpoints to be defined through multiple channels without altering the core logic.
The Config Loader
The load_config function (lines 75-88) processes configuration through a priority chain: default values, JSON config file entries, environment variables, and CLI arguments. By default, the configuration contains base_url: "https://api.openai.com/v1/" (lines 10-12), but this is immediately overrideable.
The configuration file supports multiple providers as a JSON list (lines 30-44). Each entry can specify distinct base_url, api_key, and model values, enabling side-by-side management of cloud and local endpoints.
Environment Variable Overrides
For CI/CD pipelines or temporary switching, nGPT respects standard OpenAI environment variables. The load_config function checks OPENAI_BASE_URL, OPENAI_API_KEY, and OPENAI_MODEL (lines 77-80), allowing zero-file configuration of custom endpoints.
CLI Configuration Handler
The ngpt/cli/handlers/api_config_handler.py manages runtime configuration selection. When using the --provider flag, the handler extracts the matching config entry (lines 106-128) and applies any CLI overrides. The --config-index flag allows direct selection by array position.
The Generic API Client
All HTTP communication is encapsulated in ngpt/api/client.py through the NGPTClient class. This client is deliberately provider-agnostic, relying entirely on injected configuration.
Client Initialization
The constructor receives api_key, base_url, provider, and model parameters (lines 11-15). The base_url is normalized to end with / (line 18) to ensure proper path concatenation. This injection pattern means the client has no hardcoded knowledge of OpenAI's servers.
Request Construction
The client constructs requests using standard headers: Content-Type: application/json and Authorization: Bearer {api_key} (lines 22-25). Notably, the check_config method only treats missing API keys as errors when the value is explicitly None (lines 75-78), allowing empty strings for unauthenticated local endpoints.
Chat completions are sent to {base_url}chat/completions (lines 70-72) with a payload following the OpenAI schema (model, messages, stream, top_p, etc.).
Streaming Support
The client handles Server-Sent Events (SSE) parsing uniformly (lines 55-66). The streaming logic processes data: ... lines without provider-specific branching, ensuring compatibility with any endpoint that follows the OpenAI streaming format.
Managing Multiple Providers Simultaneously
The JSON configuration structure supports multiple provider entries simultaneously. The default configuration is a list (DEFAULT_CONFIG = [DEFAULT_CONFIG_ENTRY]), where each element can define a unique endpoint.
When invoking nGPT, the --provider flag selects the appropriate configuration block by matching the provider field (lines 106-128 in api_config_handler.py). This enables workflows such as:
- Development: Using a local Ollama instance for cost-free testing
- Production: Switching to Azure OpenAI or AWS Bedrock via the same CLI
- Fallback: Maintaining multiple cloud providers for redundancy
Duplicate provider names trigger warnings in show_config (lines 13-15), preventing configuration ambiguity.
Security-Aware Defaults
The codebase implements several security patterns for API key management:
- Masking: The
show_configfunction displays[Set]rather than the actual API key value, preventing shoulder-surfing or accidental logging of credentials. - Optional Authentication: Empty strings are accepted for
api_key, enabling integration with unauthenticated local development servers without dummy values. - Environment Isolation: Sensitive values can be kept entirely in environment variables (
OPENAI_API_KEY) rather than persisted to disk in JSON files.
Extending the Pattern
For endpoints requiring custom headers, additional parameters, or different API routes (such as embeddings), the existing architecture supports extension without structural changes:
- Configuration Extension: Add new fields (e.g.,
"extra_headers": {...}) to the JSON schema inngpt/core/config.pyand expose via CLI flags if desired. - Header Injection: Update
NGPTClient.__init__inngpt/api/client.pyto mergeself.headerswith custom configuration values. - New Methods: Implement additional API methods (e.g.,
def embeddings(self, ...)) that construct appropriate payloads and POST toself.base_url + "embeddings".
Because the base URL and authentication headers are already injectable, these extensions require no changes to the core CLI logic or configuration loading mechanisms.
Practical Implementation Examples
Local Ollama Server Integration
To integrate with a local Ollama instance exposing an OpenAI-compatible API:
# Interactive configuration
ngpt --config
# Enter values:
# Provider: Ollama
# Base URL: http://localhost:11434/v1/
# API Key: (press Enter to leave blank)
# Model: llama3:8b
# Usage
ngpt --provider Ollama "Explain quantum tunneling in one paragraph."
Behind the scenes, the CLI invokes load_config to extract the Ollama entry, instantiates NGPTClient with base_url="http://localhost:11434/v1/", and posts to http://localhost:11434/v1/chat/completions (see client.chat implementation).
Azure OpenAI via Environment Variables
For CI/CD pipelines or temporary switching to Azure OpenAI:
export OPENAI_BASE_URL="https://my-resource.openai.azure.com/openai/deployments/gpt-4"
export OPENAI_API_KEY="my-azure-api-key"
export OPENAI_MODEL="gpt-4"
ngpt "Summarize the latest security advisory."
The load_config function (lines 77-80) detects these environment variables and overrides any file-based configuration, enabling zero-file deployment scenarios.
Programmatic Python Usage
For embedding nGPT in existing Python applications:
from ngpt.api.client import NGPTClient
from ngpt.core.config import load_config
# Load configuration with explicit overrides
cfg = load_config(base_url="http://localhost:8000/v1/", model="custom-model")
# Initialize client
client = NGPTClient(
api_key=cfg["api_key"], # None or empty string for unauthenticated endpoints
base_url=cfg["base_url"],
provider=cfg["provider"],
model=cfg["model"]
)
# Execute chat completion
response = client.chat(
prompt="Write a haiku about machine learning.",
stream=False,
temperature=0.7
)
print(response)
This pattern works with any endpoint implementing the OpenAI Chat Completions schema, including local inference servers, cloud alternatives, and enterprise deployments.
Summary
-
Configuration-first design: nGPT separates endpoint configuration from implementation logic through
ngpt/core/config.py, enabling custom base URLs via JSON files, environment variables (OPENAI_BASE_URL), or CLI flags without code changes. -
Generic client architecture: The
NGPTClientclass inngpt/api/client.pyaccepts injectedbase_urlandapi_keyparameters, constructing standard OpenAI-compatible requests to{base_url}chat/completionswith support for streaming SSE responses. -
Multi-provider support: The configuration system maintains a list of provider entries selectable via
--providerflag inngpt/cli/handlers/api_config_handler.py, facilitating seamless switching between cloud and local endpoints. -
Security flexibility: Empty API keys are permitted for unauthenticated local servers, while environment variables prevent credential persistence in configuration files.
-
Extensibility: New API methods or custom headers require only local modifications to the client and config modules without structural changes to the CLI or loading mechanisms.
Frequently Asked Questions
How do I configure nGPT to use a local LLM server like Ollama or llama.cpp?
Create a new configuration entry using ngpt --config and specify your local server's OpenAI-compatible endpoint (typically http://localhost:11434/v1/ for Ollama). Leave the API key blank for unauthenticated local servers, then invoke with ngpt --provider YourLocalName "prompt". The load_config function in ngpt/core/config.py will route requests to your custom base_url.
Can I use nGPT with Azure OpenAI or AWS Bedrock?
Yes. Set the OPENAI_BASE_URL environment variable to your Azure OpenAI endpoint (e.g., https://my-resource.openai.azure.com/openai/deployments/gpt-4) and provide your Azure API key via OPENAI_API_KEY. Alternatively, add these values to your JSON configuration file. The NGPTClient in ngpt/api/client.py treats these as standard OpenAI-compatible endpoints.
Why does nGPT allow empty API keys?
The check_config method in ngpt/api/client.py (lines 75-78) only rejects API keys when they are explicitly None, allowing empty strings to pass through. This design supports local development servers (like Ollama or text-generation-webui) that run without authentication on localhost, while still requiring explicit configuration for production cloud APIs.
How do I switch between multiple providers without editing files?
Use the --provider CLI flag to select from your configured providers, or set environment variables (OPENAI_BASE_URL, OPENAI_API_KEY, OPENAI_MODEL) to override file-based configuration temporarily. The api_config_handler.py module (lines 106-128) resolves provider names against your configuration list, enabling instant switching between cloud and local endpoints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →