# Integrating nGPT with Custom OpenAI-Compatible API Endpoints: Patterns and Best Practices

> Learn how to integrate nGPT with custom OpenAI-compatible API endpoints using configuration options. Discover patterns and best practices for seamless integration without code changes.

- Repository: [nazDridoy/ngpt](https://github.com/nazdridoy/ngpt)
- Tags: best-practices
- Published: 2026-03-07

---

**nGPT integrates with any OpenAI-compatible endpoint through a configuration-first architecture that injects custom base URLs, API keys, and models via JSON configs, environment variables, or CLI flags without requiring code changes.**

nGPT is an extensible CLI tool designed for seamless interaction with large language model APIs. When integrating nGPT with custom OpenAI-compatible API endpoints—such as local Ollama servers, Azure OpenAI deployments, or self-hosted inference services—the tool's modular architecture separates configuration from implementation, allowing you to switch providers instantly without modifying source code.

## Configuration-First Architecture

The foundation of nGPT's flexibility lies in [`ngpt/core/config.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/config.py), which implements a hierarchical configuration system. This design allows custom endpoints to be defined through multiple channels without altering the core logic.

### The Config Loader

The `load_config` function (lines 75-88) processes configuration through a priority chain: default values, JSON config file entries, environment variables, and CLI arguments. By default, the configuration contains `base_url: "https://api.openai.com/v1/"` (lines 10-12), but this is immediately overrideable.

The configuration file supports **multiple providers** as a JSON list (lines 30-44). Each entry can specify distinct `base_url`, `api_key`, and `model` values, enabling side-by-side management of cloud and local endpoints.

### Environment Variable Overrides

For CI/CD pipelines or temporary switching, nGPT respects standard OpenAI environment variables. The `load_config` function checks `OPENAI_BASE_URL`, `OPENAI_API_KEY`, and `OPENAI_MODEL` (lines 77-80), allowing zero-file configuration of custom endpoints.

### CLI Configuration Handler

The [`ngpt/cli/handlers/api_config_handler.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/handlers/api_config_handler.py) manages runtime configuration selection. When using the `--provider` flag, the handler extracts the matching config entry (lines 106-128) and applies any CLI overrides. The `--config-index` flag allows direct selection by array position.

## The Generic API Client

All HTTP communication is encapsulated in [`ngpt/api/client.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/api/client.py) through the `NGPTClient` class. This client is deliberately provider-agnostic, relying entirely on injected configuration.

### Client Initialization

The constructor receives `api_key`, `base_url`, `provider`, and `model` parameters (lines 11-15). The `base_url` is normalized to end with `/` (line 18) to ensure proper path concatenation. This injection pattern means the client has no hardcoded knowledge of OpenAI's servers.

### Request Construction

The client constructs requests using standard headers: `Content-Type: application/json` and `Authorization: Bearer {api_key}` (lines 22-25). Notably, the `check_config` method only treats missing API keys as errors when the value is explicitly `None` (lines 75-78), allowing empty strings for unauthenticated local endpoints.

Chat completions are sent to `{base_url}chat/completions` (lines 70-72) with a payload following the OpenAI schema (`model`, `messages`, `stream`, `top_p`, etc.).

### Streaming Support

The client handles Server-Sent Events (SSE) parsing uniformly (lines 55-66). The streaming logic processes `data: ...` lines without provider-specific branching, ensuring compatibility with any endpoint that follows the OpenAI streaming format.

## Managing Multiple Providers Simultaneously

The JSON configuration structure supports multiple provider entries simultaneously. The default configuration is a list (`DEFAULT_CONFIG = [DEFAULT_CONFIG_ENTRY]`), where each element can define a unique endpoint.

When invoking nGPT, the `--provider` flag selects the appropriate configuration block by matching the `provider` field (lines 106-128 in [`api_config_handler.py`](https://github.com/nazdridoy/ngpt/blob/main/api_config_handler.py)). This enables workflows such as:

- **Development**: Using a local Ollama instance for cost-free testing
- **Production**: Switching to Azure OpenAI or AWS Bedrock via the same CLI
- **Fallback**: Maintaining multiple cloud providers for redundancy

Duplicate provider names trigger warnings in `show_config` (lines 13-15), preventing configuration ambiguity.

## Security-Aware Defaults

The codebase implements several security patterns for API key management:

- **Masking**: The `show_config` function displays `[Set]` rather than the actual API key value, preventing shoulder-surfing or accidental logging of credentials.
- **Optional Authentication**: Empty strings are accepted for `api_key`, enabling integration with unauthenticated local development servers without dummy values.
- **Environment Isolation**: Sensitive values can be kept entirely in environment variables (`OPENAI_API_KEY`) rather than persisted to disk in JSON files.

## Extending the Pattern

For endpoints requiring custom headers, additional parameters, or different API routes (such as embeddings), the existing architecture supports extension without structural changes:

1. **Configuration Extension**: Add new fields (e.g., `"extra_headers": {...}`) to the JSON schema in [`ngpt/core/config.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/config.py) and expose via CLI flags if desired.
2. **Header Injection**: Update `NGPTClient.__init__` in [`ngpt/api/client.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/api/client.py) to merge `self.headers` with custom configuration values.
3. **New Methods**: Implement additional API methods (e.g., `def embeddings(self, ...)`) that construct appropriate payloads and POST to `self.base_url + "embeddings"`.

Because the base URL and authentication headers are already injectable, these extensions require no changes to the core CLI logic or configuration loading mechanisms.

## Practical Implementation Examples

### Local Ollama Server Integration

To integrate with a local Ollama instance exposing an OpenAI-compatible API:

```bash

# Interactive configuration

ngpt --config

# Enter values:

#   Provider: Ollama

#   Base URL: http://localhost:11434/v1/

#   API Key: (press Enter to leave blank)

#   Model: llama3:8b

# Usage

ngpt --provider Ollama "Explain quantum tunneling in one paragraph."

```

Behind the scenes, the CLI invokes `load_config` to extract the Ollama entry, instantiates `NGPTClient` with `base_url="http://localhost:11434/v1/"`, and posts to `http://localhost:11434/v1/chat/completions` (see `client.chat` implementation).

### Azure OpenAI via Environment Variables

For CI/CD pipelines or temporary switching to Azure OpenAI:

```bash
export OPENAI_BASE_URL="https://my-resource.openai.azure.com/openai/deployments/gpt-4"
export OPENAI_API_KEY="my-azure-api-key"
export OPENAI_MODEL="gpt-4"
ngpt "Summarize the latest security advisory."

```

The `load_config` function (lines 77-80) detects these environment variables and overrides any file-based configuration, enabling zero-file deployment scenarios.

### Programmatic Python Usage

For embedding nGPT in existing Python applications:

```python
from ngpt.api.client import NGPTClient
from ngpt.core.config import load_config

# Load configuration with explicit overrides

cfg = load_config(base_url="http://localhost:8000/v1/", model="custom-model")

# Initialize client

client = NGPTClient(
    api_key=cfg["api_key"],        # None or empty string for unauthenticated endpoints

    base_url=cfg["base_url"],
    provider=cfg["provider"],
    model=cfg["model"]
)

# Execute chat completion

response = client.chat(
    prompt="Write a haiku about machine learning.",
    stream=False,
    temperature=0.7
)
print(response)

```

This pattern works with any endpoint implementing the OpenAI Chat Completions schema, including local inference servers, cloud alternatives, and enterprise deployments.

## Summary

- **Configuration-first design**: nGPT separates endpoint configuration from implementation logic through [`ngpt/core/config.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/config.py), enabling custom base URLs via JSON files, environment variables (`OPENAI_BASE_URL`), or CLI flags without code changes.

- **Generic client architecture**: The `NGPTClient` class in [`ngpt/api/client.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/api/client.py) accepts injected `base_url` and `api_key` parameters, constructing standard OpenAI-compatible requests to `{base_url}chat/completions` with support for streaming SSE responses.

- **Multi-provider support**: The configuration system maintains a list of provider entries selectable via `--provider` flag in [`ngpt/cli/handlers/api_config_handler.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/cli/handlers/api_config_handler.py), facilitating seamless switching between cloud and local endpoints.

- **Security flexibility**: Empty API keys are permitted for unauthenticated local servers, while environment variables prevent credential persistence in configuration files.

- **Extensibility**: New API methods or custom headers require only local modifications to the client and config modules without structural changes to the CLI or loading mechanisms.

## Frequently Asked Questions

### How do I configure nGPT to use a local LLM server like Ollama or llama.cpp?

Create a new configuration entry using `ngpt --config` and specify your local server's OpenAI-compatible endpoint (typically `http://localhost:11434/v1/` for Ollama). Leave the API key blank for unauthenticated local servers, then invoke with `ngpt --provider YourLocalName "prompt"`. The `load_config` function in [`ngpt/core/config.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/core/config.py) will route requests to your custom base_url.

### Can I use nGPT with Azure OpenAI or AWS Bedrock?

Yes. Set the `OPENAI_BASE_URL` environment variable to your Azure OpenAI endpoint (e.g., `https://my-resource.openai.azure.com/openai/deployments/gpt-4`) and provide your Azure API key via `OPENAI_API_KEY`. Alternatively, add these values to your JSON configuration file. The `NGPTClient` in [`ngpt/api/client.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/api/client.py) treats these as standard OpenAI-compatible endpoints.

### Why does nGPT allow empty API keys?

The `check_config` method in [`ngpt/api/client.py`](https://github.com/nazdridoy/ngpt/blob/main/ngpt/api/client.py) (lines 75-78) only rejects API keys when they are explicitly `None`, allowing empty strings to pass through. This design supports local development servers (like Ollama or text-generation-webui) that run without authentication on localhost, while still requiring explicit configuration for production cloud APIs.

### How do I switch between multiple providers without editing files?

Use the `--provider` CLI flag to select from your configured providers, or set environment variables (`OPENAI_BASE_URL`, `OPENAI_API_KEY`, `OPENAI_MODEL`) to override file-based configuration temporarily. The [`api_config_handler.py`](https://github.com/nazdridoy/ngpt/blob/main/api_config_handler.py) module (lines 106-128) resolves provider names against your configuration list, enabling instant switching between cloud and local endpoints.