# How Chat2DB Integrates Its AI Assistant with Custom LLM Models

> Discover how Chat2DB integrates its AI assistant with custom LLM models using a three-layer architecture. Connect securely, stream responses, and personalize your AI experience.

- Repository: [OtterMind/Chat2DB](https://github.com/OtterMind/Chat2DB)
- Tags: deep-dive
- Published: 2026-07-28

---

**The AI assistant connects to custom LLM endpoints through a configurable three-layer architecture that persists user-defined API keys and base URLs, then streams chat responses via Server-Sent Events.**

Chat2DB enables users to plug in their own Large Language Models (LLMs) through a flexible integration layer. This article examines how the OtterMind/Chat2DB repository implements custom model support, from TypeScript service definitions to Java domain services that construct runtime HTTP clients for any OpenAI-compatible endpoint.

## Three-Layer Integration Architecture

The integration follows a clean separation between the React-based frontend, REST controllers, and domain services.

### Frontend Service Layer

Located in the client module, the frontend defines TypeScript wrappers that call backend endpoints. The [[`startup.ts`](https://github.com/OtterMind/Chat2DB/blob/main/startup.ts)](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-client/src/service/llm/startup.ts) service handles CRUD operations for custom LLM configurations, while [[`fineTuning.ts`](https://github.com/OtterMind/Chat2DB/blob/main/fineTuning.ts)](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-client/src/service/llm/fineTuning.ts) manages fine-tuning workflows. Type definitions in [[`typings/llm/index.ts`](https://github.com/OtterMind/Chat2DB/blob/main/typings/llm/index.ts)](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-client/src/typings/llm/index.ts) enforce the `ILLMStartup` interface, ensuring that `modelName`, `baseUrl`, and `apiKey` are type-safe across the UI.

### Backend Controllers

The [[`AiChatController.java`](https://github.com/OtterMind/Chat2DB/blob/main/AiChatController.java)](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-web/src/main/java/ai/chat2db/community/web/api/controller/AiChatController.java) exposes REST endpoints under `/api/v3/ai/**`. Key routes include:

- `GET /api/v3/ai/model/list` – Retrieves built-in and user-defined models.
- `POST /api/v3/ai/model/config/save` – Persists a new custom LLM configuration.
- `POST /api/v3/ai/chat/stream` – Initiates a streaming chat session using Server-Sent Events (SSE).

### Domain Services

The business logic resides in the domain layer. [`IAiModelConfigService`](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-domain-api/src/main/java/ai/chat2db/community/domain/api/service/ai/IAiModelConfigService.java) defines the contract for model configuration CRUD, while [`AiModelConfigServiceImpl`](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-domain-core/src/main/java/ai/chat2db/community/domain/core/impl/ai/AiModelConfigServiceImpl.java) implements the actual persistence and runtime client construction.

## Configuring Custom LLM Endpoints

Users define custom models by submitting a payload containing the endpoint URL, authentication credentials, and model identifier.

### Model Configuration Persistence

When the frontend calls `/api/v3/ai/model/config/save`, the controller delegates to `IAiModelConfigService.saveCurrentUserConfig()`. The implementation stores the `AiModelConfigParam`—which includes `modelName`, `baseUrl`, `apiKey`, and optional `requestHeaders`—in the workspace storage layer. This allows multiple custom configurations per user.

### Runtime Client Construction

Before each chat request, `AiModelConfigServiceImpl` retrieves the stored configuration by ID and constructs a generic HTTP client. This client injects the `apiKey` into the `Authorization` header (or custom headers) and targets the user-provided `baseUrl`. Because the client is built at runtime from persisted JSON, Chat2DB can connect to OpenAI, Azure, Anthropic, or any third-party endpoint without code changes.

## Streaming Chat Implementation

### The Chat Stream Endpoint

The `POST /api/v3/ai/chat/stream` endpoint in [[`AiChatController.java`](https://github.com/OtterMind/Chat2DB/blob/main/AiChatController.java)](https://github.com/OtterMind/Chat2DB/blob/main/chat2db-community-web/src/main/java/ai/chat2db/community/web/api/controller/AiChatController.java) accepts a `ChatRequest` and returns an `SseEmitter`. The controller passes the request to `IAiChatStreamService`, which:

1. Resolves the `modelId` to a concrete `AiModelConfigResponse`.
2. Instantiates the HTTP LLM client using the saved configuration.
3. Forwards the prompt to the remote endpoint and streams chunks back through the emitter.

### Type-Safe Frontend Consumption

The frontend consumes this SSE stream using the chat service. The TypeScript layer handles message parsing and UI updates as chunks arrive.

## Practical Code Examples

### Saving a Custom Model Configuration (TypeScript)

```typescript
import aiModelService from '@/service/aiModel';

// Persist a custom OpenAI-compatible endpoint
await aiModelService.saveModelConfig({
  modelName: 'CustomGPT-4',
  baseUrl: 'https://api.custom-provider.com/v1/chat/completions',
  apiKey: 'sk-custom-key-12345',
  requestHeaders: { 'X-Custom-Header': 'value' }
});

```

### Initiating a Streaming Chat (TypeScript)

```typescript
import chatService from '@/service/chat';

const response = await chatService.stream({
  modelId: 42, // ID of the saved custom model
  messages: [{ role: 'user', content: 'Optimize this SQL query...' }],
  stream: true
});

// Handle Server-Sent Events
response.onmessage = (event) => {
  const chunk = JSON.parse(event.data);
  console.log('Received:', chunk.content);
};

```

### Backend Client Resolution (Java)

```java
// Simplified excerpt from AiModelConfigServiceImpl
public AiLlmClient buildClient(Long modelId) {
    AiModelConfig config = repository.findById(modelId)
        .orElseThrow(() -> new ModelNotFoundException(modelId));
    
    return GenericHttpLlmClient.builder()
        .baseUrl(config.getBaseUrl())
        .apiKey(config.getApiKey())
        .headers(config.getRequestHeaders())
        .build();
}

```

## Summary

- Chat2DB separates concerns into frontend TypeScript services, REST controllers, and Java domain services.
- Custom LLM configurations are persisted via `IAiModelConfigService` and `AiModelConfigServiceImpl`, allowing per-user API keys and endpoints.
- The `AiChatController` exposes endpoints for model management (`/model/list`, `/model/config/save`) and streaming chat (`/chat/stream`).
- Runtime client construction enables connection to any OpenAI-compatible endpoint without recompiling the application.
- Server-Sent Events deliver real-time responses from custom models to the React frontend.

## Frequently Asked Questions

### How do I add a private LLM endpoint to Chat2DB?

Navigate to the AI settings panel in the UI, select "Add Custom Model," and provide the base URL, API key, and model name. This calls `POST /api/v3/ai/model/config/save`, which persists the configuration through `AiModelConfigServiceImpl` for your user account.

### Can I use multiple custom LLM providers simultaneously?

Yes. The `IAiModelConfigService` supports multiple configurations per user. Each chat request specifies a `modelId`, allowing you to switch between providers like Azure, Anthropic, and self-hosted models on a per-conversation basis.

### Does Chat2DB support streaming responses from custom models?

Yes. The `/api/v3/ai/chat/stream` endpoint returns an `SseEmitter` that streams tokens as they arrive from the custom endpoint. The frontend service consumes these chunks via the browser's EventSource API or fetch-based streaming.

### Where is the API key for my custom model stored?

The API key is stored in the workspace database via the domain service layer (`AiModelConfigServiceImpl`). It is retrieved at runtime when constructing the HTTP client for each chat request, ensuring credentials are not exposed to the frontend after initial configuration.