# How to Integrate DS4 with OpenAI-Compatible API Endpoints

> Integrate DS4 with OpenAI-compatible API endpoints easily. DS4 offers a built-in HTTP server at localhost:9333/v1 for seamless replacement of OpenAI's API.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**DS4 ships with a built-in HTTP server that exposes OpenAI-compatible REST endpoints on port 9333, allowing drop-in replacement of OpenAI's API by simply pointing clients to `http://localhost:9333/v1/`.**

The antirez/ds4 repository provides a lightweight, self-hosted LLM inference engine that implements the OpenAI REST API specification. When you integrate DS4 with OpenAI-compatible API endpoints, you can run local models using existing client libraries like the official OpenAI Python package or LangChain without modifying application code.

## DS4 HTTP Server Architecture

The server implementation resides in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c), with request routing logic located around line 12526. This module implements a minimal but complete HTTP server that parses incoming requests and dispatches them to the appropriate handler. The underlying socket layer and HTTP protocol handling are abstracted in [`ds4_web.c`](https://github.com/antirez/ds4/blob/main/ds4_web.c) (see line 296), which manages connection acceptance, header parsing, and response serialization via `http_response()` at line 13220.

When the server starts, it loads the specified GGUF model (default `./ds4flash.gguf`) into memory and keeps it resident for low-latency inference. All endpoint handlers forward requests to the core inference engine in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) and format responses as JSON objects matching the OpenAI schema, including fields like `id`, `object`, `created`, `model`, and `choices`.

## Available OpenAI-Compatible Endpoints

DS4 exposes five REST endpoints that mirror OpenAI's API surface:

### Model Information

**GET `/v1/models`** returns metadata about the currently loaded model, including the model ID and supported parameters.

### Chat Completions

**POST `/v1/chat/completions`** generates conversational responses compatible with OpenAI's chat format. This endpoint accepts `messages`, `temperature`, `max_tokens`, and other standard parameters.

### Text Completions

**POST `/v1/completions`** provides traditional text completion functionality for single-turn prompts, maintaining full compatibility with legacy OpenAI completion clients.

### Legacy and Internal Endpoints

**POST `/v1/messages`** serves the original DS4 client protocol, while **POST `/v1/responses`** provides diagnostic capabilities primarily used for internal testing and debugging tooling.

## Starting the DS4 Server

Before integrating client applications, you must build and launch the DS4 server binary. The build process supports multiple GPU backends:

```bash

# macOS with Metal acceleration

make

# Linux with NVIDIA CUDA

make cuda-generic

# Run the server (default port 9333)

./ds4-server

```

By default, the server listens on `http://127.0.0.1:9333`. You can specify an alternative model using the `-m` flag followed by the path to your GGUF file.

## Client Integration Examples

Once the server is running, any HTTP client can communicate with DS4 using standard OpenAI API patterns.

### Python Requests Example

```python
import requests
import json

BASE = "http://127.0.0.1:9333/v1"

# List available models

models = requests.get(f"{BASE}/models").json()
print(f"Loaded model: {models['data'][0]['id']}")

# Chat completion

chat_payload = {
    "model": "ds4flash.gguf",
    "messages": [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain quantum entanglement briefly."}
    ],
    "temperature": 0.7,
    "max_tokens": 128
}
response = requests.post(f"{BASE}/chat/completions", json=chat_payload)
print(json.dumps(response.json(), indent=2))

```

### cURL Examples

```bash

# List models

curl http://127.0.0.1:9333/v1/models

# Chat completion

curl -X POST http://127.0.0.1:9333/v1/chat/completions \
     -H "Content-Type: application/json" \
     -d '{"model":"ds4flash.gguf","messages":[{"role":"user","content":"Hello"}],"max_tokens":50}'

# Text completion

curl -X POST http://127.0.0.1:9333/v1/completions \
     -H "Content-Type: application/json" \
     -d '{"model":"ds4flash.gguf","prompt":"The capital of France is","max_tokens":10}'

```

## Multi-GPU and Batching Support

For high-throughput deployments, DS4 leverages [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) and [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c) to distribute inference across multiple GPUs. The server automatically queues incoming requests and processes them using micro-batching strategies (tested via [`test_server_batching.py`](https://github.com/antirez/ds4/blob/main/test_server_batching.py)), maximizing hardware utilization without requiring client-side changes.

The request dispatch block in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) handles concurrent connections asynchronously, while the web layer in [`ds4_web.c`](https://github.com/antirez/ds4/blob/main/ds4_web.c) manages keep-alive connections and HTTP/1.1 persistent connections for optimal performance.

## Summary

- **DS4** provides a built-in HTTP server in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) that implements OpenAI-compatible endpoints on port 9333 by default.
- **Three primary endpoints** (`/v1/models`, `/v1/chat/completions`, `/v1/completions`) enable drop-in replacement of OpenAI's hosted API.
- **Zero client modifications** are required; existing OpenAI client libraries work by changing the base URL to `http://localhost:9333/v1`.
- **Multi-GPU support** via [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) and automatic batching provide enterprise-grade throughput for local deployments.
- **JSON schema compatibility** ensures responses contain standard fields like `choices`, `usage`, and `finish_reason` that existing applications expect.

## Frequently Asked Questions

### What port does the DS4 server use by default?

The DS4 server listens on port **9333** by default when you run `./ds4-server`. You can verify the server is active by visiting `http://127.0.0.1:9333/v1/models` in a browser or using curl to check model availability.

### Do I need to modify my existing OpenAI Python code to use DS4?

No. You only need to change the `base_url` parameter in your OpenAI client configuration to point to `http://localhost:9333/v1` (or your server's IP address). The DS4 endpoints in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) return JSON structures that match the OpenAI specification exactly, including identical field names and response formats.

### Which file handles the HTTP request routing in DS4?

Request routing is implemented in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) around line 12526, where the server inspects the URL path and HTTP method to dispatch requests to the appropriate handler. The low-level HTTP protocol implementation resides in [`ds4_web.c`](https://github.com/antirez/ds4/blob/main/ds4_web.c) at line 296, which handles socket management and response formatting.

### Can DS4 handle multiple concurrent requests?

Yes. The server architecture supports micro-batching and multi-GPU inference through [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h). Requests are automatically queued and processed in parallel batches when hardware permits, making DS4 suitable for production workloads requiring concurrent API access.