# How to Set Up the DS4 Server with OpenAI-Compatible API Endpoints

> Set up the DS4 server with OpenAI compatible API endpoints for local inference. Easily replace OpenAI API calls with DS4's built-in HTTP server.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**DS4 ships with a built-in HTTP server that exposes OpenAI-compatible REST endpoints on port 9333, allowing drop-in replacement of OpenAI API calls with local inference.**

The **antirez/ds4** repository provides a lightweight, high-performance LLM inference engine that includes a fully functional HTTP server implementing the OpenAI REST API specification. By running the `ds4-server` binary, you can host local models and interact with them using the same JSON schema and endpoint structure as OpenAI's hosted service.

## Understanding the DS4 Server Architecture

The server implementation resides primarily in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c), where a request-dispatch block around **line 12526** routes incoming HTTP traffic to the appropriate handlers. The web layer relies on [`ds4_web.c`](https://github.com/antirez/ds4/blob/main/ds4_web.c) (see the HTTP handling utilities starting at **line 296**) for socket management, request parsing, and response formatting.

When a request arrives, the server invokes the core inference engine defined in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) to generate logits and decode tokens. The handlers then format these results as JSON objects that mirror the OpenAI schema, including fields such as `id`, `object`, `created`, `model`, and `choices`. For response serialization, the code calls `http_response()` (located around **line 13220** in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)) to stream the JSON payload back to the client.

## Building and Starting the Server

Before exposing the API, compile the server for your target platform. DS4 supports Metal on macOS, CUDA on Linux, and ROCm for AMD GPUs.

```bash

# macOS with Metal

make

# Linux with CUDA

make cuda-generic

```

Launch the server with the default configuration:

```bash
./ds4-server

```

By default, the process binds to `http://127.0.0.1:9333` and loads the model file `ds4flash.gguf` from the current directory. Specify a custom model path using the `-m` flag:

```bash
./ds4-server -m /path/to/your/model.gguf

```

The server keeps the loaded model resident in memory, ensuring low-latency inference for subsequent requests.

## Available OpenAI-Compatible Endpoints

Once running, `ds4-server` exposes the following REST routes that follow the OpenAI API specification:

- **`/v1/models`** (GET) — Returns metadata about the currently served model, including the model ID and supported parameters.

- **`/v1/chat/completions`** (POST) — Generates chat-style completions compatible with OpenAI's Chat Completion API. Accepts a JSON payload with `messages`, `temperature`, `max_tokens`, and other standard parameters.

- **`/v1/completions`** (POST) — Generates plain-text completions compatible with the legacy OpenAI Completions API. Requires a `prompt` field and supports standard sampling controls.

- **`/v1/messages`** (POST) — Legacy endpoint used by the original DS4 client (retained for backward compatibility).

- **`/v1/responses`** (POST) — Internal diagnostic endpoint primarily used for debugging and monitoring.

## Making Requests to the DS4 Server

### Using Python and Requests

The following Python script demonstrates how to query all major endpoints using the standard `requests` library:

```python
import requests
import json

BASE = "http://127.0.0.1:9333/v1"

# List available models

resp = requests.get(f"{BASE}/models")
print("Models:", resp.json())

# Chat completion (OpenAI-compatible)

payload = {
    "model": "ds4flash.gguf",
    "messages": [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain quantum entanglement in one sentence."}
    ],
    "temperature": 0.7,
    "max_tokens": 128
}
chat_resp = requests.post(f"{BASE}/chat/completions", json=payload)
print(json.dumps(chat_resp.json(), indent=2))

# Text completion (OpenAI-compatible)

payload = {
    "model": "ds4flash.gguf",
    "prompt": "Write a haiku about sunrise:",
    "temperature": 0.9,
    "max_tokens": 32
}
comp_resp = requests.post(f"{BASE}/completions", json=payload)
print(json.dumps(comp_resp.json(), indent=2))

```

### Using cURL

For quick testing or shell scripts, use `curl` to interact with the endpoints directly:

```bash

# List available models

curl http://127.0.0.1:9333/v1/models

# Chat completion

curl -X POST http://127.0.0.1:9333/v1/chat/completions \
     -H "Content-Type: application/json" \
     -d '{"model":"ds4flash.gguf","messages":[{"role":"user","content":"Translate \"hello\" to French."}],"max_tokens":16}'

# Text completion

curl -X POST http://127.0.0.1:9333/v1/completions \
     -H "Content-Type: application/json" \
     -d '{"model":"ds4flash.gguf","prompt":"The future of AI is","max_tokens":32}'

```

These requests return JSON objects that can be consumed by any OpenAI-compatible client library, including the official `openai` Python package and LangChain, without requiring code modifications.

## Advanced Configuration and Batching

For high-throughput deployments, DS4 supports **micro-batching** and **multi-GPU** inference. The server automatically queues incoming requests and processes them in the most efficient order, routing computation across available GPUs using the logic defined in [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) and [`ds4_gpu_args.c`](https://github.com/antirez/ds4/blob/main/ds4_gpu_args.c).

When running with multiple GPUs, the server distributes layers across devices to maximize memory bandwidth and compute utilization. This architecture enables the `ds4-server` to handle concurrent requests while maintaining the simple, stateless REST interface required for OpenAI compatibility.

## Summary

- **ds4-server** implements OpenAI-compatible endpoints in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c), listening on port **9333** by default.
- The server supports **`/v1/models`**, **`/v1/chat/completions`**, and **`/v1/completions`** using standard JSON schemas.
- Build with `make` (macOS) or `make cuda-generic` (Linux), then launch with `./ds4-server -m <model.gguf>`.
- The underlying HTTP layer in [`ds4_web.c`](https://github.com/antirez/ds4/blob/main/ds4_web.c) handles request parsing, while [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) manages the inference engine.
- Multi-GPU batching is supported via [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) for scalable production deployments.

## Frequently Asked Questions

### What port does ds4-server use by default?

By default, the server listens on **port 9333** bound to `127.0.0.1`. You can modify the listening address and port by passing command-line flags or environment variables supported by the underlying socket implementation in [`ds4_web.c`](https://github.com/antirez/ds4/blob/main/ds4_web.c).

### Is the DS4 API fully compatible with OpenAI's Python SDK?

Yes. Because `ds4-server` returns JSON payloads that match the OpenAI schema exactly—including `id`, `object`, `created`, `model`, and `choices` fields—you can point the official `openai` Python library or LangChain connectors to `http://127.0.0.1:9333/v1` and use standard methods like `ChatCompletion.create()` without code changes.

### How do I load a custom model in ds4-server?

Pass the path to your GGUF file using the **`-m`** flag when starting the server: `./ds4-server -m /path/to/custom-model.gguf`. If no model is specified, the server attempts to load `ds4flash.gguf` from the current working directory.

### Does ds4-server support multi-GPU inference?

Yes. The server includes support for multi-GPU routing and micro-batching through the headers [`ds4_gpu_mgpu.h`](https://github.com/antirez/ds4/blob/main/ds4_gpu_mgpu.h) and related configuration files. When multiple GPUs are detected, the server automatically distributes model layers across devices to optimize throughput for concurrent API requests.