How to Set Up the DS4 Server with OpenAI-Compatible API Endpoints
DS4 ships with a built-in HTTP server that exposes OpenAI-compatible REST endpoints on port 9333, allowing drop-in replacement of OpenAI API calls with local inference.
The antirez/ds4 repository provides a lightweight, high-performance LLM inference engine that includes a fully functional HTTP server implementing the OpenAI REST API specification. By running the ds4-server binary, you can host local models and interact with them using the same JSON schema and endpoint structure as OpenAI's hosted service.
Understanding the DS4 Server Architecture
The server implementation resides primarily in ds4_server.c, where a request-dispatch block around line 12526 routes incoming HTTP traffic to the appropriate handlers. The web layer relies on ds4_web.c (see the HTTP handling utilities starting at line 296) for socket management, request parsing, and response formatting.
When a request arrives, the server invokes the core inference engine defined in ds4.c to generate logits and decode tokens. The handlers then format these results as JSON objects that mirror the OpenAI schema, including fields such as id, object, created, model, and choices. For response serialization, the code calls http_response() (located around line 13220 in ds4_server.c) to stream the JSON payload back to the client.
Building and Starting the Server
Before exposing the API, compile the server for your target platform. DS4 supports Metal on macOS, CUDA on Linux, and ROCm for AMD GPUs.
# macOS with Metal
make
# Linux with CUDA
make cuda-generic
Launch the server with the default configuration:
./ds4-server
By default, the process binds to http://127.0.0.1:9333 and loads the model file ds4flash.gguf from the current directory. Specify a custom model path using the -m flag:
./ds4-server -m /path/to/your/model.gguf
The server keeps the loaded model resident in memory, ensuring low-latency inference for subsequent requests.
Available OpenAI-Compatible Endpoints
Once running, ds4-server exposes the following REST routes that follow the OpenAI API specification:
-
/v1/models(GET) — Returns metadata about the currently served model, including the model ID and supported parameters. -
/v1/chat/completions(POST) — Generates chat-style completions compatible with OpenAI's Chat Completion API. Accepts a JSON payload withmessages,temperature,max_tokens, and other standard parameters. -
/v1/completions(POST) — Generates plain-text completions compatible with the legacy OpenAI Completions API. Requires apromptfield and supports standard sampling controls. -
/v1/messages(POST) — Legacy endpoint used by the original DS4 client (retained for backward compatibility). -
/v1/responses(POST) — Internal diagnostic endpoint primarily used for debugging and monitoring.
Making Requests to the DS4 Server
Using Python and Requests
The following Python script demonstrates how to query all major endpoints using the standard requests library:
import requests
import json
BASE = "http://127.0.0.1:9333/v1"
# List available models
resp = requests.get(f"{BASE}/models")
print("Models:", resp.json())
# Chat completion (OpenAI-compatible)
payload = {
"model": "ds4flash.gguf",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum entanglement in one sentence."}
],
"temperature": 0.7,
"max_tokens": 128
}
chat_resp = requests.post(f"{BASE}/chat/completions", json=payload)
print(json.dumps(chat_resp.json(), indent=2))
# Text completion (OpenAI-compatible)
payload = {
"model": "ds4flash.gguf",
"prompt": "Write a haiku about sunrise:",
"temperature": 0.9,
"max_tokens": 32
}
comp_resp = requests.post(f"{BASE}/completions", json=payload)
print(json.dumps(comp_resp.json(), indent=2))
Using cURL
For quick testing or shell scripts, use curl to interact with the endpoints directly:
# List available models
curl http://127.0.0.1:9333/v1/models
# Chat completion
curl -X POST http://127.0.0.1:9333/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"ds4flash.gguf","messages":[{"role":"user","content":"Translate \"hello\" to French."}],"max_tokens":16}'
# Text completion
curl -X POST http://127.0.0.1:9333/v1/completions \
-H "Content-Type: application/json" \
-d '{"model":"ds4flash.gguf","prompt":"The future of AI is","max_tokens":32}'
These requests return JSON objects that can be consumed by any OpenAI-compatible client library, including the official openai Python package and LangChain, without requiring code modifications.
Advanced Configuration and Batching
For high-throughput deployments, DS4 supports micro-batching and multi-GPU inference. The server automatically queues incoming requests and processes them in the most efficient order, routing computation across available GPUs using the logic defined in ds4_gpu_mgpu.h and ds4_gpu_args.c.
When running with multiple GPUs, the server distributes layers across devices to maximize memory bandwidth and compute utilization. This architecture enables the ds4-server to handle concurrent requests while maintaining the simple, stateless REST interface required for OpenAI compatibility.
Summary
- ds4-server implements OpenAI-compatible endpoints in
ds4_server.c, listening on port 9333 by default. - The server supports
/v1/models,/v1/chat/completions, and/v1/completionsusing standard JSON schemas. - Build with
make(macOS) ormake cuda-generic(Linux), then launch with./ds4-server -m <model.gguf>. - The underlying HTTP layer in
ds4_web.chandles request parsing, whileds4.cmanages the inference engine. - Multi-GPU batching is supported via
ds4_gpu_mgpu.hfor scalable production deployments.
Frequently Asked Questions
What port does ds4-server use by default?
By default, the server listens on port 9333 bound to 127.0.0.1. You can modify the listening address and port by passing command-line flags or environment variables supported by the underlying socket implementation in ds4_web.c.
Is the DS4 API fully compatible with OpenAI's Python SDK?
Yes. Because ds4-server returns JSON payloads that match the OpenAI schema exactly—including id, object, created, model, and choices fields—you can point the official openai Python library or LangChain connectors to http://127.0.0.1:9333/v1 and use standard methods like ChatCompletion.create() without code changes.
How do I load a custom model in ds4-server?
Pass the path to your GGUF file using the -m flag when starting the server: ./ds4-server -m /path/to/custom-model.gguf. If no model is specified, the server attempts to load ds4flash.gguf from the current working directory.
Does ds4-server support multi-GPU inference?
Yes. The server includes support for multi-GPU routing and micro-batching through the headers ds4_gpu_mgpu.h and related configuration files. When multiple GPUs are detected, the server automatically distributes model layers across devices to optimize throughput for concurrent API requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →