How to Integrate DS4 with OpenAI-Compatible API Endpoints
DS4 ships with a built-in HTTP server that exposes OpenAI-compatible REST endpoints on port 9333, allowing drop-in replacement of OpenAI's API by simply pointing clients to http://localhost:9333/v1/.
The antirez/ds4 repository provides a lightweight, self-hosted LLM inference engine that implements the OpenAI REST API specification. When you integrate DS4 with OpenAI-compatible API endpoints, you can run local models using existing client libraries like the official OpenAI Python package or LangChain without modifying application code.
DS4 HTTP Server Architecture
The server implementation resides in ds4_server.c, with request routing logic located around line 12526. This module implements a minimal but complete HTTP server that parses incoming requests and dispatches them to the appropriate handler. The underlying socket layer and HTTP protocol handling are abstracted in ds4_web.c (see line 296), which manages connection acceptance, header parsing, and response serialization via http_response() at line 13220.
When the server starts, it loads the specified GGUF model (default ./ds4flash.gguf) into memory and keeps it resident for low-latency inference. All endpoint handlers forward requests to the core inference engine in ds4.c and format responses as JSON objects matching the OpenAI schema, including fields like id, object, created, model, and choices.
Available OpenAI-Compatible Endpoints
DS4 exposes five REST endpoints that mirror OpenAI's API surface:
Model Information
GET /v1/models returns metadata about the currently loaded model, including the model ID and supported parameters.
Chat Completions
POST /v1/chat/completions generates conversational responses compatible with OpenAI's chat format. This endpoint accepts messages, temperature, max_tokens, and other standard parameters.
Text Completions
POST /v1/completions provides traditional text completion functionality for single-turn prompts, maintaining full compatibility with legacy OpenAI completion clients.
Legacy and Internal Endpoints
POST /v1/messages serves the original DS4 client protocol, while POST /v1/responses provides diagnostic capabilities primarily used for internal testing and debugging tooling.
Starting the DS4 Server
Before integrating client applications, you must build and launch the DS4 server binary. The build process supports multiple GPU backends:
# macOS with Metal acceleration
make
# Linux with NVIDIA CUDA
make cuda-generic
# Run the server (default port 9333)
./ds4-server
By default, the server listens on http://127.0.0.1:9333. You can specify an alternative model using the -m flag followed by the path to your GGUF file.
Client Integration Examples
Once the server is running, any HTTP client can communicate with DS4 using standard OpenAI API patterns.
Python Requests Example
import requests
import json
BASE = "http://127.0.0.1:9333/v1"
# List available models
models = requests.get(f"{BASE}/models").json()
print(f"Loaded model: {models['data'][0]['id']}")
# Chat completion
chat_payload = {
"model": "ds4flash.gguf",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum entanglement briefly."}
],
"temperature": 0.7,
"max_tokens": 128
}
response = requests.post(f"{BASE}/chat/completions", json=chat_payload)
print(json.dumps(response.json(), indent=2))
cURL Examples
# List models
curl http://127.0.0.1:9333/v1/models
# Chat completion
curl -X POST http://127.0.0.1:9333/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"ds4flash.gguf","messages":[{"role":"user","content":"Hello"}],"max_tokens":50}'
# Text completion
curl -X POST http://127.0.0.1:9333/v1/completions \
-H "Content-Type: application/json" \
-d '{"model":"ds4flash.gguf","prompt":"The capital of France is","max_tokens":10}'
Multi-GPU and Batching Support
For high-throughput deployments, DS4 leverages ds4_gpu_mgpu.h and ds4_gpu_args.c to distribute inference across multiple GPUs. The server automatically queues incoming requests and processes them using micro-batching strategies (tested via test_server_batching.py), maximizing hardware utilization without requiring client-side changes.
The request dispatch block in ds4_server.c handles concurrent connections asynchronously, while the web layer in ds4_web.c manages keep-alive connections and HTTP/1.1 persistent connections for optimal performance.
Summary
- DS4 provides a built-in HTTP server in
ds4_server.cthat implements OpenAI-compatible endpoints on port 9333 by default. - Three primary endpoints (
/v1/models,/v1/chat/completions,/v1/completions) enable drop-in replacement of OpenAI's hosted API. - Zero client modifications are required; existing OpenAI client libraries work by changing the base URL to
http://localhost:9333/v1. - Multi-GPU support via
ds4_gpu_mgpu.hand automatic batching provide enterprise-grade throughput for local deployments. - JSON schema compatibility ensures responses contain standard fields like
choices,usage, andfinish_reasonthat existing applications expect.
Frequently Asked Questions
What port does the DS4 server use by default?
The DS4 server listens on port 9333 by default when you run ./ds4-server. You can verify the server is active by visiting http://127.0.0.1:9333/v1/models in a browser or using curl to check model availability.
Do I need to modify my existing OpenAI Python code to use DS4?
No. You only need to change the base_url parameter in your OpenAI client configuration to point to http://localhost:9333/v1 (or your server's IP address). The DS4 endpoints in ds4_server.c return JSON structures that match the OpenAI specification exactly, including identical field names and response formats.
Which file handles the HTTP request routing in DS4?
Request routing is implemented in ds4_server.c around line 12526, where the server inspects the URL path and HTTP method to dispatch requests to the appropriate handler. The low-level HTTP protocol implementation resides in ds4_web.c at line 296, which handles socket management and response formatting.
Can DS4 handle multiple concurrent requests?
Yes. The server architecture supports micro-batching and multi-GPU inference through ds4_gpu_mgpu.h. Requests are automatically queued and processed in parallel batches when hardware permits, making DS4 suitable for production workloads requiring concurrent API access.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →