Configuring ds4 Chat Mode: Server, Agent, and CLI Architecture for Tool Execution

ds4 provides three distinct chat interfaces—an interactive CLI (ds4_cli.c), an OpenAI-compatible HTTP server (ds4_server.c), and a native coding agent (ds4_agent.c)—that share a unified configuration interface through the core engine flags.

The antirez/ds4 repository implements a vertical inference stack where chat functionality is exposed through dedicated entry points built atop the single-file engine core (ds4.c). Whether running local interactive sessions, hosting OpenAI-compatible API endpoints, or deploying autonomous coding agents, configuration relies on standardized parameters for context length, GPU backend selection, and KV-cache management.

Interactive CLI Chat Mode

The primary command-line interface is implemented in ds4_cli.c, providing direct user interaction with DeepSeek V4 Flash/Pro and GLM 5.2 models.

Configuration is handled entirely through command-line flags parsed by the CLI entry point:


# Basic interactive mode with resident GPU memory

./ds4 -m ./ds4flash.gguf --temp 0

# High-context coding session with SSD streaming

./ds4 -m ./ds4flash.gguf \
  --ctx 32768 \
  --ssd-streaming \
  --ssd-streaming-cache-experts 32GB \
  --nothink

Key flags include --ctx for context window sizing, --temp for sampling temperature, and backend selectors like --cuda or --metal (auto-detected on macOS).

OpenAI-Compatible Server Mode

ds4_server.c exposes an HTTP endpoint compatible with the OpenAI Chat Completions API, enabling tool-calling workflows via standard REST clients.

The server accepts identical engine flags as the CLI, with additional networking parameters:


# Start server on port 8080 with extended context

./ds4-server --ctx 100000 --host 0.0.0.0 --port 8080

Client interaction follows the standard schema:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"Write a poem about AI"}],"max_tokens":128}'

The server implementation supports batched-session workloads through on-disk KV-cache storage. Configure the storage location using:

./ds4-server --kv-disk-dir ~/.ds4/kvcache --ctx 100000

This allows persistent conversation states across multiple HTTP requests without reloading the model.

Native Coding Agent Mode

ds4_agent.c provides an autonomous coding agent with KV-based session state stored in ~/.ds4/kvcache, optimized for long-running tool-augmented workflows.

Agent configuration extends the standard engine flags with multi-GPU and distributed inference support:


# Agent with tensor parallelism across 8 CUDA devices

./ds4-agent --cuda --cuda-tensor-parallel \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --model "gguf/DeepSeek-V4-Flash-Q4KExperts-…-imatrix-0731.gguf" \
  --ctx 100000

The agent maintains conversation context persistently on disk, enabling resumable coding sessions that survive process restarts.

Common Configuration Parameters

All three chat modes—CLI, Server, and Agent—accept these core parameters defined in the engine (ds4.c and ds4.h):

  • --ctx: Sets the maximum context length (e.g., 32768, 100000)
  • --power: Controls performance/power trade-offs on mobile devices
  • --prefill-chunk: Optimizes prompt processing throughput
  • --temp: Sampling temperature (greedy decoding at 0)
  • --ssd-streaming: Activates expert tensor streaming for models exceeding GPU memory
  • --role coordinator/worker: Enables distributed pipeline parallelism across machines

Backend-Specific Chat Configuration

Chat mode performance depends on the selected backend compiled into the binary:

Metal (macOS): Automatically selected when building on Apple Silicon. Supports resident GPU inference and automatic SSD-streaming activation.

CUDA (Linux): Enabled with make cuda. Supports tensor parallelism via --cuda-tensor-parallel and device selection via --gpu-devices.

ROCm (Linux/AMD): Compiled with DS4_ROCM_BUILD defined. Mirrors CUDA functionality for AMD Strix Halo and Framework Desktop hardware.

Summary

  • Three entry points: ds4 (CLI), ds4-server (HTTP API), and ds4-agent (autonomous coding) provide flexible chat interfaces.
  • Unified configuration: All modes share engine flags (--ctx, --ssd-streaming, --temp) defined in ds4.c.
  • OpenAI compatibility: The server (ds4_server.c) exposes standard chat completion endpoints suitable for function-calling integrations.
  • Persistent state: Both server and agent support disk-backed KV caches (--kv-disk-dir) for multi-turn conversations.
  • Hardware flexibility: Identical chat configurations work across Metal, CUDA, and ROCm backends with appropriate compilation flags.

Frequently Asked Questions

How do I enable tool calling in ds4 chat mode?

The OpenAI-compatible server (ds4_server.c) exposes the /v1/chat/completions endpoint, which accepts standard JSON payloads including tool definitions. Configure the server with ./ds4-server --ctx 100000, then send requests containing the tools array following the OpenAI API specification. The engine processes these through the chat completion pipeline implemented in the server module.

What is the difference between ds4-agent and ds4-server?

ds4-agent (ds4_agent.c) is a native coding agent with persistent KV-based session state stored in ~/.ds4/kvcache, designed for autonomous tool execution workflows. ds4-server (ds4_server.c) provides a stateless (or optionally stateful with --kv-disk-dir) HTTP API compatible with OpenAI clients. Both use the same core inference engine but target different integration patterns—standalone automation versus API service.

How do I configure multi-GPU chat inference?

Pass --cuda-tensor-parallel (or --rocm-tensor-parallel) along with --gpu-devices listing the device IDs. For example: ./ds4-agent --cuda-tensor-parallel --gpu-devices 0,1,2,3 --ctx 100000. This splits transformer layers across the specified GPUs while maintaining a unified chat interface.

Where does ds4 store chat conversation history?

The CLI maintains history in memory only. The server and agent persist KV-cache states (not raw text) to disk when --kv-disk-dir is specified, typically defaulting to ~/.ds4/kvcache. This allows resuming conversations without recomputing prompt embeddings, as implemented in ds4_server.c and ds4_agent.c.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →