# Configuring ds4 Chat Mode: Server, Agent, and CLI Architecture for Tool Execution

> Configure ds4 chat mode with its CLI server and agent architecture. Learn how to define and execute tools for enhanced assistant capabilities with this unified interface.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: how-to-guide
- Published: 2026-08-08

---

**ds4 provides three distinct chat interfaces—an interactive CLI ([`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c)), an OpenAI-compatible HTTP server ([`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)), and a native coding agent ([`ds4_agent.c`](https://github.com/antirez/ds4/blob/main/ds4_agent.c))—that share a unified configuration interface through the core engine flags.**

The **antirez/ds4** repository implements a vertical inference stack where chat functionality is exposed through dedicated entry points built atop the single-file engine core ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)). Whether running local interactive sessions, hosting OpenAI-compatible API endpoints, or deploying autonomous coding agents, configuration relies on standardized parameters for context length, GPU backend selection, and KV-cache management.

## Interactive CLI Chat Mode

The primary command-line interface is implemented in **[`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c)**, providing direct user interaction with DeepSeek V4 Flash/Pro and GLM 5.2 models.

Configuration is handled entirely through command-line flags parsed by the CLI entry point:

```bash

# Basic interactive mode with resident GPU memory

./ds4 -m ./ds4flash.gguf --temp 0

# High-context coding session with SSD streaming

./ds4 -m ./ds4flash.gguf \
  --ctx 32768 \
  --ssd-streaming \
  --ssd-streaming-cache-experts 32GB \
  --nothink

```

Key flags include `--ctx` for context window sizing, `--temp` for sampling temperature, and backend selectors like `--cuda` or `--metal` (auto-detected on macOS).

## OpenAI-Compatible Server Mode

**[`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)** exposes an HTTP endpoint compatible with the OpenAI Chat Completions API, enabling tool-calling workflows via standard REST clients.

The server accepts identical engine flags as the CLI, with additional networking parameters:

```bash

# Start server on port 8080 with extended context

./ds4-server --ctx 100000 --host 0.0.0.0 --port 8080

```

Client interaction follows the standard schema:

```bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"Write a poem about AI"}],"max_tokens":128}'

```

The server implementation supports **batched-session workloads** through on-disk KV-cache storage. Configure the storage location using:

```bash
./ds4-server --kv-disk-dir ~/.ds4/kvcache --ctx 100000

```

This allows persistent conversation states across multiple HTTP requests without reloading the model.

## Native Coding Agent Mode

**[`ds4_agent.c`](https://github.com/antirez/ds4/blob/main/ds4_agent.c)** provides an autonomous coding agent with **KV-based session state** stored in `~/.ds4/kvcache`, optimized for long-running tool-augmented workflows.

Agent configuration extends the standard engine flags with multi-GPU and distributed inference support:

```bash

# Agent with tensor parallelism across 8 CUDA devices

./ds4-agent --cuda --cuda-tensor-parallel \
  --gpu-devices 0,2,4,6,1,3,5,7 \
  --model "gguf/DeepSeek-V4-Flash-Q4KExperts-…-imatrix-0731.gguf" \
  --ctx 100000

```

The agent maintains conversation context persistently on disk, enabling resumable coding sessions that survive process restarts.

## Common Configuration Parameters

All three chat modes—CLI, Server, and Agent—accept these core parameters defined in the engine ([`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c) and [`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h)):

- **`--ctx`**: Sets the maximum context length (e.g., `32768`, `100000`)
- **`--power`**: Controls performance/power trade-offs on mobile devices
- **`--prefill-chunk`**: Optimizes prompt processing throughput
- **`--temp`**: Sampling temperature (greedy decoding at `0`)
- **`--ssd-streaming`**: Activates expert tensor streaming for models exceeding GPU memory
- **`--role coordinator/worker`**: Enables distributed pipeline parallelism across machines

## Backend-Specific Chat Configuration

Chat mode performance depends on the selected backend compiled into the binary:

**Metal (macOS)**: Automatically selected when building on Apple Silicon. Supports resident GPU inference and automatic SSD-streaming activation.

**CUDA (Linux)**: Enabled with `make cuda`. Supports tensor parallelism via `--cuda-tensor-parallel` and device selection via `--gpu-devices`.

**ROCm (Linux/AMD)**: Compiled with `DS4_ROCM_BUILD` defined. Mirrors CUDA functionality for AMD Strix Halo and Framework Desktop hardware.

## Summary

- **Three entry points**: `ds4` (CLI), `ds4-server` (HTTP API), and `ds4-agent` (autonomous coding) provide flexible chat interfaces.
- **Unified configuration**: All modes share engine flags (`--ctx`, `--ssd-streaming`, `--temp`) defined in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c).
- **OpenAI compatibility**: The server ([`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)) exposes standard chat completion endpoints suitable for function-calling integrations.
- **Persistent state**: Both server and agent support disk-backed KV caches (`--kv-disk-dir`) for multi-turn conversations.
- **Hardware flexibility**: Identical chat configurations work across Metal, CUDA, and ROCm backends with appropriate compilation flags.

## Frequently Asked Questions

### How do I enable tool calling in ds4 chat mode?

The OpenAI-compatible server ([`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)) exposes the `/v1/chat/completions` endpoint, which accepts standard JSON payloads including tool definitions. Configure the server with `./ds4-server --ctx 100000`, then send requests containing the `tools` array following the OpenAI API specification. The engine processes these through the chat completion pipeline implemented in the server module.

### What is the difference between ds4-agent and ds4-server?

**`ds4-agent`** ([`ds4_agent.c`](https://github.com/antirez/ds4/blob/main/ds4_agent.c)) is a native coding agent with persistent KV-based session state stored in `~/.ds4/kvcache`, designed for autonomous tool execution workflows. **`ds4-server`** ([`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)) provides a stateless (or optionally stateful with `--kv-disk-dir`) HTTP API compatible with OpenAI clients. Both use the same core inference engine but target different integration patterns—standalone automation versus API service.

### How do I configure multi-GPU chat inference?

Pass `--cuda-tensor-parallel` (or `--rocm-tensor-parallel`) along with `--gpu-devices` listing the device IDs. For example: `./ds4-agent --cuda-tensor-parallel --gpu-devices 0,1,2,3 --ctx 100000`. This splits transformer layers across the specified GPUs while maintaining a unified chat interface.

### Where does ds4 store chat conversation history?

The CLI maintains history in memory only. The server and agent persist **KV-cache states** (not raw text) to disk when `--kv-disk-dir` is specified, typically defaulting to `~/.ds4/kvcache`. This allows resuming conversations without recomputing prompt embeddings, as implemented in [`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c) and [`ds4_agent.c`](https://github.com/antirez/ds4/blob/main/ds4_agent.c).