Using the Native ds4-agent for Local Coding Tasks: A Complete Guide

The native ds4-agent runs local inference in-process with ultra-low latency, persisting KV-cache sessions to disk at ~/.ds4/kvcache for seamless context restoration.

The ds4 repository by Salvatore Sanfilippo (antirez) implements DwarfStar, a compact native inference engine optimized for DeepSeek V4 and GLM 5.2 models. When using the native ds4-agent for local coding tasks, you eliminate network overhead entirely—the agent drives inference directly from the same process, making it ideal for iterative development workflows that demand immediate feedback.

What is the Native ds4-agent?

Unlike the HTTP server mode (ds4-server), the native agent embeds the inference loop directly into your terminal session. Located in ds4_agent.c, this component shares the same backend-agnostic graph API as the CLI and server, but removes all socket boundaries between the user interface and the inference engine.

The agent maintains its KV-cache on local disk rather than in GPU memory alone, enabling you to save, list, and restore conversation sessions across restarts. This architecture ensures that your coding context persists exactly as you left it, with zero serialization overhead or API compatibility layers.

Key Benefits of the Native Agent

Running ds4-agent provides several distinct advantages for local coding workflows:

  • Ultra-low latency – Pre-fill and generation phases are limited only by your backend speed (Metal, CUDA, or ROCm), with no network round-trips.
  • Live progress bar – Visual feedback during the pre-fill phase keeps you informed of context processing.
  • Native tool handling – No DSML conversion required; tool calls execute directly within the agent process.
  • KV-cache consistency – Session state mismatches are impossible because the cache is the session state, stored persistently under ~/.ds4/kvcache.

Launching the ds4-agent

You can start the agent using the compiled binary or via helper scripts for distributed setups.

For standard single-GPU or Metal inference:

./ds4-agent

For tensor-parallelism across multiple NVIDIA GPUs, use the provided helper:

./run-nvidia-tp-agent.sh

Once launched, the agent presents an interactive ds4> prompt where you can enter coding queries or control session persistence.

Managing KV-Cache Sessions

The agent provides three primary commands for session management, allowing you to treat long coding contexts as persistent workspaces:

  • /save – Persists the current KV-cache to ~/.ds4/kvcache with a unique SHA identifier.
  • /list – Displays all available saved sessions with their identifiers and metadata.
  • /switch <sha> – Loads a previously saved KV-cache, restoring the exact conversation state.

This workflow is particularly valuable when working with large codebases across multiple files. You can /save after ingesting a complex repository structure, then /switch back to that context days later without re-processing the source code.

Architecture: Agent vs. Server

Understanding the distinction between ds4_agent.c and ds4_server.c helps you choose the right tool for your workflow:

ds4_agent.c (Native Agent)

  • Runs inference in the same process as the UI
  • Maintains a single mutable KV-state backed by disk storage
  • No HTTP overhead; direct function calls to the engine in ds4.c
  • Ideal for personal coding sessions requiring maximum responsiveness

ds4_server.c (HTTP Server)

  • Exposes OpenAI-compatible endpoints at /v1/chat/completions
  • Supports batched sessions (--batched-session N) for multi-user scenarios
  • Maintains KV-state in memory or on disk depending on configuration
  • Better suited for IDE integrations or shared development environments

Both components utilize the same core inference graph built by ds4.c, ensuring identical output quality regardless of interface.

Practical Usage Examples

Interactive Coding Session with Persistence


# Launch the agent

./ds4-agent

# Inside the agent console

ds4> /list
ds4> Analyze this Python function for potential race conditions: [paste code]

# ... generation appears with live progress bar ...

ds4> /save
Session saved to ~/.ds4/kvcache/<sha>

Tensor Parallelism with the Agent

For large models requiring multi-GPU setup:


# Worker (first machine)

./ds4 -m gguf/model.gguf --tensor-parallel --role worker \
  --coordinator 10.0.0.5 9911 --transport rdma

# Coordinator with agent interface

./ds4-agent --tensor-parallel --role coordinator \
  --listen 10.0.0.5 9911 --transport rdma

SSD-Streaming for Large Models

When your model exceeds available RAM, combine the agent with SSD-streaming:

./ds4-agent -m ./ds4flash.gguf \
  --ssd-streaming \
  --ssd-streaming-cache-experts 32GB

The agent will stream routed MoE experts from the GGUF file on cache miss while maintaining your conversation state on disk.

Summary

  • The ds4-agent (ds4_agent.c) provides the lowest-latency interface to the DwarfStar engine by running inference in-process.
  • Session persistence is handled via /save, /list, and /switch commands, with data stored in ~/.ds4/kvcache.
  • Unlike the HTTP server (ds4_server.c), the agent eliminates API boundaries, making it optimal for local coding tasks.
  • The agent supports all backend optimizations including tensor parallelism, SSD-streaming, and DSpark speculative decoding.
  • Launch via ./ds4-agent or ./run-nvidia-tp-agent.sh for distributed GPU setups.

Frequently Asked Questions

What is the difference between ds4-agent and ds4-server?

The ds4-agent runs the inference loop inside your terminal process, eliminating network latency and providing direct control over KV-cache sessions. The ds4-server (ds4_server.c) exposes an HTTP API compatible with OpenAI's specification, designed for remote clients or multi-user batched sessions. Use the agent for personal coding; use the server for IDE plugins or team environments.

Where does the ds4-agent store conversation history?

The agent persists KV-cache data to ~/.ds4/kvcache/ on your local disk. Each /save command creates a new entry with a unique SHA hash that you can list with /list and restore using /switch <sha>. This storage survives process restarts and system reboots.

Can I use the agent with multiple GPUs?

Yes. For tensor-parallelism across NVIDIA GPUs, launch the agent using ./run-nvidia-tp-agent.sh. For distributed setups across multiple machines, start worker processes with --role worker and the coordinator with ./ds4-agent --role coordinator over RDMA or TCP transport.

Does the agent support models larger than my GPU memory?

Yes, through SSD-streaming. Launch the agent with the --ssd-streaming flag to enable out-of-RAM model support. The engine keeps non-routed weights resident while streaming routed MoE experts from the GGUF file on demand, configurable via --ssd-streaming-cache-experts to limit disk cache size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →