Performance Trade-offs Between Prefill and Generation in Distributed Mode

In DS4's distributed runtime, prefill processes the entire prompt in one bulk operation to populate the KV cache upfront, while generation decodes tokens one-by-one with higher per-step network overhead; choosing between them depends on prompt length, available GPU memory, and latency requirements.

DS4 is a lightweight LLM inference engine that supports distributed execution across multiple workers. When running in distributed mode, the inference pipeline splits into two distinct phases—prefill and generation—each with different performance characteristics that directly impact throughput, latency, and memory consumption. This article examines the trade-offs between these phases based on the actual implementation in the antirez/ds4 repository.

Prefill Phase: Bulk Prompt Processing

The prefill phase in DS4 handles the entire prompt (or a large "prefill-chunk") in a single forward pass. This populates the KV cache up-front, allowing subsequent generation steps to proceed without recomputing attention keys and values for the prompt tokens.

How Prefill Works in Distributed Mode

In ds4_distributed.c, the coordinator constructs a prefill work packet using the DS4_DIST_WORK_F_INPUT_HC flag. This flag indicates that the work message contains hidden-state input for the first token batch:

/* From ds4_distributed.c lines 55-60 */
#define DS4_DIST_WORK_F_INPUT_HC     (1<<0)  /* Hidden state input follows */
#define DS4_DIST_WORK_F_OUTPUT_LOGITS (1<<1)  /* Return logits */
#define DS4_DIST_WORK_F_ACK_ONLY      (1<<2)  /* Just ack, no compute */

All workers receive the same prefilling slice, copy the KV entries locally, and acknowledge completion. The KV layout is stored per-layer as defined in ds4_distributed.c lines 10-13.

Prefill Performance Characteristics

  • Latency: One-off large latency for the entire prompt, but subsequent generation steps are fast because the KV cache is already populated
  • Memory: Reserves KV memory for the entire prompt on every worker, increasing per-GPU memory usage proportionally to prompt length
  • Network: Single bulk transfer of hidden-state and KV data; cost is amortized over many tokens

The default prefill chunk size is 2048 tokens for CUDA tensor-parallel mode, as shown in tests/test_engine_mgpu_placement.c lines 549-555. This cap prevents unbounded KV allocation:

/* Demonstrated in test_engine_mgpu_placement.c */
int prefill_chunk = 2048;  /* Default to keep KV allocation reasonable */

Generation Phase: Token-by-Token Decoding

The generation (or decode) phase produces tokens one-by-one after the prompt is cached. Each step requires network coordination between the coordinator and workers.

How Generation Works in Distributed Mode

Generation work packets carry different flags than prefill. The DS4_DIST_WORK_F_OUTPUT_LOGITS flag requests logits output, while DS4_DIST_WORK_F_ACK_ONLY enables pure "ping-pong" coordination for synchronization without computation:

/* From ds4_distributed.c lines 62-65 */
#define DS4_DIST_RESULT_OK           0
#define DS4_DIST_RESULT_LOGITS       1   /* Logits follow */
#define DS4_DIST_RESULT_ERROR        2

For each decode step, the coordinator sends a work message containing recently generated token(s). Workers compute their slice and return results using the DS4_DIST_RESULT_LOGITS result kind.

Generation Performance Characteristics

  • Latency: Higher per-token latency because every token triggers a new round-trip across the network
  • Memory: Only KV entries for already-generated tokens stay resident, so memory footprint is lower than full-prompt prefill on long contexts
  • Network: Small, frequent messages; overhead becomes noticeable with more workers or higher network latency

Key Trade-offs and Decision Criteria

When to Use Prefill

Prefill-heavy workloads—such as long static prompts—benefit from bulk prefill because the one-time cost of copying KV data spreads across many subsequent tokens. The memory investment pays off when generating many tokens from the same prompt.

The server option DS4_TEST_SERVER_PREFILL=1 in tests/test_cuda_session_batch.c (lines 110-118) demonstrates how the runtime can engage a "progress-split prefill" path, allowing mixed strategies based on prompt length.

When to Use Generation-Only or Small Prefill Chunks

Interactive or short-prompt scenarios gain from avoiding the memory blow-up of a large KV cache. When memory is constrained or low-latency interactive response is required, accepting higher per-token network overhead becomes preferable.

Distributed-Mode Specific Considerations

Workers in DS4 are independent processes holding only their assigned layers. This architecture creates unique pressures:

  1. Prefill forces every worker to materialize hidden-state for the entire prompt, which can saturate the NIC with very large prompts
  2. Generation lets each worker stay idle between tokens, reducing sustained bandwidth but increasing per-token round-trip time

Practical Code Examples

Prefill Path Implementation

/* Prefill path - bulk hidden-state copy */
ds4_session *prefill = NULL;
ds4_session_create(&prefill, engine, ctx_len);
ds4_session_sync(prefill, prompt, err, sizeof(err));

/* After this the KV cache holds the whole prompt */

This mirrors the test harness in tests/test_metal_session_batch.c lines 284-317.

Generation Path Implementation

/* Generation (decode) path */
ds4_session *decode = NULL;
ds4_session_create(&decode, engine, ctx_len);

while (need_more_tokens) {
    ds4_session_eval(decode, &next_token, err, sizeof(err));
    /* Coordinator sends work packet with DS4_DIST_WORK_F_OUTPUT_LOGITS */
}

See tests/test_cuda_mixed_batch.c lines 166-175 for a mixed prefill + decode workflow used to compare latency and accuracy.

Core Implementation Files

File Purpose
ds4_distributed.c Core distributed transport, work-message flags, and result kinds
ds4_tp.c Tensor-parallel and distributed option parsing; shows mutual exclusivity of TP and distributed roles (lines 506-514)
ds4_server.c Coordinator orchestration, worker launch, and prefill path invocation via ds4_dist_prepare_engine_options
tests/test_engine_mgpu_placement.c Default prefill chunk sizing and memory impact demonstration
tests/test_cuda_mixed_batch.c Mixed prefill + decode workflow for latency/accuracy comparison

Summary

  • Prefill minimizes per-token latency for long prompts by absorbing network transfer costs upfront and populating the KV cache, at the expense of higher initial latency and memory usage
  • Generation reduces memory footprint and avoids large bulk transfers, but incurs network round-trip overhead for every token
  • The default 2048-token prefill chunk in DS4 provides a configurable balance between these extremes
  • Distributed mode amplifies these trade-offs because workers must coordinate over the network for every operation, making prompt length and worker count critical tuning parameters

Frequently Asked Questions

How does DS4 decide between prefill and generation mode?

DS4 selects the execution path based on the DS4_DIST_WORK_F_INPUT_HC flag in work messages. When this flag is set, workers execute the bulk prefill path with hidden-state input. Without it, they enter token-generation mode. The coordinator in ds4_server.c manages this orchestration through ds4_dist_prepare_engine_options.

What is the default prefill chunk size and why does it matter?

The default prefill chunk size is 2048 tokens for CUDA tensor-parallel execution. This limit prevents unbounded KV cache growth that would exhaust GPU memory on long prompts. Smaller chunks reduce peak memory usage but increase the number of prefill operations needed.

Can prefill and generation be mixed in the same session?

Yes. DS4 supports mixed workflows where a partial prefill populates the cache, followed by generation steps. The test file tests/test_cuda_mixed_batch.c demonstrates this pattern, and the DS4_TEST_SERVER_PREFILL environment variable enables configuration of progress-split prefill behavior.

How does network latency affect distributed inference performance?

Network latency has asymmetric impact: prefill tolerates latency well because transfer costs amortize across many tokens, while generation suffers linearly since every token requires a round-trip. High-latency networks favor larger prefill chunks or full-prompt prefill to minimize coordination steps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →